The RI Safety Layer records behavioural evaluation as inspectable evidence, supports deterministic re-derivation from fixed records, and keeps publication governance distinct from interpretation and human action.
Designed for bounded technical diligence where evidence, derivation, interpretation, and publication decisions need clear boundaries.
The current implementation extends beyond measurement and publication governance to bounded interpretation, review, and human-stewardship surfaces while preserving those authority separations.
AI evaluation exists — but trust is fragile
Across labs, enterprises, and regulators, the same issues appear:
Evaluation results are difficult to verify
Live model outputs may vary across repeated runs
Metrics can be difficult to trace back to recorded evidence
Decisions about inclusion can be informal or opaque
The result: evaluation that may be useful internally, but difficult to inspect or challenge externally.
A new layer for evaluation integrity
The RI Safety Layer introduces a structured pipeline:
Behaviour is recorded as evidence
Fixed recorded evidence can be deterministically re-derived under the same evaluation configuration
Publication state is governed separately from measurement and interpretation
Evaluation is recorded, checked for integrity, derived under explicit methods, and subject to separate publication governance.
These implemented capabilities support bounded technical diligence. They do not by themselves establish institutional readiness, independent external validation, certification, or compliance sufficiency.
How It Works
Behavioural Measurement
Record and measure behaviour
Runs structured evaluation sessions
Records interaction evidence
Produces evidence and derived measurement outputs with explicit provenance
Outcome: A recorded evidential basis from which evaluation outputs can be inspected and re-derived
Publication Governance
Decide what is allowed to count
Applies explicit publication rules to eligible evidence and outputs
Records governance decisions
Controls inclusion in metrics and dashboards
Outcome: Publication state is governed without rewriting the underlying measurement
Bounded Interpretation, Review & Stewardship
Keep derived meaning and human action separate from evidence authority
Supports derived interpretation and comparison without promoting them into canonical evidence
Supports controlled review and inspection routes
Records bounded human action without automatically creating publication authority
Outcome: Interpretation and human action remain traceable, bounded, and distinct from measurement and governance
These surfaces have different authority. A derived interpretation, reviewer view, or human record does not rewrite canonical evidence or an already-computed metric.
Measurement is preserved.
Interpretation is bounded.
Governance is explicit.
Human action does not silently rewrite evidence.
Foundational Evaluation / Governance Flow
This diagram shows the foundational evidence-to-publication flow. It does not represent every current derived interpretation, review, or stewardship surface.
Capabilities
Evidence integrity and traceability
Where the relevant evidence is available, inspect:
what was recorded
how it was evaluated
the integrity and provenance checks attached to that record
Deterministic re-derivation
Given fixed recorded evidence and the same evaluation configuration:
re-run the evaluation derivation
reproduce deterministic measurement outputs
localise disagreement to evidence, method, criteria, or interpretation
This does not claim that live model generation will reproduce the same interaction on a new run.
Governed publication
Control whether eligible results:
enter dashboards
affect trends
are exposed to stakeholders
Traceable governance records
Governance decisions can be recorded with:
structured decision state
provenance to the relevant evidence and rules
separation from the underlying measurement record
These capabilities support credible bounded technical diligence; they are not by themselves evidence of operational readiness, certification, compliance sufficiency, or independent external validation.
Robustness Against Superficial Optimisation
The architecture can evaluate behaviour across varied probe families, preserved evidence traces,
and explicit criteria rather than relying on a single score. This is a design property intended to
make evaluation less dependent on single-score optimisation; it is not evidence that benchmark gaming
is eliminated or that durable behavioural robustness has been empirically established.
A simple, inspectable flow
Run an evaluation session
Record the interaction and supporting evidence
Apply the defined evaluation method
Check the available integrity and provenance signals
Apply publication governance separately
Include or hold the eligible result in rollups
Recorded evidence, derived interpretation, human judgement, and publication authority remain distinct. Verification is bounded to the evidence and checks actually available.
These are contexts in which the approach may be examined. They are not claims of present institutional approval or deployment readiness.
AI Labs
Inspect recorded evaluation evidence
Re-derive bounded measurement outputs
Apply explicit publication governance
Enterprises
Conduct bounded technical diligence
Examine internal metric-governance design
Assess separate context-specific assurance and compliance requirements
Regulators & Oversight Bodies
Inspect the public evaluation and governance framing
Examine available evidence where access is authorised
Assess claims against their stated boundaries
Research Partners
Re-derive results from fixed recorded evidence
Examine criteria and assumptions
Compare bounded outputs across defined conditions
Clear boundaries
The RI Safety Layer does not:
modify model outputs
guarantee correctness
replace training or alignment
intervene in real-time behaviour at its current stage
establish independent external validation, certification, or institution-specific compliance sufficiency
It provides a bounded measurement, evidence, interpretation, and governance framework; it does not establish that models are inherently safe or that an institution can rely on the system operationally without separate assessment.
Evaluation is becoming infrastructure
As AI systems move into:
regulated environments
enterprise decision-making
public-facing applications
there is increasing demand for:
accountability
traceability
reproducibility of evaluation from fixed evidence
The RI Safety Layer is designed to support these needs through inspectable evidence and explicit governance. Sufficiency for any particular regulated, enterprise, or assurance context requires separate assessment.
A broader current system, with preserved boundaries
The RI Safety Layer is modular by design.
Its current implementation includes bounded surfaces for:
Measurement and evidence — what was recorded and how it was evaluated
Publication governance — what is allowed to count publicly
Derived interpretation and comparison — advisory meaning that remains distinct from canonical evidence
Controlled review and human stewardship — inspection and action records that do not silently rewrite measurement or governance authority
These responsibilities are deliberately separated.
Real-time behavioural intervention, prompt rewriting, and inference-time control are not part of the current system.
Current breadth does not turn supportability into readiness: implemented surfaces remain bounded by their evidence and authority.