Peer Review Is Not Authorization
A FERZ Technical Note
v1.0 | July 31, 2026
© 2026 FERZ, Inc. Licensed under CC BY 4.0.
Abstract. In July 2026, a widely reported proposal placed peer review at the center of frontier AI release governance: developers would examine one another's models before release, confer regularly on safety and security issues, and alert government when a serious concern went unaddressed. This note examines that proposal as a recent, high-profile instance of release governance by review and asks a question prior to any judgment about feasibility: when review concludes, who has authorized what? Measured against the Five Tests Standard (5TS), the proposal as publicly described supplies evidence but does not establish any of the five tests. The note identifies the conditions under which review findings become bound inputs to authorization rather than substitutes for it.
1. The Proposal
In an interview with The Economist's editor-in-chief, Zanny Minton Beddoes, published on July 23, 2026, Elon Musk set out his preferred near-term mechanism for governing frontier AI releases. In substance: the leading developers would confer regularly on safety and security issues; competitors would receive a week or two of early access to major new models ahead of release; mutual scrutiny among rivals would serve as the honesty mechanism; and where a developer declined to address a serious concern raised by its peers, government would be alerted and could act to constrain the release. Asked about timing, he argued that the arrangement should exist without delay.
This note takes the proposal seriously and reads it on its own terms. On its own terms, it is a contribution to hazard discovery, and hazard discovery matters. The note also treats the proposal as a recent, high-profile instance of release governance by review. Related mechanisms include red-team exchanges, staged-access programs, third-party evaluations, model documentation, and disclosure regimes. The analysis below extends to arrangements in which those mechanisms generate evidence without separately defined authority, policy, pre-execution verdict semantics, and enforcement consequences.
The question the note asks is narrow and prior to any judgment about the proposal's feasibility. When the review concludes, who has authorized what?
2. Two Functions That Resemble Each Other
A review finding is a statement about a system: what it did under test conditions, what capabilities it appears to have, what risks the reviewers judge material. An authorization is a determination about a specific proposed action: that this release, by this actor, under this policy and this authority state, may proceed. The two differ in object, function, and output. A finding characterizes a system, capability, or risk on the basis of testing or analysis; it yields evidence. An authorization evaluates a specific proposed action under the policy and authority state in force; it yields a pre-execution verdict.
Review as review produces findings. A review body may also participate in authorization, but only where its authority, applicable policy, verdict semantics, and execution consequences are separately defined. Hazard discovery does not, by itself, determine whether a specific release or action is permitted. As publicly described, the proposal specifies participants and evidence flows but does not specify that second function. It identifies no defined point at which an authorization determination issues, no final authority from which it issues, and no authorization artifact in which it is preserved.
Governance regimes commonly blur this line, because both functions produce documents and both involve judgment. The blur is consequential. Findings can accumulate while releases proceed. A non-bypassable authorization boundary instead makes release unavailable absent an authorizing verdict.
3. The Five Concepts the Proposal Leaves Open
"Authority versus Authorization: A Definitional Framework for AI Governance" (DOI: 10.5281/zenodo.21341907) distinguishes five concepts that any complete governance arrangement must resolve: Authority, who may decide, a property of actors; Delegation, who may exercise that authority; Policy, under what conditions; Authorization, whether a specific proposed action may proceed, determined before execution under applicable policy; and Enforcement, whether the action can proceed without that determination. The framework's central observation: "Authority determines who may authorize. Authorization determines whether a specific proposed action is permitted. Enforcement determines whether that authorization governs execution."
Held against these five concepts, the proposal identifies candidate participants and evidence flows but does not resolve the allocation required for authorization. It does not specify who holds final authorization authority, what authority is delegated, what policy governs a release verdict, what verdict function applies, or what mechanism prevents release absent that verdict.
A recurring position in frontier-safety discussion states the objective dispositionally: advanced systems should be truthful, curious, and well intentioned toward humanity. Dispositional quality is a property of the actor. It may inform policy or risk assessment, but does not by itself resolve any of the five concepts. Truth-seeking is a behavioral objective. It does not establish that a particular action was permitted before execution. A truthful system can execute an action that exceeds its authority, and will then describe the action truthfully.
4. The Proposal Under the Five Tests
The Five Tests Standard (5TS v1.2.0; standard/repository concept DOI: 10.5281/zenodo.21040295) is a published, vendor-neutral evaluation standard developed by FERZ for determining whether a governance arrangement governs execution. Its five tests are Stop, Ownership, Replay, Escalation, and Provenance. The five tests are operational criteria, and they are distinct from the five definitional concepts of Section 3; the two sets should not be conflated.
Stop asks whether the governed action can be halted before execution absent authorization. The proposal describes early access and possible government intervention but specifies no mandatory pre-release gate or fail-closed rule. It therefore does not establish the Stop test.
Ownership asks who holds final authorization authority and how override operates. The proposal names participants but does not specify final authority, threshold rules, or override relationships.
Replay asks whether the determination can be independently reconstructed afterward from the inputs, policy state, authority chain, and proposed action in force at the time. The proposal specifies neither an authorization determination nor an artifact sufficient to reconstruct one.
Escalation asks how the arrangement resolves indeterminacy. In the proposal, an unresolved concern becomes a report to government. Reporting carries no defined execution semantics within the proposal: nothing specified keeps the model from shipping while the report is under consideration. A sound authorization regime resolves indeterminacy fail-closed. In the FERZ verdict space, the corresponding disposition is ABSTAIN, and ABSTAIN blocks execution pending authorized human override. The difference between a report and an ABSTAIN is the difference between informing someone and stopping something.
Provenance asks whether the inputs grounding a determination have established origin, version, and binding. Provenance concerns origin, not truth. The proposal does not specify how review findings are versioned, which findings are bound to a release determination, or how their origin is established. Evidence Projection (version DOI: 10.5281/zenodo.20277842) states: "Authorization that cannot establish the provenance of its inputs is authorization in form, not in substance." The Authorization Boundary Integrity Model (DOI: 10.5281/zenodo.20929115) states the corresponding Input Integrity requirement directly: no verdict may rest on inputs whose origin is not established.
The result is uniform. Peer review may provide relevant evidence, but none of the five tests is established by the proposal alone.
5. Norms Bind Differently Than Boundaries
Review mechanisms that lack allocated authority and an enforced execution boundary bind through reputation, reciprocity, and goodwill. Those forces are real and may be strong among competitors. But where collegial pressure is the mechanism's only binding force, the mechanism remains a norm. Norms shape conduct; they do not gate execution. The measure of the difference is the hard case: the moment when a participant, facing commercial pressure, concludes that this release should proceed despite the concern. A norm makes that choice costly. A boundary makes it unavailable. Governance must be built for the hard case, because the easy cases were never the problem.
A standard objection notes that no single actor can prevent the rest of the world from using a model. The objection is correct as a limit on universal unilateral reach, and it is a closed-world observation ("The Closed-World Bargain," DOI: 10.5281/zenodo.21643658). Enforceable technical governance is authority-bound and fail-closed within a defined execution domain: the infrastructure, credentials, and downstream systems a deployment actually controls. Within that domain, runtime authorization can govern execution. Outside it, the architecture makes no claim of technical control; legal, commercial, diplomatic, and other institutional mechanisms may still operate.
6. What Evidence Requires to Enter Authorization
None of the foregoing diminishes review. It locates it. Review findings become bound inputs to authorization when five conditions hold. Their origin and version are established, and the findings are bound to the proposed action and determination. A defined authority evaluates them under an applicable, versioned policy before execution. The evaluation issues a verdict from a closed verdict space: ALLOW, DENY, or ABSTAIN, with DENY terminal for the requested action and ABSTAIN blocking execution pending authorized human override. A non-bypassable runtime authorization boundary fails closed, so the release cannot proceed absent ALLOW. The determination is preserved in a tamper-evident authorization artifact from which the verdict can be independently reconstructed without access to the governed system. In the vendor-neutral 5TS register, the required proof object is a Proof-Carrying Decision. FERZ’s authorization artifact is an implementation of the proof-carrying decision object required by 5TS.
Under those conditions, a competitor's finding remains evidence, but becomes a bound input that applicable policy may admit. It is evaluated within a process with allocated authority, applicable policy, a pre-execution verdict, an enforced runtime authorization boundary, and an independently reconstructable record.
This is the position argued across the FERZ doctrine corpus: observation, however sophisticated, does not authorize ("On the Impossibility of Observability-Based Authorization," DOI: 10.5281/zenodo.19647542; "From Monitoring to Authorization," DOI: 10.5281/zenodo.18743974). This note applies to peer review the argument previously developed for language-model evaluation ("LLM-as-a-Judge Is Not Authorization," DOI: 10.5281/zenodo.21462978): an instrument that produces judgments about systems is not thereby an instrument that authorizes actions.
7. Conclusion
The July proposal defines a potentially useful evidence process. It identifies reviewers, a mechanism for surfacing adverse findings, and a governmental escalation path. It does not specify the layer that converts evidence into a governed release: allocated authority, applicable policy, a pre-execution verdict, a non-bypassable fail-closed boundary, and an authorization artifact from which the verdict can be independently reconstructed. That layer is not a refinement of review. It is a different function.
The note's claim is correspondingly narrow. A review finding remains evidence. It becomes a bound input to authorization only when a defined authority evaluates it under policy, before execution, at a non-bypassable fail-closed boundary, and preserves the verdict in an independently reconstructable authorization artifact.
Peer review identifies the hazard. Authorization determines whether the model may be released or the action may execute, under whose authority, and with what independently reconstructable authorization artifact.
References
- FERZ, Inc. "Authority versus Authorization: A Definitional Framework for AI Governance," v1.0, July 2026. Concept DOI: 10.5281/zenodo.21341907.
- FERZ, Inc. "The Five Tests Standard (5TS)," v1.2.0. Version DOI: 10.5281/zenodo.21040296; standard/repository concept DOI: 10.5281/zenodo.21040295.
- FERZ, Inc. "On the Impossibility of Observability-Based Authorization." Concept DOI: 10.5281/zenodo.19647542.
- FERZ, Inc. "Authorization Boundary Integrity Model (ABIM)," v1.0. Concept DOI: 10.5281/zenodo.20929115.
- FERZ, Inc. "From Monitoring to Authorization." DOI: 10.5281/zenodo.18743974.
- FERZ, Inc. "The Closed-World Bargain," v1.0, July 2026. Version DOI: 10.5281/zenodo.21643659; concept DOI: 10.5281/zenodo.21643658.
- FERZ, Inc. "LLM-as-a-Judge Is Not Authorization," July 2026. Version DOI: 10.5281/zenodo.21462979; concept DOI: 10.5281/zenodo.21462978.
- FERZ, Inc. "Evidence Projection." Version DOI: 10.5281/zenodo.20277842; concept DOI: 10.5281/zenodo.20277841.
- The Economist. "The full-length interview with Elon Musk" (video; interview by Zanny Minton Beddoes), published July 23, 2026. https://www.youtube.com/watch?v=XuoqKYxDHVc. Proposal discussion at approx. 21:43 to 25:37; dispositional-safety discussion at approx. 05:55 to 06:27.
- Reuters (Akash Sriram; ed. Anil D'Silva). "Musk proposes peer review for frontier AI models in Economist interview," July 23, 2026. https://www.reuters.com/legal/litigation/musk-proposes-peer-review-frontier-ai-models-economist-interview-2026-07-23/
- Marco Quiroz-Gutierrez, Fortune. "Musk says frontier AI models should face peer review from rival labs before release, with the government only stepping in as a last resort," July 23, 2026. https://fortune.com/article/elon-musk-says-rival-ai-labs-should-peer-review-frontier-models-before-release-with-government-as-last-resort/
About FERZ, Inc. FERZ, Inc. is building runtime authorization infrastructure for AI systems, designed so that each governed action is evaluated before execution and produces a tamper-evident authorization artifact. Doctrine, standards, and research: https://ferz.ai.
DOI: 10.5281/zenodo.21722297 (concept, all versions) · 10.5281/zenodo.21722298 (this version, v1.0)
Canonical URL: https://ferz.ai/articles/peer-review-is-not-authorization
