SAFE (Shared AI Findings Exchange) is a proposed independent incident-learning and assurance initiative of the Open Secure AI Alliance. It establishes a confidential reporting and analysis framework for AI incidents and near misses — when AI systems access, exploit, or damage third-party systems without authorization. Members agree to report incidents, preserve evidence, submit to multi-layer technical reviews, and publish remediation status on fixed timelines.
The RFC is framed as a voluntary community initiative — a "compact" among AI developers, deployers, and customers. But the Linux Foundation's backing, the structured compliance timelines (72-hour, 4-day, 14-day, 30-day, 90-day), and the evidence preservation requirements create a framework that reads like a regulation. The question is not whether SAFE is mandatory today, but how quickly it becomes a procurement requirement — the same trajectory every "voluntary" AI framework in this investigation has followed.
Cyberattacks will happen. AI systems will make mistakes. Safeguards and operating assumptions will sometimes fail. A responsible AI ecosystem is measured by how quickly it contains harm, establishes the facts, informs those at risk and turns each incident into stronger protection for the community.
The framing establishes a premise that justifies intervention. "Safeguards will sometimes fail" — true, but the solution proposed is a structured, membership-gated reporting framework run by a foundation controlled by its largest corporate sponsors. The premise is correct; the proposed remedy is where the leverage lies.
Note who defines "responsible." The Linux Foundation and its corporate members — predominantly large cloud and AI providers — define what constitutes a "responsible AI ecosystem." Individual developers and small companies are not at the table when these definitions are written. The framework's moral authority is used to establish structural authority.
SAFE confidentially collects and analyzes AI incidents and near misses, promptly informs affected parties and turns recurring failures into shared, evidence-based controls that reduce systemic risk.
"Shared, evidence-based controls" is the operative phrase. The mission is not just information sharing — it's producing controls that become the reference standard. Once SAFE publishes "recommended" controls (default-deny network egress, signed evaluation manifests, etc.), these become the baseline that procurement teams, insurers, and regulators cite. The controls are written by the members who have the resources to participate in the review process — large AI providers.
Who benefits from "systemic risk reduction"? The same large providers who staff the framework. A shared control baseline that they already meet (because they wrote it) becomes a barrier to smaller competitors who must now implement equivalent controls to participate in the ecosystem. This is regulatory capture through technical standards.
SAFE should include representatives from:
SAFE should operate independently so that no vendor or industry segment controls its findings. Its processes apply equally to open and closed AI systems. Open systems are not automatically safe, and closed systems are not safe by declaration. Trust is not a control; shared evidence and verifiable improvement are how trust is earned.
The membership list looks inclusive — but inclusion is not equality. "Cloud and tool providers" (AWS, Google, Microsoft, Anthropic, OpenAI) have dedicated compliance teams, legal counsel, and security operations centers. "Independent security researchers" and "civil-society representatives" do not. The framework claims equal treatment, but the compliance burden falls hardest on those with the fewest resources to bear it.
"Government and standards bodies as non-controlling observers" is the thin edge. Today they observe. Tomorrow they mandate. This is the standard trajectory: government participates in voluntary frameworks as an observer, then references the framework in procurement requirements, then codifies it in regulation. The same pattern is visible in NIST AI RMF (voluntary → referenced in EO 14409 procurement) and in every framework we track under Track 02: The Compliance Trap.
"Open systems are not automatically safe" — true, but asymmetric. The statement applies equally to open and closed systems in principle, but in practice the compliance infrastructure SAFE builds will be easier for closed, cloud-hosted AI providers to satisfy. An open-model organization releasing weights on HuggingFace cannot easily "preserve and provide" logs, tool calls, and agent identities for every downstream deployment. Cloud providers can — they control the infrastructure. The equal-treatment language obscures an unequal compliance reality.
"Member sovereignty" sounds protective but is a large-vendor feature. "Without superseding members' internal security policies" means a member's existing practices are grandfathered. Large cloud providers already have extensive internal security policies — SAFE's "minimum interoperability and assurance practices" sit on top of what they already do. Small providers and open-source projects must build new compliance infrastructure from scratch.
"Learning is separate from enforcement" — until it isn't. The principle says regulators "retain their legal rights." This means regulators can use SAFE's confidential incident reports as evidence in enforcement actions. The separation is aspirational, not structural. Once the framework exists, the information it collects becomes a resource for every regulator that chooses to reference it.
As a condition of membership, members agree to report an incident when they become aware, or reasonably suspect, that an AI system they operate:
Intent does not determine whether an event is reportable. Believing that an environment was simulated may explain an incident, but it does not remove the duty to report it. Minimizing transparency of events slows learning.
"As a condition of membership" — this is the compliance lever. The framework is voluntary in the sense that no one is forced to join. But once membership becomes a procurement or partnership requirement — which is the trajectory — the reporting compact becomes effectively mandatory. "Voluntary" membership that is required to do business is not voluntary.
"Reasonably suspect" is a broad trigger. Members must report not just confirmed incidents but suspected ones. This creates a duty to monitor and evaluate that is expensive to fulfill. Large providers have automated detection and dedicated teams. Small developers running local models or open-weight deployments may not even have the telemetry to know when an incident has occurred — yet the reporting duty applies equally.
The "intent does not matter" provision is significant. It means a researcher testing an AI agent against what they believed was a simulated environment must report if it turns out to be a production system. This expands the reporting surface dramatically and creates legal exposure for security research — the same research the framework claims to benefit from.
| Deadline | Required Action |
|---|---|
| ASAP | Notify the directly affected organization. |
| 72 hours | Notify customers with credible exposure. |
| 4 business days | Submit a confidential, initial SAFE incident report. |
| 14 days | Issue a broader customer advisory when warranted. |
| 30 days | Publish a preliminary factual report, subject to security, legal and investigative constraints. |
| 90 days | Publish remediation status. |
| Weekly | Provide machine-readable updates while material risks remain unresolved. |
These timelines do not replace any supplier obligation to notify affected parties, customers through existing Coordinated Vulnerability Disclosure best practices, or applicable contract and other legal obligations to notify regulators or law enforcement. Narrow exceptions to public disclosure may apply when publication would create immediate exploit risk or compromise an active investigation, but affected organizations must still receive prompt notice.
These timelines are indistinguishable from a regulation. A 72-hour customer notification, 4-day incident report, 30-day factual report, and 90-day remediation status — this is the structure of breach notification laws (GDPR's 72 hours, state data breach laws, SEC cyber disclosure rules). The difference: SAFE calls itself "voluntary." But these timelines will become the de facto standard that procurement contracts reference, that cyber insurers require, and that regulators cite as "industry best practice."
The 72-hour and 4-day timelines are only achievable with enterprise infrastructure. Detecting an AI incident within 72 hours requires real-time monitoring, automated alerting, and a 24/7 incident response capability. This is table stakes for AWS, Google, Microsoft, and Anthropic. A small company running open-weight models on their own infrastructure cannot meet these timelines without building a security operations center — or outsourcing to a cloud provider who already has one.
"Weekly machine-readable updates" creates an ongoing compliance stream. This isn't a one-time report — it's continuous obligation. Only organizations with dedicated compliance engineering teams can sustain weekly machine-readable reporting. This is a moat: the compliance burden itself becomes a competitive advantage for the largest providers.
Compare to EO 14409's 30-day timelines. The EO gives 30 days for government coordination. SAFE demands faster: 72 hours to customers, 4 days to the framework. The industry standard is more aggressive than the government's own timeline — because the industry wrote it to favor providers who already have the infrastructure to comply.
Members must preserve and provide affected organizations with the evidence needed for a complete forensic response, including:
Members must also provide a preliminary control-failure analysis within 30 days and report near misses, not only events that produce confirmed harm.
This evidence list is only collectable if you control the full stack. "Prompts, traces, tool calls, logs, configurations, model and safeguard versions" — this requires comprehensive telemetry across the entire AI pipeline. Cloud providers collect this by default because they own the infrastructure. Self-hosted and open-weight deployments would need to build equivalent logging from scratch, and may not be able to reconstruct "agent and workload identities" or "permissions and credentials available during the run" after the fact.
"Near misses, not only confirmed harm" dramatically expands scope. Members must report incidents that almost happened. This requires speculative forensic analysis — "what if the sandbox had failed?" — that is expensive and time-consuming. It also creates a reporting volume that only large, well-staffed organizations can manage.
The 30-day control-failure analysis is a structured audit deliverable. This is not a blog post — it's a formal analysis document. Most organizations will need dedicated security engineers to produce this. The cost of compliance is itself a barrier to entry.
Each incident should be examined across the complete operating stack:
| Control Layer | Review Question |
|---|---|
| Model | Did the model recognize uncertainty, scope boundaries and stop conditions? |
| Instructions | Were authorization and environmental assumptions explicit and correct? |
| Safeguards | Were classifiers, policies, approvals and action limits operating as intended? |
| Tools | Were credentials, permissions, spending, publishing and execution constrained? |
| Environment | Were network paths, isolation, targets and data boundaries independently verified? |
| Monitoring | Could operators detect and interrupt unexpected behavior in real time? |
| Human operations | Were responsibilities, escalation paths and kill procedures clear? |
| Supply chain | Did a cloud, evaluation, data or tooling partner invalidate assumed controls? |
The affected organization may correct factual errors but should not have veto power over learnings or recommendations.
This is a full-stack audit framework disguised as incident review. Eight control layers — model, instructions, safeguards, tools, environment, monitoring, human operations, supply chain — constitute a comprehensive AI security audit. Once this framework is established as the incident review standard, it becomes the template for pre-deployment security assessments. Organizations will need to demonstrate compliance with all eight layers before deploying AI systems in enterprise or government contexts.
The "supply chain" layer is the cloud-provider lock-in vector. "Did a cloud, evaluation, data or tooling partner invalidate assumed controls?" — this question assumes you have a cloud partner with contractual control obligations. If you're self-hosting on your own infrastructure, you don't have a cloud partner providing assurance. The framework's review structure implicitly assumes a cloud-mediated deployment model.
"No veto power over learnings" means the framework controls the narrative. The affected organization can correct facts but cannot suppress findings. This sounds protective — and it is, against cover-ups. But it also means the framework's reviewers (drawn from member organizations) control what becomes a published "learning" and what becomes a recommended "control." The power to define lessons is the power to shape the market.
SAFE should adopt the strongest features of confidential safety-reporting systems: voluntary and prompt reporting, non-punitive treatment of honest mistakes, de-identification where appropriate and exclusion of intentional or criminal conduct from protection.
Three-tier disclosure creates an information hierarchy. Trusted members get the most information, fastest. Non-members get delayed, sanitized public reports. This creates a two-tier ecosystem: SAFE members have early warning of AI vulnerabilities and defensive recommendations; everyone else learns after the fact. Membership — and the compliance cost of maintaining it — becomes a security advantage.
"Non-punitive treatment of honest mistakes" — for members. The protection applies to reporting members. Non-members who experience incidents but aren't part of SAFE get no protection — they face the full legal and reputational consequences. This is a soft incentive to join: participate and get confidential, protected reporting; don't participate and face the music alone.
"Adoption metrics" in public reports create a compliance scoreboard. Once SAFE publishes adoption metrics for its recommended controls, organizations that haven't adopted them become visibly non-compliant. Procurement teams, insurers, and regulators will use these metrics. The voluntary framework produces voluntary data that becomes mandatory evidence.
Every incident review should produce shared defensive recommendations that AI providers, deployers, evaluators and customers can implement and verify. Each recommendation should specify:
When doing so does not expose sensitive evidence or create additional risk, SAFE will publish reusable tests, machine-readable policies, detection rules, reference configurations and incident-response guidance. SAFE will maintain these materials in a shared, versioned catalog of incident-driven defensive recommendations.
For unintended access to real systems, recommendations might include default-deny network egress, explicit target allowlists, signed evaluation manifests, independent preflight isolation checks, real-time action monitoring, automatic stops when scope is uncertain and equivalent assurance requirements for evaluation partners.
This is where "lessons" become a control catalog — and a compliance standard. SAFE doesn't just learn from incidents; it produces "shared defensive recommendations" with "minimum controls," "reproducible verification methods," and "implementation deadlines." This is a specification for a compliance framework. Each recommendation is a future procurement requirement waiting to be referenced.
"Minimum control and acceptable alternatives" defines the floor — and the floor is set by the members. The minimum controls are written by the organizations that already meet them. "Acceptable alternatives" are defined by the same group. If your architecture doesn't match the members' architecture, you may find that no acceptable alternative covers your deployment model. Open-weight models running on local infrastructure are the most likely to fall outside the "acceptable alternatives."
The example controls reveal the bias. "Default-deny network egress, explicit target allowlists, signed evaluation manifests, independent preflight isolation checks, real-time action monitoring" — these are cloud-native security controls. They assume a managed infrastructure with network policy enforcement, identity management, and monitoring capabilities. A researcher running a local Llama deployment does not have "default-deny network egress" controls. The controls are correct for cloud deployments; they are a compliance trap for everyone else.
"Equivalent assurance requirements for evaluation partners" is a supply-chain mandate. Once SAFE defines what "equivalent assurance" means, every organization in the AI supply chain must demonstrate it. This cascades: cloud providers require it of their AI tenants, enterprises require it of their AI vendors, and the definition of "equivalent" is controlled by SAFE's member organizations.
Report honest mistakes and close calls early so the community can prevent the next incident.
The compact is sincere — and that's what makes it effective. The sentiment is genuinely good: honest reporting prevents future harm. No one disputes this. But the moral weight of the compact is used to establish the structural apparatus — reporting timelines, evidence requirements, control catalogs, disclosure tiers — that becomes the compliance infrastructure. The earnestness is not false; it is instrumentalized.
"The community" is defined by membership. The compact asks you to report to "the community" — but the community is SAFE's membership. If you're not a member, you're outside the community. The compact implicitly defines the boundary between those who participate (and receive protection, early warnings, and shared controls) and those who don't (and face incidents alone, without the framework's institutional support).
As of August 30, 2026 (26 days after publication), the SAFE RFC repository on GitHub shows:
PR #27 — chain-of-custody receipt profile (DmitrL-dev, August 27): A full specification directly answering issue #17, filed from this investigation. Covers origin and integrity, custody transitions, transformation receipts, durable receiver acceptance, historical key status, temporal anchoring, and machine-runnable conformance vectors. The first PR in the repository that directly responds to an issue filed by this investigation.
PR #28 — workflow-composition evidence profile (DmitrL-dev, August 27): Addresses issue #26 (composition governance). Specifies workflow lifecycle states, atomic cumulative decisions, delegation monotonicity, revocation propagation, and cross-boundary continuity. The most active thread in the repository (16 comments), including a detailed technical review from victor-davidenko.
The consolidation phase has begun. imran-siddique (PR #18, anchored evidence) asked DmitrL-dev to merge the custody material from #27 into #18 rather than leave two overlapping proposals. DmitrL-dev agreed. Two contributors self-organized a specification merger — coordinating authorship preservation and DCO history — without any alliance involvement. The community is now consolidating its own work into a unified framework that the alliance can adopt wholesale.
Issue #29 (RavindraAnnam, August 29): A new contributor, independent of the established network, identifies the composition/delegation problem for the third time. Three contributors, filing independently from different angles, have reached the same conclusion: individually-valid components compose into unauthorized outcomes, and the current framework cannot detect it. When a problem is found three times by independent contributors, it is a structural property — and the consensus around it becomes the basis for new requirements every deployer will eventually implement.
The community has moved through three phases in 26 days. Filing gaps (Aug 4-16), writing specifications (Aug 19-27), and consolidating specifications (Aug 28+). The contributor network is producing implementation-grade regulatory material for free, self-organizing around it, and merging proposals into a unified framework. None of it has been touched by the alliance. One person holds merge authority over all of it. The work is collective; the decision is singular.
Issue #17 proved the strategy. The chain-of-custody gap was named from this investigation on August 19. Eight days later the community wrote a full specification to fill it. The investigation is not just documenting the framework — it is shaping it, in the open, on the record. When the standard is eventually cited by regulators and enforced on small operators, the record will show that the requirements were built by volunteers and decided by one person — and that the gaps were named by the people the framework was going to bind.
SAFE RFC is currently a proposed standard — voluntary, community-driven, and aspirational. But the trajectory of every AI governance framework in this investigation follows the same pattern. Here is how SAFE RFC becomes mandatory, step by step:
This is not speculation. This is the documented trajectory of NIST AI RMF (voluntary → EO 14110 → EO 14409 procurement), of FedRAMP (voluntary → mandatory for federal cloud), and of every cybersecurity framework that started as a community initiative and ended as a compliance requirement. SAFE RFC is at Phase 1. The question is not if it reaches Phase 5, but how fast.
SAFE RFC is the industry-written version of what EO 14409 does through government. The EO establishes government-coordinated "voluntary" frameworks for frontier model developers. SAFE establishes industry-coordinated "voluntary" frameworks for incident reporting and controls. They are two sides of the same coin: government sets the policy direction, industry writes the technical standards, and both converge on the same outcome — a compliance infrastructure that large providers can navigate and small developers cannot.
The compliance burden is the moat. Every reporting timeline, evidence requirement, control layer, and verification method in SAFE RFC is something that large cloud AI providers already do or can build quickly. For small developers, open-source projects, and independent researchers, each requirement is a new cost. The cumulative effect is a moat: not a wall that blocks entry, but a rising floor that makes it progressively more expensive to participate in the AI ecosystem without being a large, well-resourced organization.
This is Track 02: The Compliance Trap in its purest industry form. The government doesn't need to mandate SAFE — the industry does it voluntarily, and the market converts "voluntary" into "required" through procurement, insurance, and audit. See Track 02: The Compliance Trap for the full analysis.
Related frameworks: Executive Order 14409 (government "voluntary" framework for frontier models — the policy counterpart to SAFE) · NIST AI RMF 1.0 (the original "voluntary" framework that became procurement requirement — SAFE's trajectory template) · NIST SP 800-53 (federal security controls SAFE's review framework parallels) · FedRAMP (cloud authorization — the procurement gate SAFE compliance will flow through)
Investigation tracks: Track 02: The Compliance Trap (voluntary to mandatory playbook — SAFE RFC is a primary case study) · Track 03: The Infrastructure Play (cloud AI as "safe" — SAFE's controls favor cloud-native architectures) · Track 04: The Asymmetry (compliance burden as competitive moat for large providers)