The Future of Life Institute released the Summer 2026 edition of its AI Safety Index this month, and the headline number is the one worth remembering: the highest overall grade awarded to any of the nine leading AI companies evaluated was a C+.
That grade went to Anthropic, at 2.66 on a four-point scale. OpenAI earned a C at 2.28, Google DeepMind a C at 2.01, and Meta a D+ at 1.32. Z.ai and Alibaba Cloud each received a D-. Three companies failed outright — xAI at 0.65, DeepSeek at 0.47, and Mistral at 0.33. An independent panel of seven researchers and governance experts, including Stuart Russell of UC Berkeley and David Krueger of the University of Montreal, scored the companies across thirty-seven indicators in six domains, using evidence collected through June 3.
The distribution matters less than what the panel found underneath it. Annual index reporting has become one of the few reliable instruments for tracking whether governance is keeping pace with capability, and Hassan Taher’s reading of the Stanford AI Index 2026 covered the capability half of that same ledger earlier this year. This report covers the other half, and the two do not describe a field moving in step.
The Finding That Should Concern Everyone
The single most consequential item in the report is not a grade. It is the observation that four of the industry’s safety leaders — Anthropic, OpenAI, Google DeepMind, and Meta — have weakened or voided their pledges to pause development unilaterally if their systems approach specified danger thresholds. Some replaced firm commitments with competitor-contingent conditions, effectively making their own restraint dependent on whether rivals exercise theirs.
The reviewers had a name for this. They called it a “moving goalpost” and concluded it has “undermined safety frameworks across the board.”
Russell put the implication plainly in remarks accompanying the report: “While there is good work being done on AI safety in the industry, the capabilities race has become more extreme. Companies have backed away from earlier commitments to release new systems only with safety measures appropriate for their capability levels; now, they’re planning to release them even if it’s demonstrably unsafe to do so.”
The panel’s top recommendation for Anthropic — the highest-scoring company in the index — was to reverse the walk-back in its most recent responsible scaling policy and restore the credibility of its commitments. When the recommendation for the class leader is “keep the promises you already made,” the class has a problem.
What the Domain Breakdown Reveals
The scores are not uniformly weak, and the pattern of strength and weakness is instructive. On Information Sharing, Anthropic earned a B+ and both OpenAI and Google DeepMind earned B-. On Governance and Accountability, Anthropic earned a B. Transparency practices — published model specifications, system prompts, incident reporting — have measurably improved.
Existential Safety is where the floor drops out. No company exceeded a C- in that domain, and the best grade awarded — a D+ — went to Anthropic and OpenAI. Panelists acknowledged specific constructive efforts, including Anthropic’s constitutional classifiers, OpenAI’s public calls for global governance institutions, Google DeepMind’s monitoring commitments, and Meta’s loss-of-control provisions, but judged the collective effort “entirely inadequate.” They also questioned whether the field’s dominant approaches — interpretability and chain-of-thought monitorability — actually address the problem, noting that “detection is not prevention.”
The panel further found that while companies are publishing and updating safety frameworks as U.S. and EU compliance deadlines approach, those frameworks “have weak teeth.” They frequently lack quantitative thresholds, genuinely independent audits, and clear internal authority over deployment decisions. For Google DeepMind, the reviewers noted it remains unclear which internal body can halt a deployment independently of executive leadership. For OpenAI, they recommended removing leadership’s ability to override its own Safety Advisory Group.
Why This Is a Procurement Document, Not Just an Ethics Report
Hassan Taher, an AI analyst and author who has written extensively on AI governance and ethics, has argued that safety index results should be read by buyers, not only by policy audiences. “Most companies treat this kind of report as a moral scorecard, and so they either nod at it or ignore it,” he has noted. “That’s the wrong frame. Read it as a vendor risk assessment. A framework without quantitative thresholds is a framework that can be reinterpreted under commercial pressure. A safety body that leadership can override is not an independent safety body. Those are contractual and operational facts about a supplier you are about to build your business on.”
He has been particularly pointed about the gap the panel identified between public messaging and revealed behavior — a finding the reviewers applied to Google DeepMind, OpenAI, and xAI, noting that reassuring leadership statements diverged from commercial conduct and legislative positions. “The report’s most useful sentence for a business audience is the one about rhetoric outpacing behavior,” Taher has observed. “It means stated commitments are an unreliable proxy for actual practice. So stop procuring on the basis of stated commitments. Ask for the audit reports. Ask who can stop a launch. Ask what threshold triggers what action, in writing. The companies that can answer those questions crisply are telling you something real, and the ones that answer with values language are telling you something too.”
The Global Picture Is Not a Regional One
Two findings complicate the familiar framing of a safety-conscious West and a reckless East. Failing grades went to one company from each of the three major AI regions — xAI in the United States, DeepSeek in China, and Mistral in Europe. And Mistral, the leading European lab operating in the world’s most aggressive AI regulatory environment, scored last overall. Regulation of the market has evidently not translated into safety leadership by the companies inside it.
The panel also flagged the industry’s pivot toward military applications as an emerging category of present-day harm, noting that between 2024 and 2026 several companies that had previously prohibited military use gradually reversed those prohibitions and began pursuing defense partnerships.
The Honest Reading
There are reasonable objections to the exercise. Grading safety practice involves judgment calls, the panel weights domains at its discretion, and comparing companies across radically different regulatory environments is genuinely difficult — the index itself devotes substantial space to explaining how Chinese regulatory context complicates direct comparison. The evidence cutoff of June 3 also means recent developments are excluded.
But the central finding does not depend on methodological precision. It is that during a period when capabilities advanced faster than at any point in the field’s history, the companies building those capabilities collectively loosened the conditions under which they had promised to stop. Krueger’s assessment was blunt: he called the lack of progress toward credible safety plans “scandalous,” and noted that recent leadership gestures toward coordinating a slowdown, while welcome, still understate how urgent the risk is and how unprepared the industry remains.
Self-regulation was the industry’s argument against binding rules. This report is the clearest available evidence on how that argument is holding up — and it lands while Congress weighs a federal framework that would suspend state AI legislation for three years, a proposal Hassan Taher has examined in the context of the Great American AI Act. Whatever replaces voluntary commitments will be shaped by evidence like this.
Sources:
- AI Safety Index — Summer 2026 — Future of Life Institute
- AI Safety Index Summer 2026, full report PDF — Future of Life Institute
- The Latest AI Safety Rankings Are In. Nobody Gets an A — TIME
- AI companies retreat from safety pledges even as capabilities grow — Axios
- Anthropic Tops 2026 AI Safety Index, But No AI Firm Earns Above a C+ — MIT Sloan Management Review Middle East



