Skip to content
EssayAI & Technology09 / 08 / ’266 min readAI

Fable 5.1 vs GPT-6 Astra: Compare the Terms, Not the Tables

Anthropic shipped Claude Fable 5.1 on September 1 and OpenAI shipped GPT-6 Astra two days later, at an identical $10/$50 list price and with each lab's own benchmark table putting itself on top. The tables cannot settle it — both labs footnote their scores into a wrapper contest. The comparison that matters for a buyer is the terms: who can have the model, what a completed task costs, where the data sits, and what is switched off by default.

Anthropic shipped Claude Fable 5.1 on September 1. OpenAI shipped GPT-6 Astra two days later. Both list at $10 per million tokens in and $50 out, both claim the top of the table, and the comparison threads wrote themselves. Most of them are comparing the wrong thing.

This Journal has argued three times this year that the model is no longer the product: the wrapper is, access is a supply-chain risk, and the race has split into a consumer half and an enterprise half. This week’s two launches put that argument in the fine print. The spec sheets have converged. The terms have not.

Claude Fable 5.1 and GPT-6 Astra set against each other across a diagonal split, with the note that both list at $10 in and $50 out
Two launches, one list price. The comparison that matters is in the terms.

Same week, same price, two different products

Anthropic’s release is two SKUs of one set of weights. Claude Fable 5.1 and Claude Mythos 5.1 are, in the company’s words, “the same model, but with different levels of safeguards.” Fable 5.1 is generally available: Pro, Max, Team and Enterprise plans, the API, AWS, Google Cloud and Microsoft Foundry. Mythos 5.1 is not. It ships through Anthropic’s trusted-access programs for cyber and life-sciences work, and “currently, it is only available to a set of US organizations.” The list price did not move from Fable 5. What moved is the cache: reads now cost $0.25 per million, 75% less than Fable 5, which Anthropic says takes roughly 25% off a typical workload and up to 45% off a heavily agentic one. Dan Shipper of Every summarized the pitch in one line: “It’s friendly Fable. Fable-level intelligence, Opus-level price, Sonnet-speed.”

OpenAI’s release is one model behind a threshold. GPT-6 Astra arrived September 3 as “the world’s most intelligent and aligned model,” at the same $10 and $50 on Standard processing, with a Fast mode that doubles the speed at double the price. It rolls out to Plus, Pro, Business and Enterprise, the API, Azure and Bedrock. For enterprise workspaces, though, “access is off by default at launch.” Astra is also the first model OpenAI has designated Critical for cybersecurity under its Preparedness Framework, so its most advanced cyber capabilities go to a group of testers first, with wider access through Daybreak Blue to follow.

Two labs, one week, an identical list price to the dollar — and the divergence is in what the price buys and who is allowed to pay it.

TermClaude Fable 5.1GPT-6 Astra
ReleasedSep 1, 2026Sep 3, 2026
List price, per 1M tokens$10 in / $50 out$10 in / $50 out
Cache reads$0.25 per 1M, 75% below Fable 5Separate rates; not on the launch page
Context / max output1M / 128KNot on OpenAI-owned pages
SpeedAdaptive thinking, always onStandard, or Fast at 2× the speed and 2× the price
Where it runsClaude plans, API, AWS, Google Cloud, Microsoft FoundryChatGPT plans, API, Azure, AWS Bedrock
Enterprise defaultGenerally availableOff by default at launch
Gated tierMythos 5.1: same weights, fewer safeguards, US organizations onlyAdvanced cyber capabilities: testers first, then Daybreak Blue
Cyber ratingLower risk category, Frontier Compliance Framework (Mythos 5.1)Critical, Preparedness Framework
Bio / chem ratingBelow the next RSP tier (Mythos 5.1)High capability
Humanity’s Last Exam, with tools65.0%57.2%
Computer-use misaligned outcomes9.5%2.4%
From each lab’s launch page and platform docs. The last two rows are OpenAI’s own numbers; treat them as vendor-reported.

Whose table do you believe

Each lab published a table, and each table has its publisher on top. Anthropic’s has Fable 5.1 at 52.6% on Terminal-Bench-Science against 22.4% for GPT-5.6 Sol. OpenAI’s has Astra at 64.6% on the same benchmark against Fable 5.1’s 52.6%, at roughly 31% lower estimated API cost. Both sets of numbers can be true, because each lab ran its own harness, and they still tell you nothing about which model does your job.

The footnotes are more useful than the headline rows. OpenAI’s own launch page lists Astra at 57.2% on Humanity’s Last Exam with tools, next to Fable 5.1’s 65.0%. It concedes a benchmark loss on the page announcing the win. Anthropic notes that on OSWorld tasks where its safeguards intervened, Fable 5.1 and Fable 5 scored zero, and that its OSWorld numbers are not comparable to earlier published ones. OpenAI notes it left Fable out of three life-science evaluations because the model refuses most of the questions. Both tables measure the model plus the wrapper, and each lab grades its own wrapper.

VentureBeat’s Carl Franzen put it plainly the day Fable 5.1 shipped: “Those numbers should be read as vendor-reported results rather than independent proof of superiority.” That is what a launch-day table is for.

The gate is the spec sheet

Anthropic’s gate is a product line. The Fable tier keeps the safeguards; the Mythos tier loosens them for vetted, US-based organizations. Between releases the company retuned those safeguards, reporting 60% fewer false positives on cyber requests and biology guardrails that fire 85% less often on benign ones, and it now lets Fable 5.1 hunt for software vulnerabilities. Enterprise Frontier Safeguards, due in phases this fall, move data custody onto infrastructure the customer controls.

OpenAI’s gate is a threshold. The company says it slowed Astra’s development to harden protections against cyber misuse, paused certain frontier training for two weeks after the Hugging Face incident, and restarted the run on August 28. It expects the launch safeguards to create “more friction than we ultimately intend,” sometimes pausing legitimate defensive work. And its safety overview carries a sentence no competitor wrote for it: “GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol.” Under adversarial testing the model can sometimes evade OpenAI’s internal monitors on sabotage tasks. The same launch reports a 2.4% misaligned-outcome rate in a computer-use stress test, against 9.5% for Fable 5.1.

OpenAI shipped the more capable cyber model and told you it is harder to watch. Anthropic shipped the same weights twice and told you which one you are allowed to have.

What this means if you’re buying

Per-token price is a tie, so stop comparing it. Compare cost per completed task and where your data sits. OpenAI claims Astra beats Fable 5.1 on Terminal-Bench 4.0 at 63% lower estimated cost per task. Anthropic’s answer is the cache price: Cognition is moving Devin’s Opus traffic to Fable 5.1 on launch day because a Fable-class model is, in its words, “finally economical for the workloads we’d kept on Opus.” Which claim wins depends on the shape of your work. Agents that reread the same context all day win on the cache price. Short, heavy reasoning wins on cost per task. The only table that settles it is the one you build on your own tasks.

For the work an agency ships, the testimonials say more than the benchmarks. Canva’s read on Fable 5.1 is that the writing is “more understandable, more meaningful,” and follows its guidance better. Hebbia ranks its decks above every other model it has tested. OpenAI says Astra is its best model at adhering to existing templates and laying out slides. Browserbase reports Fable 5.1 completing 82% of browser tasks in about ten minutes each; OpenAI reports Astra at 72.6% on its computer-use suite in roughly 40 minutes per task, down from 75 for Sol. Those figures are not comparable to each other. They point at the same thing: both labs now sell hours-long unattended runs, and price for them.

VentureBeat, citing the Financial Times and Ramp’s transaction data, reports that Fable 5 drew only about 11% of Anthropic model spending across roughly 70,000 companies, and that Fable 5.1 remains expensive against a market where GPT-5.6 Sol is promoting at $4 in and $20 out. Buyers had the frontier model and mostly did not put it on their workloads. The cache cut is Anthropic’s answer to that. Astra off by default in enterprise workspaces, Mythos limited to US organizations, and EFS still a season away are the terms that decide whether either answer reaches your team this quarter.