Disclosure: Some links in this review are affiliate links. We may earn a commission at no extra cost to you. This never influences our ratings or editorial opinions.

Claude Fable 5.1 Review

Anthropic's Frontier Model for Long-Horizon Work
Coding FRONTIER Published September 2, 2026 · Updated September 2, 2026

Quick Verdict

Claude Fable 5.1 is the model Anthropic reserves for its hardest work — and after a month, we understand why. It is not the fastest or the cheapest at the frontier, but it is the most reliable when the job runs for hours. For long-horizon agentic coding, deep research, and knowledge work that stretches across a full day, it is the best we have tested. 9.2/10.

What is Claude Fable 5.1?

Claude Fable 5.1 is Anthropic's most capable generally available model, sitting in the Mythos-class tier above Claude Opus. Released September 1, 2026, it keeps the same $10/$50 token pricing as Fable 5 but cuts cache-read costs by 75%, making long context sessions significantly cheaper. Its sibling, Claude Mythos 5.1, is the same model with relaxed safeguards for vetted cybersecurity and life-science work.

It is built for the work that consumes whole days: multi-file refactors, deep research that follows up on its own findings, and document, spreadsheet, and slide production from a blank prompt. It is not the fastest model at the frontier, but it is the one we trusted with tasks we could not babysit.

Our Testing Process

We spent four weeks using Fable 5.1 daily on the work that actually consumed our year:

  • A full-codebase refactor across 40+ files, run as a single agent session
  • Multi-day autonomous coding with test-writing and self-checking
  • Long-windowed research briefs that followed up on their own findings
  • Document, spreadsheet, and slide production from a blank prompt
  • Dense PDF and chart reading with crop-and-zoom vision
  • Comparison runs against GPT-5.6 Sol and Claude Opus 5

Reasoning Depth

This is where Fable 5.1 earns its score. On a complex refactor, it kept the whole architecture in mind, caught a subtle coupling two reviewers missed, and explained its reasoning in a diff we could actually review. On ambiguous prompts it asked rather than guessed — the mark of a model that knows what it does not know.

"Fable 5.1 doesn't just write code — it reasons through why the change is right, and says so when it isn't sure. That honesty is worth more than a faster answer."

Long-Horizon Agentic Work

The standout feature. Fable 5.1 runs for hours without losing the thread, writes its own tests to check its work, and — crucially — reports honestly when stuck rather than declaring false success. In a multi-day session, that honesty saved us an afternoon of debugging a confidently wrong solution.

Knowledge Work & Vision

It took a blank prompt to a finished document with live formulas and a slide deck without hand-holding. Its crop-and-zoom vision read dense financial filings and charts more accurately than any model we have tested, spotting a footnote-level discrepancy that had slipped past a human analyst.

Pricing

Fable 5.1 costs $10 per million input tokens and $50 per million output tokens — the same as Fable 5. The headline change is cache reads at $0.25 per million tokens, 75% cheaper, which cuts typical workload cost by ~25% and highly agentic work by up to ~45%. For long-context sessions that re-read a cached prefix, the economics are far better than the sticker price implies.

Who Should Use It

Fable 5.1 is for teams doing ambitious, long-running, asynchronous work: engineers on multi-day codebase projects, researchers on deep multi-step investigations, and knowledge workers producing complex documents. If your work is short and varied, GPT-5.6 will serve you better for less.

Final Verdict

Claude Fable 5.1 is the depth pick — the most reliable model we have tested when the job runs for hours. It costs a premium, but for long-horizon agentic work, the cache-read savings and honest self-reporting make it worth it. 9.2/10.

★★★★½
9.2
out of 10 · Excellent
Compare vs GPT-5.6 →

StackHK head-to-head test

Quick Info

Price$10 / $50 per MTok
Cache Read$0.25 / MTok (−75%)
Context1M tokens
Max Output128K tokens
Best ForLong-horizon agentic work
ReleaseSep 1, 2026

At a Glance

Best for
Long-horizon agentic work
Starting Price
$10 / $50 per MTok
Context Window
1M tokens
Overall Score
9.2/10
Max Output
128K tokens
Cache Read
$0.25 / MTok

Score Breakdown: Our 5 Dimensions

DimensionScoreNotes
Quality of Output9.5Best deep reasoning and root-cause fixes we have measured
Ease of Use8.8Powerful but slower; effort controls need tuning
Value for Money8.9Premium price, but cache reads close the gap on long sessions
Speed & Reliability9.0Slower per answer, but fewer confident wrong answers
Support & Docs8.8Strong docs; smaller ecosystem than OpenAI
Overall9.2Across 5 dimensions of hands-on testing

Benchmark Reference Data

Independent, third-party scores for Claude Fable 5.1 from public evaluation sources. Source attribution in brackets reflects where each figure was published — official vendor results are marked accordingly.

SWE-bench Verified[Official System Card]
Fable 5.1
95.0%
SWE-bench Pro[Official System Card]
Fable 5.1
80.0%
GPQA Diamond[Official System Card]
Fable 5.1
92.6%
LiveCodeBench[Official System Card]
Fable 5.1
90.52%
Terminal-Bench 4.0[Official System Card]
Fable 5.1
55.8%
Humanity's Last Exam (tools)[Official System Card]
Fable 5.1
65.0%

Source: Anthropic Claude Fable 5.1 System Card (production safeguards enabled). Benchmarks are self-reported or third-party as noted; scores reflect the publishing vendor's configuration and may not be directly comparable across different evaluation setups.

Frequently Asked Questions

What is Claude Fable 5.1?

Claude Fable 5.1 is Anthropic's most capable generally available model, a Mythos-class model that sits above Claude Opus. It is built for long-horizon agentic coding, deep research, and knowledge work that runs for hours.

How much does Claude Fable 5.1 cost?

Fable 5.1 costs $10 per million input tokens and $50 per million output tokens, the same as Fable 5. Cache reads are $0.25 per million tokens, 75% cheaper, cutting typical workload cost by ~25% and highly agentic work by up to ~45%.

Is Claude Fable 5.1 better than GPT-5.6?

For long-horizon agentic coding and deep reasoning, Fable 5.1 leads on depth and stamina. For breadth, speed, and dramatically lower cost per task, GPT-5.6 Sol is the stronger choice. The pick depends on your work.

What is Claude Mythos 5.1?

Mythos 5.1 is the same underlying model as Fable 5.1 with more permissive safeguards for cybersecurity and life sciences, available only through trusted access programs.

Related Reviews

Claude ProClaude Pro9.4

Hands-on testing, score breakdown, pros & cons.

StackHKFull review
ChatGPT-4oChatGPT-4o8.7

Hands-on testing, score breakdown, pros & cons.

StackHKFull review
Fable 5.1vs GPT-5.69.2

Head-to-head: depth vs breadth, pricing, and stamina.

StackHKComparison