anthropic

Claude Fable 5.1

Claude Fable 5.1 is Anthropic's flagship model for agentic coding and science, featuring a 1M context window, adaptive thinking, and 128K output tokens.

AnthropicAgentic CodingLong ContextReasoningMultimodal
anthropic logoanthropicClaudeSeptember 1, 2026
Context
1Mtokens
Max Output
128Ktokens
Input Price
$10.00/ 1M
Output Price
$50.00/ 1M
Modality:TextImage
Capabilities:VisionToolsStreamingReasoning
Benchmarks
GPQA
93.7%
GPQA: Graduate-Level Science Q&A. A rigorous benchmark with 448 multiple-choice questions in biology, physics, and chemistry created by domain experts. PhD experts only achieve 65-74% accuracy, while non-experts score just 34% even with unlimited web access (hence 'Google-proof'). Claude Fable 5.1 scored 93.7% on this benchmark.
HLE
60.9%
HLE: High-Level Expertise Reasoning. Tests a model's ability to demonstrate expert-level reasoning across specialized domains. Evaluates deep understanding of complex topics that require professional-level knowledge. Claude Fable 5.1 scored 60.9% on this benchmark.
MMLU
89.8%
MMLU: Massive Multitask Language Understanding. A comprehensive benchmark with 16,000 multiple-choice questions across 57 academic subjects including math, philosophy, law, and medicine. Tests broad knowledge and reasoning capabilities. Claude Fable 5.1 scored 89.8% on this benchmark.
MMLU Pro
78.4%
MMLU Pro: MMLU Professional Edition. An enhanced version of MMLU with 12,032 questions using a harder 10-option multiple choice format. Covers Math, Physics, Chemistry, Law, Engineering, Economics, Health, Psychology, Business, Biology, Philosophy, and Computer Science. Claude Fable 5.1 scored 78.4% on this benchmark.
SimpleQA
44.5%
SimpleQA: Factual Accuracy Benchmark. Tests a model's ability to provide accurate, factual responses to straightforward questions. Measures reliability and reduces hallucinations in knowledge retrieval tasks. Claude Fable 5.1 scored 44.5% on this benchmark.
IFEval
88.6%
IFEval: Instruction Following Evaluation. Measures how well a model follows specific instructions and constraints. Tests the ability to adhere to formatting rules, length limits, and other explicit requirements. Claude Fable 5.1 scored 88.6% on this benchmark.
AIME 2025
88%
AIME 2025: American Invitational Math Exam. Competition-level mathematics problems from the prestigious AIME exam designed for talented high school students. Tests advanced mathematical problem-solving requiring abstract reasoning, not just pattern matching. Claude Fable 5.1 scored 88% on this benchmark.
MATH
94.2%
MATH: Mathematical Problem Solving. A comprehensive math benchmark testing problem-solving across algebra, geometry, calculus, and other mathematical domains. Requires multi-step reasoning and formal mathematical knowledge. Claude Fable 5.1 scored 94.2% on this benchmark.
GSM8k
97.6%
GSM8k: Grade School Math 8K. 8,500 grade school-level math word problems requiring multi-step reasoning. Tests basic arithmetic and logical thinking through real-world scenarios like shopping or time calculations. Claude Fable 5.1 scored 97.6% on this benchmark.
MGSM
92.1%
MGSM: Multilingual Grade School Math. The GSM8k benchmark translated into 10 languages including Spanish, French, German, Russian, Chinese, and Japanese. Tests mathematical reasoning across different languages. Claude Fable 5.1 scored 92.1% on this benchmark.
MathVista
71.5%
MathVista: Mathematical Visual Reasoning. Tests the ability to solve math problems that involve visual elements like charts, graphs, geometry diagrams, and scientific figures. Combines visual understanding with mathematical reasoning. Claude Fable 5.1 scored 71.5% on this benchmark.
SWE-Bench
58.2%
SWE-Bench: Software Engineering Benchmark. AI models attempt to resolve real GitHub issues in open-source Python projects with human verification. Tests practical software engineering skills on production codebases. Top models went from 4.4% in 2023 to over 70% in 2024. Claude Fable 5.1 scored 58.2% on this benchmark.
HumanEval
92.4%
HumanEval: Python Programming Problems. 164 hand-written programming problems where models must generate correct Python function implementations. Each solution is verified against unit tests. Top models now achieve 90%+ accuracy. Claude Fable 5.1 scored 92.4% on this benchmark.
LiveCodeBench
74.8%
LiveCodeBench: Live Coding Benchmark. Tests coding abilities on continuously updated, real-world programming challenges. Unlike static benchmarks, uses fresh problems to prevent data contamination and measure true coding skills. Claude Fable 5.1 scored 74.8% on this benchmark.
MMMU
71.2%
MMMU: Multimodal Understanding. Massive Multi-discipline Multimodal Understanding benchmark testing vision-language models on college-level problems across 30 subjects requiring both image understanding and expert knowledge. Claude Fable 5.1 scored 71.2% on this benchmark.
MMMU Pro
58.7%
MMMU Pro: MMMU Professional Edition. Enhanced version of MMMU with more challenging questions and stricter evaluation. Tests advanced multimodal reasoning at professional and expert levels. Claude Fable 5.1 scored 58.7% on this benchmark.
ChartQA
89.2%
ChartQA: Chart Question Answering. Tests the ability to understand and reason about information presented in charts and graphs. Requires extracting data, comparing values, and performing calculations from visual data representations. Claude Fable 5.1 scored 89.2% on this benchmark.
DocVQA
94.1%
DocVQA: Document Visual Q&A. Document Visual Question Answering benchmark testing the ability to extract and reason about information from document images including forms, reports, and scanned text. Claude Fable 5.1 scored 94.1% on this benchmark.
Terminal-Bench
55.8%
Terminal-Bench: Terminal/CLI Tasks. Tests the ability to perform command-line operations, write shell scripts, and navigate terminal environments. Measures practical system administration and development workflow skills. Claude Fable 5.1 scored 55.8% on this benchmark.
ARC-AGI
12.4%
ARC-AGI: Abstraction & Reasoning. Abstraction and Reasoning Corpus for AGI - tests fluid intelligence through novel pattern recognition puzzles. Each task requires discovering the underlying rule from examples, measuring general reasoning ability rather than memorization. Claude Fable 5.1 scored 12.4% on this benchmark.

About Claude Fable 5.1

Learn about Claude Fable 5.1's capabilities, features, and how it can help you achieve better results.

Architectural Overview

Claude Fable 5.1 is Anthropic's flagship frontier model built for long-horizon agentic workflows, autonomous software engineering, and scientific research. It succeeds the original Fable 5 model while maintaining identical base token pricing and introducing a 75% reduction in prompt cache read rates. The model is built on an adaptive thinking architecture that allows it to plan, execute, and self-correct across multi-step execution loops without dropping context across its million-token window.

Core Capabilities

Architecturally, Fable 5.1 is optimized for agentic persistence, enabling it to handle tasks that span hours or days of execution, such as full-stack application development or multi-stage scientific data modeling. It incorporates refined safeguards that reduce false-positive refusals on technical security and biology tasks by up to 60-85%.

Target Workloads

This model is designed for developers and enterprises building autonomous agents where reliability in multi-file refactoring and expert-level reasoning is the primary requirement. Native vision support also allows it to parse technical diagrams, UI mockups, and nested tabular charts directly within codebases.

Claude Fable 5.1

Use Cases

Discover the different ways you can use Claude Fable 5.1 to achieve great results.

Autonomous Software Engineering

Navigates complex multi-directory repositories to locate architectural flaws and execute multi-file refactors coherently.

Scientific Data Modeling

Automates the parsing of raw observational data, differential equation fitting, and neural network training for research.

Defensive Security Auditing

Inspects internal APIs and source code to identify software vulnerabilities and suggest remediation patches.

Visual UI/UX Prototyping

Generates functional 3D simulations, interactive WebGL scenes, and production-ready frontend code from text descriptions.

Multi-Source Knowledge Synthesis

Ingests massive datasets of SEC filings, transcripts, and charts to build detailed financial or analytical valuation models.

Long-Horizon Agent Orchestration

Acts as a primary controller for asynchronous workflows, delegating subtasks to tools and monitoring progress over dozens of turns.

Strengths

Limitations

Elite Scientific Reasoning: Achieves 93.7% on GPQA Diamond and 55.8% on Terminal-Bench for research and technical automation.
Premium Non-Cached Pricing: Base pricing of $10/M input and $50/M output tokens makes un-cached single calls expensive.
Massive Agentic Capacity: Provides a 1,000,000 token context window and 128,000 max output tokens for long autonomous execution.
Token-Heavy Internal Drafting: Extended adaptive thinking cycles consume high output token counts before final emission.
Optimized Cache Economy: Cache read pricing at $0.25/M tokens reduces total operational costs on multi-step prompt loops.
No Native Video Support: The model processes static images and documents but cannot ingest raw video files.
Refined Safety System: Reduces benign false-positive blocks by up to 60% on defensive cybersecurity and vulnerability tasks.
Full-File Replacement Bias: Tends to rewrite entire source files for simple edits unless explicitly instructed to output diffs.

API Quick Start

anthropic/claude-fable-5.1

View Documentation
anthropic SDK
import Anthropic from '@anthropic-ai/sdk';

const anthropic = new Anthropic({
  apiKey: process.env.ANTHROPIC_API_KEY,
});

const msg = await anthropic.messages.create({
  model: "claude-fable-5.1",
  max_tokens: 1024,
  messages: [{ role: "user", content: "Analyze this codebase for architectural flaws." }],
});

console.log(msg.content[0].text);

Install the SDK and start making API calls in minutes.

Community Feedback

See what the community thinks about Claude Fable 5.1

Fable 5.1 is the first model that actually feels like a co-pilot for long refactors without losing the plot halfway through.
DevOpsGuru
reddit
While prior models became hard to follow the longer they worked, Fable 5.1 remains readable over long, multi-step tasks.
Craig Falls
hackernews
Claude Fable 5.1 is now available in Cursor! It's the most capable model we've run on CursorBench 3.2, scoring 73.4%.
cursor_ai
twitter
Anthropic's Mythos-class models are finally cracking the long agent problem. Fable 5.1 persistence is unmatched.
HackernewsUser42
hackernews
Fable 5.1 is insane guys. It one shotted my session usage limit but finished the entire migration cleanly.
TechLeadDev
twitter
The prompt cache price cut makes running continuous test-fix loops actually viable in production.
VibeCoder
youtube

Related Videos

Watch tutorials, reviews, and discussions about Claude Fable 5.1

Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads

For highly agentic work, the savings will often be much larger, up to approximately 50%

Claude Mythos 5.1 is identical to Fable 5.1, but it offers more permissive safeguards for vetted individuals and organizations

The shadows look way more realistic, the physics feel more real

being asked to. And what's also interesting about this is we've been praising its desire or its ability to get all the little details right without you asking. And this is one place where for whatever reason it's not as good.

just mute yourself if you're typing. It happens to be uh about uh 50% more token efficient than Opus 5. It's also faster than Opus 5. So, if

polished. Um, I think in terms of like the future here, there's more that we're going to be uh exploring in terms of how we can make this even easier. I think

More than just prompts

Supercharge your workflow with AI Automation

Automatio combines the power of AI agents, web automation, and smart integrations to help you accomplish more in less time.

AI Agents
Web Automation
Smart Workflows

What to get right first

The decisions that are painful to change later in Claude Fable 5.1.

Maximize Cache Savings

Keep system instructions and large repository trees at the top of your prompt to take advantage of the $0.25/M cache read pricing.

Control Execution Effort

Use the effort parameter to switch between 'low' for standard tasks and 'high' for complex debugging to balance cost and speed.

Prompt for Targeted Edits

Explicitly instruct the model to provide unified diffs or targeted line changes to prevent it from rewriting entire large files.

Nudge Autonomous Persistence

In agentic loops, include instructions prompting the model to complete multi-step tasks directly rather than pausing for permission.

Testimonials

What Our Users Say

Join thousands of satisfied users who have transformed their workflow

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Jonathan Kogan

Jonathan Kogan

Co-Founder/CEO, rpatools.io

Automatio is one of the most used for RPA Tools both internally and externally. It saves us countless hours of work and we realized this could do the same for other startups and so we choose Automatio for most of our automation needs.

Mohammed Ibrahim

Mohammed Ibrahim

CEO, qannas.pro

I have used many tools over the past 5 years, Automatio is the Jack of All trades.. !! it could be your scraping bot in the morning and then it becomes your VA by the noon and in the evening it does your automations.. its amazing!

Ben Bressington

Ben Bressington

CTO, AiChatSolutions

Automatio is fantastic and simple to use to extract data from any website. This allowed me to replace a developer and do tasks myself as they only take a few minutes to setup and forget about it. Automatio is a game changer!

Sarah Chen

Sarah Chen

Head of Growth, ScaleUp Labs

We've tried dozens of automation tools, but Automatio stands out for its flexibility and ease of use. Our team productivity increased by 40% within the first month of adoption.

David Park

David Park

Founder, DataDriven.io

The AI-powered features in Automatio are incredible. It understands context and adapts to changes in websites automatically. No more broken scrapers!

Emily Rodriguez

Emily Rodriguez

Marketing Director, GrowthMetrics

Automatio transformed our lead generation process. What used to take our team days now happens automatically in minutes. The ROI is incredible.

Related AI Models

openai

GPT-5.5

OpenAI

GPT-5.5 is OpenAI's flagship frontier model with a 1M context window and five reasoning effort levels, optimized for autonomous agentic workflows and coding.

1M context
$5.00/$30.00/1M
xai

Grok-3

xAI

Grok-3 is xAI's flagship reasoning model, featuring deep logic deduction, a 128k context window, and real-time integration with X for live research and coding.

1M context
$3.00/$15.00/1M
moonshot

Kimi K3

Moonshot

Kimi K3 is Moonshot AI's 2.8T MoE model with a 1M token context window, native multimodal vision, and frontier-tier coding performance for complex agents.

1M context
$3.00/$15.00/1M
google

Gemini 3.1 Flash Live Preview

Google

Gemini 3.1 Flash Live Preview is Google's ultra-low-latency, audio-to-audio model featuring a 131K context window, high-fidelity multimodal reasoning, and...

131K context
$0.75/$4.50/1M
anthropic

Claude Opus 4.7

Anthropic

Claude Opus 4.7 is Anthropic's flagship model with a 1-million-token context, adaptive reasoning, and 3.3x vision resolution for enterprise-scale agents.

1M context
$5.00/$25.00/1M
openai

GPT-5.2 Pro

OpenAI

GPT-5.2 Pro is OpenAI's 2025 flagship reasoning model featuring Extended Thinking for SOTA performance in mathematics, coding, and expert knowledge work.

400K context
$21.00/$168.00/1M
google

Gemini 3.1 Pro

Google

Gemini 3.1 Pro is Google's elite multimodal model featuring the DeepThink reasoning engine, a 1M+ context window, and industry-leading ARC-AGI logic scores.

1M context
$2.00/$12.00/1M
alibaba

Qwen 3.7 Max

alibaba

Qwen 3.7 Max is Alibaba’s flagship AI model for deep reasoning and autonomous agent tasks, featuring a 256k context window and top-tier coding performance.

256K context
$1.20/$6.00/1M

Frequently Asked Questions

Find answers to common questions about Claude Fable 5.1