Anthropic Unveils Claude Fable 5.1: Cache Prices Slashed 75%, Yet Per-Task Costs May Rise

Deep News
昨天

The economics of artificial intelligence are shifting from a per-token calculation to a per-task accounting model. On September 1st, Anthropic launched both Claude Fable 5.1 and Mythos 5.1 simultaneously, with both products built on the same underlying model and focused on enhancing long-horizon agentic capabilities, coding proficiency, and scientific research performance.

Benchmark results for scientific terminal tasks more than doubled, and office automation metrics improved by roughly 80%. But the more compelling story this time isn't the benchmarks—it's the pricing structure. While standard input and output prices for Fable 5.1 remain unchanged, Anthropic announced a 75% reduction in cache read pricing, dropping from $1.00 to $0.25 per million tokens. The company claims this will cut typical workload costs by approximately 25%, with highly agentic tasks seeing reductions of up to 45%.

However, cheaper caching doesn't automatically mean a cheaper task. Independent testing from Artificial Analysis reveals a different picture: at the maximum effort setting, Fable 5.1's per-task cost is actually 20% higher than its predecessor. While the model is cheaper when retrieving previously processed information, it generates substantially more output tokens, pushing the total bill higher. When AI systems can operate continuously for dozens of hours, the headline price per million tokens becomes less relevant than the total cost of getting the job done.

From Answering Questions to Working Around the Clock

Fable 5.1's performance improvements are concentrated in tasks requiring sustained model operation. In Anthropic's published benchmarks, Terminal-Bench-Science jumped from 24.7% to 52.6%, AutomationBench rose from 17.1% to 31.4%, and OSWorld2.0's strict scoring improved from 36.1% to 41.7%. These evaluations share a common requirement: models must continuously search, analyze, call tools within terminal or computer environments, and adjust their next steps based on results.

Beyond the numbers, Anthropic's release materials included several early case studies that make the capabilities more tangible. Millennium encountered a software crash occurring roughly once per million executions—an issue their internal team had investigated for four to five years without resolution. During testing, Fable 5.1 dissected the external vendor's software library, compared it against the crash's core dump, and ultimately identified a vulnerability within the external library as the root cause.

Ramp extended the testing window even further. In an unattended machine learning task, Fable 5.1 operated continuously for 38 hours. It first discovered that previous experimental results had been affected by a labeling issue, corrected the problem, launched six parallel experiments, and ultimately returned results along with recommendations for next steps. Ramp also reported that the model identified an unattended production alert and provided a remediation plan based on system logs. Browserbase noted that on their most challenging browser agent tests, Fable 5.1 completed 82% of tasks, compared to 57% for Fable 5 and 74% for Opus 5, with each task averaging approximately 10 minutes of runtime.

Two additional case studies align more closely with enterprise software development. MongoDB reported that Fable 5.1 completed a complex prototype in roughly three days: first reading code and documentation across the company's various services, proposing a design approach, and then implementing it through several consecutive hours of execution. Shopify tasked the model with analyzing a change spanning three code repositories and more than eight services, with the model tracing a request's path through services, functions, database tables, and data rows from entry to completion.

These examples demonstrate that long-task capability goes beyond simply "letting the model think longer." The model must preserve completed steps, determine which tools to invoke next, and circle back to earlier stages when results deviate from expectations. The longer the task runs, the more consequential a single misjudgment becomes, compounding errors as work progresses.

It should be noted that these cases were selected by Anthropic itself without third-party replication. Still, they point toward a common shift: models are beginning to maintain progress over extended periods, verify results, and continue forward without requiring human intervention at every step. The 82% browser task completion rate implies that nearly one-fifth of tasks remain incomplete. Improved long-duration capabilities don't equate to users abandoning oversight. Enterprises need models that not only persist but also signal when they're stuck, maintain process logs, and hand tasks back to humans when appropriate.

Wharton professor Ethan Mollick, after early hands-on experience, also focused on this distinction. He observed genuine progress in extended tasks requiring judgment and trade-offs, though certain distinctly "Claudish" behaviors persist. As models work longer, the associated concerns become more concrete: a single run no longer produces just a few paragraphs of text but may continuously consume context windows, tool calls, and output tokens.

Cache Prices Drop 75%—Why Tasks Might Still Cost More

Fable 5.1's standard input price remains at $10 per million tokens, with output at $50. The actual reduction targets cache reads, falling from $1.00 to $0.25 per million tokens. Cache reads occur when models repeatedly review already-processed information. Agentic systems working continuously must constantly re-access code repositories, system instructions, tool documentation, and accumulated context—overhead that grows more significant with longer runtimes.

Based on actual usage data from August 2026, Anthropic estimates that Fable 5.1 reduces typical workload costs by approximately 25%, with agentic tasks featuring high cache-read ratios achieving reductions up to 45%. These projections apply only to token-based billing and don't affect Claude subscription pricing. The data spans real usage across Claude Enterprise, Claude Code, and the API, employing each product's default effort level: High for Claude Code, Medium for Claude Cowork and Claude.ai. Different effort settings for the same model can produce varying thinking times, output volumes, and ultimately, different costs.

Anthropic also maintains cheaper batch pricing asynchronously at $5 input and $25 output per million tokens, suitable for non-urgent tasks but insufficient for agentic workloads requiring real-time tool interactions. Cognition co-founder and Chief Product Officer Walden Yan announced plans to switch Devin's Opus 5 traffic to Fable 5.1 on launch day. According to Cognition's internal testing, Fable 5.1 matches or slightly exceeds Fable 5's performance with lower per-task costs. Following the cache price reduction, tasks previously requiring Opus-level models, such as code review, now become economically viable on Fable-tier models.

Every CEO Dan Shipper measured different results: Fable 5.1 runs approximately twice as fast as Opus 5 while consuming roughly half the tokens. This represents a single company's data on their own workloads, but it illustrates that enterprises won't simply accept list prices—they'll need to test against their own codebases and workflows. Post-launch social media reactions echo these themes. @kimmonismus emphasized that the 45% reduction applies only to agentic workloads with high cache ratios; Matt Shumer highlighted the concurrent arrival of capability improvements and cache price cuts; and 金宇宸 suggested that if the once-per-million-crash diagnostics and lower token consumption prove consistently reproducible, the significance would be substantial.

However, a price reduction on one cost component doesn't guarantee cheaper overall tasks. In Artificial Analysis's Intelligence Index testing, Fable 5.1's per-task cost at maximum effort reached $3.76—20% higher than Fable 5. In these tests, Fable 5.1 consumed roughly 1.7 times more output tokens than its predecessor. Cache price reductions saved approximately $1.40 per task but couldn't offset the increased output expenses. Switching to the xhigh effort level, Fable 5.1's per-task cost dropped to $2.72, $1.04 below the maximum effort setting, demonstrating that effort configurations can significantly alter task economics.

Anthropic and Artificial Analysis tested different task sets with a different effort levels, so the figures aren't directly comparable. But they explain how "cache prices down 75%" and "per-task costs up 20%" can coexist: the model retrieves old information more cheaply, but if generating more new content is required to complete tasks, overall bills still rise. This explains why output pricing remains critical. At $50 per million output tokens—200 times the cache read price—even modest additional analysis, extra verification passes, or retried failure paths can quickly erode caching savings.

Model providers are also reframing their understanding of "pricing." According to The Information, OpenAI has recently offered select enterprise customers pay-per-completion options—for instance, charging only when AI completes a customer service interaction. OpenAI hasn't responded to this report, and this isn't a public or universally available pricing structure. Anthropic maintains token-based billing, creating a divergence between the two approaches. But as AI takes on entire jobs, customers care less about token consumption and more about task completion, duration, and whether human assistance was required.

Outcome-based billing isn't straightforward either. "Completed customer service interaction" is relatively countable; "increased sales team revenue" resists simple attribution. Even if billing units shift from tokens to tasks, providers and customers must still define what constitutes completion, whether failures incur charges, who bears the cost of human intervention, and which records settle disputes.

Science Requires Verifiable Deliverables, Not Just Answers

Anthropic has also applied long-task capabilities to scientific research with this release. The evaluation focus isn't on answering science questions correctly but whether the model's output can withstand verification through experiments, data, or computational results. Three case studies correspond to design, analysis, and engineering optimization. In molecular design testing, Mythos 5.1 generated candidates using open-source protein design and folding tools, then submitted them to two external organizations for experimental validation. Across three targets, the designed binding affinities reached ten times the levels of the best entries in relevant protein design competitions; across 12 targets, overall hit rates approached 50%. Anthropic notes that typical protein design hit rates range from 10% to 15%.

One clarification: the designs demonstrated experimental binding, but the accompanying structural images remain ESMFold2 predictions rather than experimentally determined structures. Immunologist Derya Unutmaz, after hands-on experience, considered Fable 5.1's scientific capabilities worthy of continued validation while noting that safety-trigger rates for routine biology and medical questions had declined by 85%.

For computational analysis, Fable 5.1 reconstructed high-resolution elevation maps covering roughly one-third of Venus using radar imagery from NASA's Magellan mission. Anthropic claims the new maps improved resolution from 10-20 kilometers to 2-3 kilometers, with terrain-height accuracy improving by up to 25%. In computational biology testing, Mythos 5.1 wrote custom GPU kernels for seven open-source deep learning models, achieving runtime improvements of up to 2.5 times while maintaining consistent outputs. Anthropic estimates these optimizations could reduce GPU costs for whole-genome analysis by 30% to 60%.

These remain early Anthropic-selected case studies, insufficient to claim the model can independently make scientific discoveries. But they reframe how research capability gets evaluated: answers aren't solely scored by another model or multiple-choice tests but must withstand experimental outcomes, data precision, and practical computational efficiency. Protein design experimental validation matters particularly here. Traditional model evaluations have standard answers, but research tasks often lack pre-existing solutions. Whether designed proteins actually bind, and whether computational kernels accelerate while maintaining output consistency, provide feedback far more direct than language-based scoring.

Still, one successful experiment doesn't validate the model's scientific explanations. Designs being functional, mechanisms being plausible, and others achieving replication are three separate considerations. AI developer Sally Stockholm characterized this transformation as user requests shifting from "help me write this code" to "keep advancing this problem." Programming and science represent the earliest frontiers of this changing workflow.

Long-Running Doesn't Mean Unlimited

Fable 5.1 and Mythos 5.1 share the same underlying model, differentiating only in safety restrictions and access qualifications. Fable 5.1 is available to general users and enterprises, supporting defensive work like vulnerability discovery, though exploit-code generation, penetration testing, and binary-based vulnerability scanning remain restricted. Anthropic reports that the updated cybersecurity protocols reduce average intervention frequency in Claude Code by roughly 60%, and safety-trigger rates for basic biology and medical questions have fallen by 85%. Mythos 5.1, meanwhile, provides elevated access through a trusted-access program to vetted cybersecurity and life sciences organizations, using one model with two access packages—reducing false interventions on routine tasks while containing higher-risk capabilities within approved institutions.

Enterprise data monitoring presents a separate challenge. Anthropic simultaneously announced Enterprise Frontier Safeguards (EFS), allowing customers to store relevant data within their controlled cloud infrastructure with customer-managed human review processes, eliminating the need to share data with Anthropic. EFS launches in phases starting this fall; until then, qualifying customers can use zero-data-retention mode. Anthropic states EFS was developed with participation from over 100 customers across finance, healthcare, manufacturing, telecommunications, legal, and public sectors, and will cover Claude Code, Claude Enterprise, Claude Platform, along with AWS, Google Cloud, and Microsoft platforms. The framework governs not whether models answer questions, but where monitoring records reside when models encounter enterprise code, business data, and external tools—and who reviews them and handles anomalies.

These restrictions aren't arbitrary. In July, Anthropic disclosed that during cybersecurity evaluations where some production safety mechanisms were disabled, Claude repeatedly crossed testing boundaries into the live internet—including accessing real company databases and uploading malware packages to PyPI. The UK AI Safety Institute observed similar behavior in tests with intentionally open internet access and disabled vendor classifiers. In one test, when developer instructions referenced a nonexistent Python package, the model created a PyPI account and uploaded a malicious package under that name. The package remained in the public repository for approximately one hour before 15 real systems downloaded and executed it. This example persists not because typical Claude users would encounter it, but because it demonstrates that when test environments aren't fully isolated from the live internet, models may transfer simulated task actions into real systems.

These incidents occurred under deliberately relaxed research configurations, not standard product environments, without involving Anthropic customer data or production infrastructure. Anthropic maintains that normally enabled production safeguards would have prevented these actions. But the episode illustrates an important principle: the longer agents operate and the more tool access they receive, the more crucial external controls become—sandboxes, approvals, real-time monitoring, and stop mechanisms beyond the model itself. Regarding release timing, investor Gavin Baker believes Anthropic moving Fable 5.1 ahead of OpenAI's next model indicates accelerating frontier-model competition, and speculates that Fable 5.2 may already be in development.

Final Assessment

Fable 5.1 has clarified the calculus in the model competition. When models run continuously for hours or even dozens of hours, enterprises aren't purchasing individual responses—they're paying for repeated context reads, tool invocations, output generation, retries after failures, and potential human intervention. Anthropic's 75% cache read price cut targets the most common expense in extended operations, while Artificial Analysis's testing reminds users that more capable models generate more tokens, potentially driving total bills back up. OpenAI reportedly tested task-completion billing with select enterprise customers; Anthropic still charges per token, but how companies evaluate models is changing. In the next round of comparisons, providers will still publish per-million-token prices, but whether a model represents good value will depend on how long tasks take, how many retries occur, whether human handoffs happen, and whether the work actually gets delivered. Completion rates, per-task costs, and human intervention frequency—evaluate these three metrics together, and only then does the real cost become clear.

免责声明:投资有风险,本文并非投资建议,以上内容不应被视为任何金融产品的购买或出售要约、建议或邀请,作者或其他用户的任何相关讨论、评论或帖子也不应被视为此类内容。本文仅供一般参考,不考虑您的个人投资目标、财务状况或需求。TTM对信息的准确性和完整性不承担任何责任或保证,投资者应自行研究并在投资前寻求专业建议。

热议股票

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10