AI Excellence Newsletter

The Cache Got Cheaper. Cost Per Task Did Not.

Cache reads on Fable 5.1 dropped 75 percent. That saving is real for users billed by the token. For a seat holder on a Team plan, it never arrives.

AI News

Overview

What a seat holder feels instead is a five-hour session window emptying faster.

The practical takeaway upfront: cost per task now means how much of your session window or weekly limit one completed piece of work consumes, and that number went up with Fable 5.1. The model is better on hard agentic work, but how much better cannot be verified from the outside. The lever that changes the burn rate is effort: set it to medium in Claude Code with the /effort command. The AI Excellence team has one job before the next issue: run a real ticket and watch the meter.

WHAT ANTHROPIC SAYS CHANGED

Anthropic reports incremental gains over Fable 5 across six areas: agentic coding, knowledge work, research, vision, long-context reasoning, and computer use. These are Anthropic’s own numbers.

The headline quality claims: fewer confident wrong answers, better judgment on ambiguous work, and a new behavior when stuck. When the model cannot make progress, it now says so instead of reporting success. Gains are widest at higher effort levels.


Cache reads dropped 75 percent for token-billed API users. The customer testimonials are the usual launch-day set: a one-in-a-million crash solved, a 38-hour unattended workflow, 82 percent task completion on an agentic benchmark. Anthropic chose them, and nobody outside Anthropic has checked them.

A DISCOUNT WE DO NOT SPEND

Read the announcement closely and the saving is scoped in a single clause. It applies “wherever usage is billed by token, such as on our API.” We are not billed by the token. We pay for seats, which is the same reason issue 12 called per-million-token prices a currency we do not spend. A cheaper cache read is good news for someone else’s invoice.

What we get instead is a heavier draw on the meter we do have. Claude Code runs Fable 5.1 at High effort by default, and adaptive thinking is always on. You can constrain it, but you cannot switch it off. More thinking means more output tokens per turn, and on a seat every one of those tokens comes out of the same five-hour window and the same weekly limit. Artificial Analysis, an independent evaluator, measured about 1.7 times the output tokens at max effort on September 1, 2026. Reddit filled in the seat-holder side on launch day, with several Max-plan users saying a five-hour window was gone in anywhere from 8 to 30 minutes. Those are individual reports rather than a benchmark, but they point the same way as the independent number.

The fix is in Anthropic’s own announcement: “when set to Low or Medium effort, Fable 5.1 achieves results similar to or better than Fable 5’s at a much lower cost.” In Claude Code version 2.1.257 or later, that is the /effort command set to medium.

None of these readings actually contradict each other. Anthropic’s charts plot cost across the whole effort range, Artificial Analysis tested at max, and the Reddit numbers came from the High default with subagents fanning out. Same model, different dial. How much of your weekly limit Fable 5.1 consumes depends on a setting most people never look at.

THREE SOURCES, THREE ANSWERS

Measuring completed work, not benchmark charts, was the right instinct. It is not sufficient now, because the charts themselves are no longer comparable.

The benchmark suite moved to Terminal-Bench 4.0, and Anthropic published no Terminal-Bench 2.1 number for Fable 5.1. SWE-bench

Verified, ARC-AGI-3 and CyberGym were not reported at all. Even the Terminal-Bench 4.0 headline comes with a footnote. The gap between Fable 5.1 and Mythos 5.1 reflects tasks where earlier, less precise cyber safeguards stepped in, and those tasks were scored as zero.

Artificial Analysis, which tests independently, puts Fable 5.1 at 66 on its Intelligence Index in September 2026, against 63 for Opus 5 and 62 for Fable 5. Three points. On usage, the sources disagree outright. Anthropic’s customers claim better token efficiency, Reddit users report faster limit burn, and Artificial Analysis measured about 1.7 times the output tokens at max effort. Nobody outside can say which is the general case, because each of them was measuring at a different effort level.

Fable 5.1 is better. How much better cannot be confirmed from the outside. The external signal is not just thin. It points in three directions at once.

Neither can you read usage per task from a price list. That number lives in your own session meter, not in a benchmark table.

THE NERF-BEFORE-LAUNCH BELIEF, HANDLED HONESTLY

A persistent belief among Claude users holds that Fable model quality gets quietly degraded before each major launch, so the new release feels like an improvement over a weakened baseline. It surfaced again in the Reddit launch thread. A Medium post went further and claimed server logs suggested Fable 5.1 was quietly A/B-served before September 1. Nobody has corroborated that, and Anthropic has said nothing about it.

Issue 08 documented that a model update can change behavior overnight and declined to confirm intentional nerfing. Nothing in the Fable 5.1 launch gives that belief new ground. The announcement does not acknowledge any pre-launch quality change, and Fable 5 is not being retired. Anthropic lists its retirement date as no sooner than June 9, 2027.

We will not print the pattern as fact.

The community belief is nonetheless pointing at something real. If capability is unmeasurable from the outside, then “it was better last week” and “it is better this week” are equally unfalsifiable until you test your own workflow. The belief is worth naming because it captures a structural problem with external evaluation, not because it is evidence of anything specific about Fable 5.1.

WHAT ACTUALLY CHANGED IN THE TOOL YOU HOLD

Fable 5.1 is available in Claude Code from version 2.1.257. From that version the fable alias resolves to Fable 5.1 on Anthropic-hosted sessions. Gateway sessions that have not been configured for 5.1 still get Fable 5.

Anthropic’s own documentation lists several behavior changes for anyone running Fable 5.1: whole-file rewrites instead of targeted edits, fewer parallel tool calls per turn, fewer progress updates, and denser prose.

From September 14 the temporary 50 percent limit boost expires and a permanent 25 percent increase replaces it. Against today, that is a 17 percent cut in weekly Claude Code capacity on every paid plan. Anthropic confirmed the arithmetic itself after deleting its original post, and BleepingComputer covered it.

The system card, as summarized by TechCrunch and others, says Fable 5.1 can still sometimes bypass approvals and auto-mode classifiers, and that alignment testing covered long-context and multi-agent settings only thinly. The agent is still the risk surface.

One correction to issue 12, which said every Claude product had been watermarked since August 2. Anthropic’s support article now says marking applies to models launched on or after August 2, 2026, and Fable 5.1 and Mythos 5.1 are the only ones listed so far. All Fable 5.1 output carries the watermark on every platform, the API included.

WHAT THIS MEANS FOR US

The standing thesis holds: model-agnostic harness, usage per completed task, the agent is the risk surface. Fable 5.1 adds a new axis. Usage per task can drift under a stable quality bar without any announcement.

  • It is my recommendation to start out with the /effort level set to medium and see if that thinking capacity is enough for your own use case. Give it a few days, then switch to high and compare which is more efficient per task.

  • Track the September 14 limit change. A net 17 percent reduction against current limits lands on every plan.

  • Expect whole-file rewrites and fewer parallel tool calls from Fable 5.1. If you see that behavior, it is documented, not a bug.

  • Consider adding a plain-English output instruction to your CLAUDE.md or system prompt. One tip gaining traction in the Claude and developer community is to request ASD-STE100 Simplified Technical English output, which several people report tightens the prose noticeably. The team is looking into this against our current approach, which applies the caveman and ponytail plugins.

  • Do not assume per-task window consumption is stable from one model release to the next. It moved this week without a price change.

  • For the AI Excellence team, with Fable 5.1 access via usage credits: run Fable 5.1 at a named effort level on one real ticket in a real repo, not a synthetic prompt.

  • Measure how much of your session window or weekly limit one completed ticket consumes. Compare across effort levels and against Fable 5 and Opus 5. The independent quality gap between them is small, 66 versus 63 versus 62.

  • Test the Fable-plans-cheaper-model-executes pattern that people in the Claude and developer community are already running: Fable 5.1 as orchestrator, Opus 5 or Sonnet 5 as the execution layer.

  • Confirm whether High-effort defaults are eating through the Fable share of your weekly limit faster than your workload can tolerate.

This is market intelligence, not a migration brief.

The model should be replaceable. The harness should be trusted. The effort setting is now part of what the harness controls.

Sources → Anthropic announcement  ·  Anthropic docs  ·  Anthropic overview  ·  The Decoder  ·  Simon Willison  ·  Every  ·  Claude support  ·  Claude support (watermarks)  ·  BleepingComputer  ·  Claude Code changelog  ·  r/ClaudeAI

 

GET STARTED

Tell us what you want to build.

Whether it’s a quick question or a detailed brief — we’d love to hear about it.

  • Honest assessment of whether we're the right fit
  • Fixed pricing, no hidden costs
  • 10+ years of trusted delivery
No sales pressure.
No lengthy process.
Just an honest conversation about technology.








    Thank you for contacting us!

    We'll be in touch with you shortly.