Go to Easy Invoice Cloud
PepperTools Guide
World News

Claude Fable 5.1 is here: what the benchmarks really say – and why your quota still shrinks on September 14

Anthropic releases Claude Fable 5.1: far ahead of Opus 5 and GPT-5.6 Sol in science and automation tests, and about 25% cheaper to run. At the same time, the Claude Code weekly quota effectively shrinks from September 14 – even though Anthropic calls it an increase.

Claude Fable 5.1 is here: what the benchmarks really say – and why your quota still shrinks on September 14

Update 4 September 2026: OpenAI has released GPT-6 Astra. What reaches your ChatGPT plan and what it costs is covered in GPT-6 Astra: availability and pricing.

Anthropic released Claude Fable 5.1 on September 1, 2026. In the published tests, the model solves considerably more tasks than its predecessor Fable 5, than the cheaper Opus 5, and than OpenAI's GPT-5.6 Sol – especially in scientific research and business-process automation. Per-token prices stay the same, but Fable 5.1 uses fewer of them: typical workloads become about 25 percent cheaper, according to Anthropic. There is still a catch, and it is not in the announcement: from September 14, subscribers effectively get about 17 percent less weekly quota in Claude Code than they have today – even though Anthropic presents the change as "25 percent more usage". The math behind that is further down.

Contents

What the percentages in benchmarks mean

A benchmark is a fixed set of exam tasks that every model gets under the same conditions – like a class test that several students sit. The percentage tells you how many tasks the model solved. "Terminal-Bench 4.0: 55.8%" therefore means: Fable 5.1 completed 55.8 percent of these test tasks correctly.

Three things help put the numbers in context:

  1. No model reaches 100 percent. The tasks are deliberately so hard that even the best systems fail. 55 percent can therefore be a top score – what matters is the gap to the other models, not the absolute number.
  2. Every benchmark tests something different. A model can be strong in the coding test and weak in the science test, just as a student can excel at math and struggle with English.
  3. One number in the table is not a percentage: GDPval-AA is an Elo rating like in chess. Raters compare two answers and pick a winner, which produces a ranking number. 1,853 versus 1,824 means: Fable 5.1 wins such duels slightly more often than Opus 5.

And one caveat that applies to every model announcement: the numbers come from the vendor itself. Anthropic chooses which tests appear in the announcement – plausibly the ones where its own model looks good. Independent re-tests usually appear days to weeks later.

The benchmark results at a glance

All numbers come from the Anthropic announcement of September 1, 2026. The best value is in bold.

Benchmark (what it tests)Fable 5.1Fable 5Opus 5GPT-5.6 Sol
Terminal-Bench-Science 0.1 (scientific research tasks)52.6%24.7%29.0%22.4%
Terminal-Bench 4.0 (working independently on the command line)55.8%42.0%52.3%37.3%
AutomationBench (automating business processes)31.4%17.1%26.9%19.6%
GDPval-AA v2 (office and knowledge work, Elo rating)1,8531,7231,8241,711
OSWorld 2.0 strict (operating a computer like a human: windows, clicks)41.7%36.1%39.6%
Humanity's Last Exam without tools (hard expert knowledge across all fields)60.9%57.8%56.6%
CursorBench 3.2.0 (programming in a code editor)73.4%70.5%70.0%67.2%

Where Fable 5.1 is truly far ahead – and where it is not

The table shows two very different pictures, depending on where you look.

There are big jumps in two task types. In scientific research tasks (Terminal-Bench-Science), Fable 5.1 solves more than twice as many tasks as its own predecessor at 52.6 percent – leaving Opus 5 (29.0%) and GPT-5.6 Sol (22.4%) far behind. Similar in business-process automation (AutomationBench): 31.4 percent versus 17.1 percent for the predecessor. These are exactly the long, multi-step tasks where the model has to plan on its own, use tools, and check intermediate results. Anthropic gives a research example: in designing protein binders, Fable 5.1 achieved a hit rate of nearly 50 percent, where 10 to 15 percent is considered typical.

The gap stays small in everyday coding and factual knowledge. In the editor test CursorBench, only 2.9 percentage points separate Fable 5.1 from its predecessor; in the knowledge test Humanity's Last Exam, 3.1. If you use the model for normal code and text tasks, you will barely notice the difference in daily work.

In short: Fable 5.1 is above all a model for long, independent chains of work – research, analysis, processes that take many steps. For the quick one-off task in between, the lead is small.

What Fable 5.1 costs – and why "same price" is still cheaper

The list prices via the developer interface (API) are unchanged – the savings come from elsewhere:

Item (per 1 million tokens)Fable 5.1Fable 5Opus 5GPT-5.6 Sol*
Input$10$10$5$4
Output$50$50$25$20
Cached input (cache reads)$0.25$1.00

\* Promotional price until at least November 21, 2026; list price $5/$30.

Tokens are the billing unit of AI providers – roughly a fragment of a word. And this is the real price lever: Fable 5.1 uses fewer tokens for the same work. Anthropic puts the savings at about 25 percent for typical workloads and up to 45 percent for long, independent chains of work. The financial data provider Bloomberg reports from its own tests that the model was about twice as fast as Opus 5 and used half as many tokens. On top of that comes the cache price: if you send the same documents along repeatedly – the normal case in longer working sessions – you now pay only a quarter of the previous price for those repetitions.

Which model for which task? Three recommendations

Price and test results together yield three simple rules of thumb. They reflect the state of September 2026 and are based on the vendor numbers above:

  1. Everyday tasks – texts, emails, translations, normal coding tasks: the mid-range is enough here. Claude Sonnet 5 ($2/$10 per million tokens) or GPT-5.6 Sol ($4/$20) handle this work at a fraction of the Fable cost; for short single tasks, the quality difference is barely noticeable. If you work on a Claude subscription, this also protects your weekly quota, because Fable weighs doubly heavy there (see below).
  2. Hard single tasks – a tricky bug, contract analysis, a complex evaluation: Claude Opus 5 ($5/$25) is the price-performance sweet spot. In most tests it trails Fable 5.1 by only a few points, but costs half as much.
  3. Long, independent chains of work – research across many sources, multi-step automation, scientific analysis: this is where Fable 5.1 plays its lead, twice over: it solves considerably more of these tasks (52.6% versus 29.0% for Opus 5 in the science test) and needs fewer tokens to do it. For this kind of work, the more expensive model can end up the cheaper one – measured by cost per completed task, not per request.

The quota fine print: +25% announced, −17% in practice

Alongside the model announcement, Anthropic communicated a change to usage limits: from September 14, 2026, the standard weekly limits in the programming tool Claude Code rise permanently by 25 percent – for the Pro, Max, Team, and Enterprise seat plans.

What the announcement leaves out: since May, a temporary 50 percent increase has been in place, extended several times, and it expires on September 13. The math therefore looks like this:

Weekly quota (base = 100)
Original limit100
Today (temporary +50%)150
From September 14 (permanent +25% on the old base)125

Compared with the original limit, that is indeed 25 percent more. Compared with what users have been accustomed to for months, it is about 17 percent less (125 instead of 150). Both statements are true – the announcement just tells the first and omits the second. Trade media such as BleepingComputer have retraced the calculation.

For Fable models on subscriptions, the rule remains unchanged: on Max and premium seats you may spend at most half of your weekly limit on Fable; after that you buy extra credits or switch to a smaller model. On the Pro plan, Fable runs on paid extra credits from the start. The fact that Fable 5.1 is more economical with tokens softens this in practice – it changes nothing about the limit itself.

If you are working out which AI tools pay off for your business: how the model families from Anthropic and OpenAI have evolved is covered in our articles on GPT-5.6 Sol, Terra and Luna and the eventful history of Claude Fable 5. This guide appears on office1.cloud, where freelancers also handle invoices and bookkeeping.

About the author

Charles Imilkowski

Software developer · PepperTools

Charles Imilkowski has been developing and selling his own software for invoicing and accounting since 2014, through his company PepperTools. He has also worked as a software developer since 2004, today for medium-sized companies, building interfaces between ERP systems such as SAP and accounting solutions such as DATEV.

Handle invoices more easily

Easy Invoice combines quotes, invoices and customer management in the cloud.

Try Easy Invoice

Language versions

DE Claude Fable 5.1 ist da: Was die Benchmarks wirklich sagen – und warum Ihr Kontingent ab 14. September trotzdem schrumpft ES Claude Fable 5.1 ya está aquí: lo que dicen de verdad los benchmarks – y por qué su cupo se reduce igualmente el 14 de septiembre PL Claude Fable 5.1 już jest: co naprawdę mówią benchmarki – i dlaczego od 14 września Twój limit i tak się kurczy FR Claude Fable 5.1 est là : ce que disent vraiment les benchmarks – et pourquoi votre quota diminue quand même le 14 septembre TR Claude Fable 5.1 çıktı: Benchmark'lar gerçekten ne söylüyor – ve kotanız 14 Eylül'den itibaren neden yine de küçülüyor RU Вышел Claude Fable 5.1: что на самом деле говорят бенчмарки – и почему ваша квота с 14 сентября всё равно уменьшится IT Claude Fable 5.1 è arrivato: cosa dicono davvero i benchmark – e perché dal 14 settembre la vostra quota si riduce comunque NL Claude Fable 5.1 is er: wat de benchmarks echt zeggen – en waarom uw tegoed vanaf 14 september toch krimpt