Fable 5.1 Leads Real-SWE at 38.8%; Code Benchmark Halves Scores

Specific Labs' new benchmark tests code agents on licensed private corpora. No frontier model exceeds 40%, with a drop of over 42 points against SWE-Bench Pro.
Specific Labs, a startup from Y Combinator's F25 batch, released on September 13 the Real-SWE, a benchmark that assesses code agents on licensed private corpora of real companies. Claude Fable 5.1 leads with 38.8% of tasks solved, ahead of GPT-6 Astra (33.8%) and Gemini 3.8 Flash (31.2%). The difficulty cut shifts the public score: in SWE-Bench Pro, Fable 5.1 records 81.2%.
What Real-SWE Measures
Each task provides the agent with a codebase it has never seen (an app with over 200,000 users, a fintech platform processing 100,000 bank statements, internal sales tools from a B2B company) and the context that a human engineer would have before starting. The median is 11 files per task. The agent needs to produce the diff that resolves the ticket; automated tests check if the intended behavior has been delivered without breaking the rest.
The point the founders of Specific Labs make clear: public benchmarks like SWE-Bench are contaminated. Frontier models have already seen, directly or indirectly, the open repositories that make up these tests. The numbers drop when the agent encounters genuinely out-of-distribution code.
What the Results Mean for Buyers of Code AI
The gap of 42.4 points between Fable 5.1 in SWE-Bench Pro (81.2%) and Real-SWE (38.8%) redefines the benchmark that enterprise buyers have been using. Purchasing an "AI engineer" plan based on public benchmark scoring means contracting performance in Django, Flask, and other repositories sampled to exhaustion. Within a 15-year-old Java monolith, or a customized SAP without documentation, Fable 5.1 currently delivers on two out of five tickets. GPT-6 Astra and Gemini 3.8 Flash deliver on one out of three.
No model exceeds half in Real-SWE. For platform teams planning to use agents as a layer for autonomous bug resolution, the real number of rollbacks, reopened tickets, and human revisions will remain high enough that the economy will depend on those who review, not those who write.
Global Reading: Who Rushed to Automate Delivery
For American model companies, the result serves a dual purpose: it validates the superiority of Fable 5.1 in coding but exposes the ceiling of any claims that the model replaces mid-level engineers. Anthropic told investors its annualized revenue surpassed $30 billion this year; the purchase made by large clients today is for marginal productivity per seat, not outsourcing entire squads.
For India, this is the curve that TCS laid off 23,460 people in FY26 to accelerate. The skepticism surrounding this bet gains ammunition: if the best model delivers 38.8% in typical client engagement private bases, TCS, Infosys, and Wipro need to keep senior engineers in the loop right where they promised to reduce bench. The margin of "AI-first delivery" does not come from the model solving the ticket alone; it comes from the remote engineer reviewing three diffs per hour instead of writing one.
For Brazil, the reading comes via BPO and shared services hubs operating for American and European clients from São Paulo and Curitiba, and from the pile of legacy systems that the Brazilian banking and insurance sector has accumulated over the past twenty years. COBOL and Java 1.6 bases running core banking would score even worse. Bradesco announces R$ 6 billion per year in technology and Itaú, R$ 2.7 billion just in the first quarter; much of this is maintenance operation in code that will never enter public model training.
Where Real-SWE Can Still Deceive
The obvious critique is sample size. Three dozen private bases of companies willing to license code do not represent the entire Fortune 500. If the mix skews toward relatively recent SaaS, the number exaggerates difficulty toward the "new" and underestimates what happens within bank mainframes. The Aider team has already pointed out a similar bias in SWE-Bench Verified.
The second point: 38.8% is the ceiling for resolution in a single attempt. Agents with multiple attempts, external retrieval, and human in the loop deliver more. The question for the CIO is not "Does Fable resolve 38%?"; it's how much each attempt that Fable misses costs. While Anthropic maintains the price list at $15 per million output tokens in Fable 5.1, each retry round counts before the economy closes.