Agentic AI in Production: 95 Percent on SWE-bench, 21 Percent on Real Work
Francis Okafor
On this page
- What Moved: Long-Horizon Task Performance
- Tool Use Got Boring, Which Is the Real Change
- Where Computer Use Stops
- The Benchmarks Are Part of the Unsolved Problem
- Cost Predictability Is Genuinely Unsolved
- Adoption Is Narrower Than the Marketing, and Narrower Than the Critics Too
- The Strongest Case Against This Reading
- What Shenzhen Is Building On
- Tools referenced
- Sources
Two numbers describe agentic AI in production more honestly than any vendor deck released this year. Frontier models now resolve roughly 95 percent of SWE-bench Verified. On OSWorld 2.0, a computer-use benchmark published in June 2026 by XLANG Lab at the University of Hong Kong, the best system completes 20.6 percent of tasks.
Same models. Same quarter. The gap is not measurement noise and it is not cherry-picking.
It is the shape of the problem. Something real changed between 2024 and now, and something else did not move at all. Most writing on this picks one and ignores the other.
What Moved: Long-Horizon Task Performance
METR's Time Horizon 1.1 report, published 29 January 2026, gives the cleanest series available. The 50 percent time horizon, meaning the human-expert task length a model completes half the time, went from 60 minutes for Claude Sonnet 3.7 to 101 minutes for Claude Opus 4, 121 for o3, 214 for GPT-5 and 320 minutes for Claude Opus 4.5.
The doubling time is 196.5 days across the full series. Since 2023 it is 130.8 days. Since 2024, 88.6 days. That acceleration is the most defensible single claim in this entire area.
Now the part that usually gets dropped. METR rebuilt the instrument while measuring with it. The task suite grew from 170 to 228 items, tasks of eight hours or longer went from 14 to 31 and the evaluation harness moved from Vivaria to Inspect, the UK AI Security Institute framework. Confidence intervals stay wide. Opus 4.5 sits at 320 minutes with a 95 percent interval running from 170 to 729 minutes.
In March 2026 METR evaluated an early Claude Mythos Preview at a 50 percent time horizon of at least 16 hours, interval 8.5 to 55 hours. METR's own methodology page states plainly that measurements above 16 hours are unreliable with the current task suite. Five of the 228 tasks run 16 hours or longer.
The headline number arrived precisely at the point where the instrument stops working. That is not a reason to dismiss the trend. It is a reason to stop quoting it to three significant figures.
The headline time-horizon number arrived precisely at the point where METR says its own instrument stops working.
Tool Use Got Boring, Which Is the Real Change
The Qwen3-Coder-Next technical report, dated 28 February 2026, contains a table that says more about production readiness than any leaderboard position. The model scores 70.6 percent on SWE-bench Verified under SWE-Agent, 71.1 under MiniSWE-Agent and 71.3 under OpenHands. A spread of 0.7 points across three different scaffolds.
In 2024 scaffold choice was worth double digits. Prompt template, tool-call format, XML against JSON: all of it moved the number, sometimes by more than the model change did. An 80 billion parameter mixture-of-experts model activating 3 billion parameters per token is now close to scaffold-independent. The same report gives 92.7 percent template-following accuracy across five different IDE and CLI environments.
I read that report in Chinese the week it appeared. English-language coverage led with the SWE-bench figure and skipped the template-following table entirely. From Shenzhen, that table is the interesting one. If you are shipping an agent into an environment you do not control, format adherence across scaffolds decides whether the thing survives contact with a real toolchain. Leaderboard position does not.
This is the change that made agentic systems buildable. Not that models got smarter. That the interface between model and tool stopped being a research problem.
Where Computer Use Stops
OSWorld launched in 2024 with a human baseline of 72.36 percent and a best model score of 12.24 percent. By June 2026 tracked leaders sit around 85 percent, comfortably above the human baseline. Read alone, that is a solved benchmark.
OSWorld 2.0 is the same team asking a harder question. 108 workflows, median human completion time about 1.6 hours, an average of 318 tool calls per task against roughly 30 in the original. Claude Opus 4.8 leads at 20.6 percent binary completion with a 54.8 percent partial score. GPT-5.5 is more token-efficient and plateaus near 13 percent.
The failure modes matter more than the score. The paper reports that agents lose track of constraints, miss information that arrives mid-task, guess rather than ask the user and skip verification. They do not fail at GUI control. They do not fail at writing code. They fail at recovering hidden state.
Partial scores across every system evaluated cluster between 20 and 55 percent. Agents get most of the way and stop. In an operations context that is worse than failing in the first minute, because a person now has to reconstruct which 55 percent actually happened.
The Benchmarks Are Part of the Unsolved Problem
UTBoost, a test-augmentation study of SWE-bench, found 345 patches across SWE-bench Lite and Verified that the evaluation harness marked as passing and that do not actually fix the issue. The corrections touched 24.4 percent of SWE-bench Verified leaderboard entries and produced 11 ranking changes on Verified and 18 on Lite. Thirty-six instances carried test coverage too thin to catch a wrong patch.
SWE-bench Verified was built through human expert review specifically to strip out bad instances. It still shipped with test suites weak enough to pass incorrect code at that rate.
SWE-bench Pro exists because of contamination pressure, and reported scores on it diverge by roughly twenty points depending on whether you read a standardised public split or a vendor-run aggregate. When one benchmark name produces a twenty-point spread, the number has become a claim about scaffolding and reporting practice as much as about the model.
Cost Predictability Is Genuinely Unsolved
An April 2026 study ran eight frontier models over SWE-bench Verified through OpenHands and measured token consumption run by run. Repeated runs of the same task by the same model differ by up to 30x in total tokens.
Two further findings from the same work. Accuracy peaks at intermediate cost and saturates above it, so spending more does not reliably buy more. And when models are asked to predict their own token consumption, correlation with actual usage reaches only 0.39, with systematic underestimation throughout.
The practical version is that you cannot quote a per-task price. You can quote a distribution. Teams I have watched work in Shenzhen adjusted to this faster than the ones I read about elsewhere, probably because hardware people already think this way. They run the same task twenty times and look at the spread rather than the mean, which is closer to how a production line gets qualified than to how software normally gets benchmarked. Pass@1 on a leaderboard tells you nothing about whether a system survives a twelve-hour shift.
Adoption Is Narrower Than the Marketing, and Narrower Than the Critics Too
The US Census Bureau's Business Trends and Outlook Survey put AI use at 19.8 percent of firms as of 3 May 2026. By size: 37 percent for firms with 250 or more employees, 32 percent for firms of 100 to 249 and under 20 percent for the smallest. Census broadened the question in November 2025, moving from AI used in producing goods or services to AI used in any business function. The number still sits under twenty.
A Federal Reserve note published 3 April 2026 lines up three surveys of the same economy over the same period. BTOS: about 18 percent of firms at the end of 2025. The Real-Time Population Survey: about 41 percent of workers reporting generative AI use at work as of November 2025. The Survey of Business Uncertainty: 78 percent of the labor force at firms that have adopted AI, and 54 percent at firms using LLMs.
Eighteen against seventy-eight. The Fed's explanation is methodological rather than contradictory. BTOS mirrors the actual firm population, which is overwhelmingly small businesses. SBU overweights large employers, and large employers adopt most. One is firm-weighted, the other employment-weighted. Both are correct answers to different questions, and quoting either one alone produces a distorted picture.
None of them measure agents. They measure AI of any kind, including a marketing team running a chatbot. Agentic deployment is a subset of a subset, and no government statistical series currently isolates it. The confident enterprise-agent percentages circulating this year trace back to vendor surveys with undisclosed sampling. Treat them accordingly.
The Strongest Case Against This Reading
The best counter-argument is that all of this is a snapshot mistaken for a ceiling. It deserves a serious answer.
The doubling is real and it compounds. 88.6 days since 2024. If that rate holds even approximately, workflows requiring 318 tool calls are a 2027 engineering problem rather than a structural limit. OSWorld itself went from 12.24 percent to above the human baseline in about two years. OSWorld 2.0 at 20.6 percent in 2026 looks statistically like OSWorld 1.0 looked in 2024. Benchmarks get constructed to measure what is currently hard, and what is currently hard has a habit of getting solved.
The second half of the counter-argument is that benchmarks lag deployment. Coding agents became useful well before their pass rates justified it, because a human reviews the diff and the cost of catching an error is low. Measured pass rate understates deployed value anywhere verification is cheap.
Both points hold. Neither touches the load-bearing objection.
The OSWorld 2.0 failures are not capability failures. Losing a constraint, skipping verification and guessing instead of asking are failures of calibration: not knowing what you do not know. Scaling has improved capability considerably faster than it has improved calibration, and the 30x cost variance is direct evidence of the same thing. A system that cannot predict its own token consumption one run ahead is not tracking its own state well enough to be trusted across three hundred steps of it.
The cheap-verification argument also inverts at length. Reviewing a 40-step diff costs minutes. Auditing a 318-step workflow that finished 55 percent of the way costs more than doing the work yourself. The economics that made coding agents land do not extend to long workflows automatically. They have to be rebuilt, usually by decomposing the long thing back into short verifiable pieces, which is itself an admission about where the technology actually sits.
What Shenzhen Is Building On
Shenzhen's embodied intelligent robotics action plan targets industrial output above 100 billion yuan, roughly 14 billion US dollars, by 2027. More than 1,200 firms in the cluster. Over ten companies valued above 10 billion yuan. I live inside that bet. Walk the industrial parks in Nanshan or Bao'an and the assumption is physical: floors, fixtures and integration work laid out for systems whose task horizon is a shift rather than a prompt.
Nobody there is hedging on the step-count problem. The capital expenditure assumes it resolves on schedule.
It might. The trend line supports that more than it undercuts it. But the honest position in August 2026 is that this industry has learned to measure capability far better than it has learned to measure reliability, and those are not the same variable. Every vendor publishes a pass rate. Not one publishes the variance.
Tools referenced
Qwen3-Coder, reviewed here: Qwen3-Coder review.
LangGraph, reviewed here: LangGraph review.
Temporal, reviewed here: Temporal review.
Langfuse, reviewed here: Langfuse review.
Arize Phoenix, reviewed here: Arize Phoenix review.
Braintrust, reviewed here: Braintrust review.
Sources
METR, Time Horizon 1.1 (29 January 2026): https://metr.org/blog/2026-1-29-time-horizon-1-1/
METR, Task-Completion Time Horizons of Frontier AI Models: https://metr.org/time-horizons/
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks (arXiv 2606.29537): https://arxiv.org/abs/2606.29537
UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench (arXiv 2506.09289): https://arxiv.org/abs/2506.09289
How Do AI Agents Spend Your Money? Token Consumption in Agentic Coding Tasks (arXiv 2604.22750): https://arxiv.org/pdf/2604.22750
Qwen3-Coder-Next Technical Report (arXiv 2603.00729): https://arxiv.org/abs/2603.00729
US Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users (May 2026): https://www.census.gov/library/stories/2026/05/ai-use-businesses.html
Federal Reserve, Monitoring AI Adoption in the U.S. Economy (3 April 2026): https://www.federalreserve.gov/econres/notes/feds-notes/monitoring-ai-adoption-in-the-u-s-economy-20260403.html
Frequently Asked Questions
Are AI agents reliable enough for production use in 2026?
It depends almost entirely on task length. On tasks requiring roughly 30 to 40 sequential tool calls, frontier agents perform well: around 95 percent on SWE-bench Verified and about 85 percent on the original OSWorld computer-use benchmark, above its 72.36 percent human baseline. On OSWorld 2.0, released in June 2026 with workflows averaging 318 tool calls and a median human completion time of 1.6 hours, the best system completes only 20.6 percent of tasks. Agents are production-ready for short, verifiable tasks where a human reviews the output cheaply. They are not yet reliable for long autonomous workflows.
Why do AI agents fail on long-horizon tasks?
The OSWorld 2.0 paper found the failures are not about capability. Agents do not fail at controlling a GUI or at writing code. They fail at calibration and state tracking: they lose track of constraints given earlier, miss information that arrives partway through a task, guess instead of asking the user for clarification and skip verification of their own work. The paper notes agents struggle most when a task depends on hidden state that requires recovery. Partial scores cluster between 20 and 55 percent, meaning agents typically get most of the way through a long workflow and then stop without finishing.
How unpredictable are AI agent costs per task?
Highly unpredictable. An April 2026 study of eight frontier models running SWE-bench Verified through the OpenHands framework found that repeated runs of the same task by the same model can differ by up to 30x in total token consumption. The study also found that accuracy peaks at intermediate cost and saturates above it, so spending more does not reliably improve results, and that models predicting their own token usage correlate with actual consumption at only 0.39 while systematically underestimating. In practice this means agentic workloads can be priced as a distribution but not as a fixed per-task cost.