
UC Berkeley benchmark tested AI agents on 1,490 real work assignments across 55 industries. Best system completed 26.2% correctly. Hardest tasks saw just 2.6% average pass rate.
Artificial intelligence agents fail roughly three out of four real-world work assignments, according to a new benchmark from researchers at UC Berkeley's Center for Responsible, Decentralized Intelligence. The study tested leading AI systems against 1,490 actual tasks submitted by more than 250 professionals across 55 industries -- the kind of projects that normally take hours to weeks: preparing a legal filing, building a financial model, designing a manufacturing part.
Each system was graded on whether it delivered a correct, complete, finished product. Partial credit did not exist. A missing deliverable or an error counted as a fail, the paper said.
OpenAI's Codex tool running on its GPT-5.5 model posted the best result, completing 26.2% of all assignments correctly -- roughly one in four. On the hardest tier of tasks, the kind requiring an AI to track many steps over a long stretch and get every piece right, pass rates across all tested systems averaged just 2.6%. Codex managed 8.6% on that tier. Anthropic's Claude Code tool, running on its Opus 4.7 model, failed every one of those hard assignments.
One version of the test threw enormous resources at the hardest problems: $630 worth of computing power and 763 million tokens, the units AI systems use to process text. It succeeded just 2.9% of the time. The researchers said the result shows that spending more money and computing power does not reliably make an AI agent better at this kind of work.
Drafted by a large language model from the source reporting linked above, then screened by automated publishing checks. It is not read by a journalist before publication. Some articles cite our Alpha Score. Verify prices and figures against the original source. Educational coverage, not personalized advice.