A research team at Princeton just tested the best AI agents money can buy, across tasks mapped to real US jobs, and the full pass rate was 2.6%.
Not 26%. Not 2.6%. Two point six.
That means on economically valuable tasks, the frontier models everyone’s excited about, the ones powering your competitor’s “AI transformation” press releases, failed completely more than 97 times out of 100.
That number should make you optimistic.
What Princeton actually found
The paper, presented at ICML 2026, tested GPT 5.5, Gemini 3.1 Pro, Gemini 3.5 Flash, and Claude Opus 4.7 on a new benchmark called ALE (Agents’ Last Exam). The tasks weren’t trivia questions or coding puzzles. They were mapped to actual occupations and economically meaningful work.
The result: these models are not meaningfully more reliable than their predecessors. Bigger context windows, more training data, fancier architectures, and the agents still break down when the task gets messy, multi-step, or requires real judgment.
They also ran something called SWE-Marathon, which tests whether coding agents can stay coherent over long, sustained workloads. The answer was not reassuring. Agents that look smart on short demos lose the plot when the job takes time.
Why this matters for your business
If you’ve been sitting on the fence about AI automation, wondering whether to wait until the models get better, this research is actually good news.
It means the gap between “impressive demo” and “reliable business tool” is still enormous. And that gap is exactly where the value lives.
Raw model access, the kind you get from an API key and a prayer, is not enough. You need scaffolding. Evaluation harnesses. Human oversight where it counts. Monitoring that catches failures before your customers do.
The businesses that figure this out now, while their competitors are still bolting ChatGPT onto their website and calling it a strategy, will have a real operational advantage.
What reliable agent systems actually look like
The pattern that works is not “give the AI a task and hope.” It’s closer to how a good manager operates:
You define the task clearly. You set boundaries on what the agent can and can’t do. You build in checkpoints where a human reviews the work before it goes live. You log everything so you can diagnose failures.
The agent does the heavy lifting on the repetitive, structured parts. A human handles the edge cases, the ambiguous calls, the stuff that needs context the model doesn’t have.
It’s not glamorous. It’s not the “fully autonomous future” that keynote speakers keep promising. But it works. And it works reliably, which is the part that actually matters when you’re running a business.
The self-improvement wrinkle
There’s another development worth watching. Sakana AI just launched a Recursive Self-Improvement Lab in Tokyo, formalising the idea that AI agents can get better at improving themselves.
That’s exciting in theory. In practice, “self-improving” also means “harder to predict.” If your automation gets better on its own, you need monitoring and guardrails that keep up. Otherwise you end up with a system that’s optimised for metrics you didn’t intend.
The combination of unreliable base models and self-improving agent layers means the businesses that invest in proper agent infrastructure now will be the ones that can safely adopt more capable systems as they arrive.
The takeaway
Don’t wait for AI agents to be perfect. They won’t be, not soon.
Instead, build the scaffolding. The evaluation. The monitoring. The human checkpoints. That’s the unsexy work that turns a 2.6% pass rate into something that actually runs your business processes.
The models will keep getting better. The businesses that win will be the ones ready to use them properly.
Sources
- Princeton ICML 2026 paper, “Towards a Science of AI Agent Reliability” via smol.ai news
- Sakana AI RSI Lab launch via smol.ai news



