Moody office desk at night with scattered invoices, receipts, calculator and laptop showing faint data visualisations under warm amber lamp light

Your AI agents are quietly burning money while you sleep

Here is something nobody told you when you set up that fancy multi-agent workflow last month: every subagent your system spawns might be running at the most expensive setting available, and you are paying for it.

This week’s AI launches made the problem impossible to ignore. OpenAI shipped GPT-5.6 with a new tiering system: Luna, Terra, and Sol, each with effort levels including Max and Ultra. The reviews are mixed, but the operational pain point that keeps surfacing is not the model quality. It is what happens when agents start spawning other agents.

Sol Ultra, OpenAI’s top-tier orchestration mode, can automatically spin up subagents to parallelise work. That sounds great until you realise those subagents inherit the same premium settings by default. Users are reporting that their quotas drained two to three times faster than expected because a single task quietly spawned a dozen Sol Ultra workers behind the scenes. Nobody asked for that. Nobody budgeted for it.

This is not an OpenAI-specific problem. It is an orchestration problem, and it affects anyone running multi-agent systems in production.

The orchestration tax

Anthropic published benchmarks this week showing that their Fable 5 model, used as an orchestrator with cheaper Sonnet 5 workers, achieves 96% of the performance of running Fable 5 everywhere at 46% of the cost. That is not a small saving. On BrowseComp, the numbers were 86.8% accuracy at $18.53 per problem versus 90.8% at $40.56 per problem.

The pattern is simple: use an expensive model to plan and verify, use cheaper models to do the grunt work. But if your orchestration layer does not enforce that split, the expensive model does everything, including the tasks a $1-per-million-token model could handle.

Perplexity’s Arav Srinivas put it well this week: the real product is now the harness around the model. He is right. Frontier models are converging in capability. What separates a cost-effective agent deployment from a money pit is the routing, memory, tool use, and delegation logic sitting on top.

What this looks like for business operations

If you are running AI agents for anything operational, quote generation, customer support triage, invoice processing, contract review, you are likely in one of two camps:

Camp one: You use a single agent per task. It works, it is predictable, and you know roughly what it costs. Good.

Camp two: You have started experimenting with multi-agent setups where one agent plans and others execute. This is where the cost explosion hides. If your planning agent is your most expensive model, and it delegates by spawning copies of itself, you have a cost multiplication problem that gets worse with task complexity.

The fix is not to avoid multi-agent patterns. They work. The fix is to build a harness that controls which model gets used for which role, what effort level each agent runs at, and how many subagents can be active at once.

The practical takeaway

The AI industry is moving fast toward a world where the harness matters more than the model. Sebastian Raschka’s latest tutorial on local coding agents makes the same point from a different angle: open-weight models like Qwen 3.6 are now good enough for real work, and running them locally gives you complete control over cost and behaviour.

For businesses, the lesson is straightforward. Before you add another AI agent to your stack, ask three questions:

  1. What model and effort level will each agent use?
  2. Can it spawn subagents, and if so, what settings do they inherit?
  3. Who is watching the cost?

If you cannot answer those, you are not running an automation. You are running a slot machine that pays the house every spin.

The companies getting value from AI agents right now are not the ones with the best models. They are the ones with the tightest orchestration. The model is the engine. The harness is the car. And right now, most businesses are driving Formula 1 engines in first gear while the fuel gauge drops.

Sources