Developer at OpenAI warns of token waste in agent swarms
Eric Provencher, a developer of the Codex tool at OpenAI, warns that agent swarms with many parallel sub-agents needlessly burn tokens without improving quality. He cites an example of refactoring a Python file for 20 000 dollars using 1 393 agents, which a single agent could have handled for a fraction of the price.
Eric Provencher, a developer of the Codex tool at OpenAI, highlighted the inefficiency of so-called agent swarms on X — systems in which a large number of AI sub-agents process a task in parallel. According to him, more than two concurrently running sub-agents almost always just burn tokens without improving the quality of the result. He calls this phenomenon a “coordination tax”: agents do not trust one another and spend time excessively rechecking the work of others instead of conducting actual reviews.
As a specific example, he cites a case in which someone spent 20 000 dollars on tokens to refactor a single Python file using 1 393 agents. According to Provencher, a single agent could have handled the same task for a fraction of that price. Although agent swarms can save time according to him, excessive token overhead is, in his words, a “trap”.
As a solution, he proposes delegating subtasks to separate threads that notify the main agent only upon completion, instead of constantly polling for status. He adds that system prompts accumulate across sub-agents, and that sub-agents make unnecessary duplicate tool calls if they lack sufficient context. Provencher also acknowledges that OpenAI does not yet have a better solution to this problem and still needs to deliver one.
Why it matters
This concerns anyone who designs or uses multi-agent AI workflows: excessive parallelization of tasks across many sub-agents can multiply token costs without producing a better result. The recommendation from Provencher — limit parallelism and replace ongoing polling with notifications upon completion — offers a concrete way to reduce costs.
Two audiences, two different impacts
What this means
For individuals
According to Provencher, a developer configuring an agent workflow themselves should limit parallelism to a maximum of two sub-agents; otherwise, they risk needlessly burning tokens without any gain in quality.
For a business
Companies deploying agent workflows risk token costs that are orders of magnitude higher without a corresponding improvement in results if they unnecessarily split tasks among dozens to thousands of parallel sub-agents.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.