In this article
August 5, 2026
August 5, 2026

Nicholas Arcolano on why 10x the tokens buys only 2x the output

Jellyfish Head of Research Nicholas Arcolano talks to Michael Grinich at AI Engineer World's Fair 2026 about token maxing, agent regimes, and real limits.

Explore with AI
Open in ChatGPT
Open in Claude
Open in Perplexity

Jellyfish watches about 300,000 developers across roughly a thousand customers spend tokens on coding agents. The headline number from that data is uncomfortable for anyone running a token-maxing program: the top 10% of engineers are about twice as productive as the median, and they spend 10x the tokens to get there.

Michael Grinich sat down with Nicholas Arcolano, Head of Research at Jellyfish, at the AI Engineer World's Fair 2026 in San Francisco to talk about what the token-maxing era has actually bought.

The constraint stopped being headcount

Jellyfish sells observability to engineering leaders: how their teams are adopting AI coding agents, and how that adoption connects to outcomes, ROI, and productivity. Before agents, that meant the old metrics: pull requests, cycle times, and the one everybody hated, lines of code shipped per quarter, which Arcolano says was always a bad measure but was all anyone had.

The planning problem underneath those metrics hasn't gone away. Businesses still reason under finite constraints. What's changed is which constraints bind. Headcount used to be the whole list; now it runs money, people and agents, and tokens. Arcolano's framing is that an engineering leader used to plan against fixed headcount and now has effectively unlimited robot contractors and the ability to spend effectively unlimited money, and nobody has figured out how to reason about that yet.

Tokens are rocket fuel

Grinich dates the token-maxing era to roughly the last six months, starting around July. Arcolano's metaphor for it: "Tokens are rocket fuel, and token maxing is rocket science". Orbital mechanics are hard. Agent workflow complexity, waste in the system, and bottlenecks like code review all start to matter more as spend goes up, and the returns thin out.

He isn't dismissive of the metric. Raw token consumption works as a pulse check: there are levels of AI sophistication you cannot reach while doing fancy autocomplete, so a low number tells you something real. It just stops being informative fast.

Jellyfish's published research says the same thing with different cuts of the data. Across 12,000 developers at 200 companies in Q1 2026, the median user burned about 51 million tokens a month while the 90th percentile burned roughly 380 million. Throughput rose with spend, from about 0.77 merged PRs per week at the low end to 2.15 at the high end, but the median developer needed about 7 million tokens per PR against roughly 69 million for the top decile. Priced out, the cost per merged PR runs $0.28 in the lowest usage tier and $89.32 in the highest.

Three regimes, not one curve

The more useful frame from the interview is that token spend clusters into distinct regimes rather than sliding along a smooth curve.

At the top, 2% of the engineers Jellyfish tracks consume 500 million tokens per developer per week, running sophisticated multi-agent workflows with high levels of autonomy. Below that sits a regime where you run multiple agents in parallel but stay in the loop on each one, which runs into a hard ceiling: "You can only babysit so many robots". Below that, engineers are still doing fancy autocomplete because they're afraid to let the robot write the code.

Grinich pushed on whether everyone eventually climbs to the top rung. Arcolano's answer was no, at least not uniformly. He does a lot of agentic development himself in data science and struggles to get highly autonomous workflows there, because you can't prescribe upfront what a data exploration needs to be. It's an interactive back-and-forth. Agents perform very differently on infrastructure work than on front-end design, and it's reasonable to expect different results from different kinds of engineering.

The frontier numbers people quote are frontier numbers. Grinich cited a conversation with someone at Anthropic in February or March putting each of their engineers at roughly $50,000 a month in tokens for writing code, about $600,000 a year, plus a figure he'd read that morning of $2 million a year in inference per engineer. Most companies Jellyfish works with are nowhere near that.

Nobody is actually price sensitive yet

Cost panic is loud right now. Uber's CTO revealed in April that the company had blown its entire annual AI budget in four months, and Uber has since capped spending at $1,500 per employee per month per agentic coding tool, Claude Code and Cursor included.

A cap like that is a budget-line reaction. It says little about what an engineer reaches for while spending under it, and the data underneath the panic looks different. Jellyfish sees people overwhelmingly choosing the top model for every use case, with little real price sensitivity despite all the public thrashing. That's rational when the work is unknown and high risk: businesses care about quality, and picking a lesser model means wondering whether you left better code on the table. Trust in cheaper models will take time to build.

The absolute numbers are also still small. Even the top 2% are spending a couple hundred dollars in tokens (Arcolano first said per month, then corrected to per week), far less than those developers are paid. About 0.5% of the population Jellyfish tracks spends at their salary in tokens.

Arcolano reads that as a category error in the accounting rather than a signal about value. CFOs balk because a few hundred dollars is a lot for developer tooling, but this is coding capacity priced against what you pay human beings. "The spreadsheets haven't caught up yet". If a manager running agents can do two, five, or ten times the work, spending more than a salary in compute is the correct answer, and he expects the industry to get there.

The bottleneck left the IDE

Jellyfish has measured large coding gains, and those gains are now stressing everything downstream. You can build the biggest software factory in the world and it won't matter if you can't figure out what to build or sell what you make.

"We see people struggling with 2x," Arcolano said. "If you could make them 10x more productive they wouldn't know what to do with that capacity, because all these other things are throttled".

That reframes the question. It stops being how many tokens an engineer should spend and becomes what a product development life cycle looks like when code is cheap.

Infrastructure and permissioning are the wall

The easy part is over, in Arcolano's telling: hand people a good IDE or a coding agent and let the rest of the machine run the way it always has. The next step means big organizational rocks: hiring and training different people, different infrastructure.

The blockers he named are concrete: "We see a lot of people struggling to get to autonomous agents at scale because they don't have the infra for it, or the permissioning". Jellyfish's own research calls this the agentic barrier, a limit you can't spend your way past, one that takes sandboxed environments, orchestration, and context engineering to clear. An agent you can't scope, audit, or safely grant access to is an agent you keep on a human leash. That puts you back in the babysitting regime no matter how large the budget is.

Getting over that hump takes one of two things: time, or crisis. Arcolano has watched companies move fast, but that happens when a wartime CEO decides the industry demands big disruptive changes right now.

What actually moves an organization

Grinich's last question was about the teams stuck partway through: people trying to AI-pill or modernize an organization where the obstacle is the infrastructure, the culture, or a CFO who doesn't understand the spend.

Arcolano's advice was concrete. Showing is better than telling and money talks, an individual developer is more empowered than they have ever been, and putting something in front of customers that they love and will pay for is what moves the needle in most organizations.

For an engineering leader sizing next quarter's token budget, the marginal token isn't the thing to buy. The gains sit in moving more of the organization into the regimes that already work, and in building the infrastructure and permissioning that let agents run without a human watching each one.

This interview was recorded at the AI Engineer World's Fair 2026 in San Francisco.