Tokenmaxxing: What It Is, When It Pays Off
Cicero Campelo, CISSP
September 28, 2026 · 13 min read
Part of our guide to AI for startups.

Table of contents
- What tokenmaxxing actually means
- The case for tokenmaxxing is a startup case
- What a real tokenmaxxing experiment cost
- The incumbent retreat is real and it is not your playbook
- Where the tokens should actually go
- When tokenmaxxing is just burn
- Budget visibility is also a security control
- What to do this week
- Sources
- Frequently asked questions
Tokenmaxxing is the practice of deliberately maximizing what you spend on AI tokens, on the logic that the output those tokens produce is worth far more than the bill. It describes a way of operating: stop rationing model access, reach for the strongest model, and treat the token line as the thing producing the work rather than as overhead to minimize.
The term also carries a second, less flattering meaning: treating raw token volume as a score. That version is the one being dismantled right now, and the two keep getting confused, which is why the honest answer to whether tokenmaxxing is dead depends entirely on which one you mean.
The most useful framing of the split comes from Mike Mignano, a general partner at Union Square Ventures, on 20VC with Harry Stebbings. Mignano joined USV in April 2026 from Lightspeed, and before investing he co-founded the podcasting company Anchor, which Spotify acquired in 2019. His position is not that everyone should spend without limit. It is that the answer is different depending on how many people you employ.
What tokenmaxxing actually means
Two things travel under the name, and they behave differently.
The first is a resource-allocation claim: for a small team, model spend is the cheapest input you have relative to what it produces, so rationing it is optimizing the wrong line. Our own write-up of revenue per employee covers the best-known version of that argument, from YC's Garry Tan, who compares the token bill to San Francisco rent. It looks expensive right up until you notice it is more expensive not to pay it.
The second is a status claim: high token consumption as evidence of being a serious builder. That one has no mechanism behind it. Tokens consumed is an input measure, and an input measure with no output measure next to it will drift, because the easiest way to raise it is to do more work of the kind nobody checks.
Everything that follows is about keeping the first and killing the second. They are not the same practice, and the companies that got burned in 2026 mostly bought the second while believing they had bought the first.
The case for tokenmaxxing is a startup case
Mignano's argument starts from headcount rather than from conviction. For large companies, he thinks the arithmetic simply does not work: "there's too many employees at a Salesforce or I don't know a Microsoft such that every employee can just have an unlimited token spend budget." His conclusion for that group is that they will pull back, and Stebbings interjects with three names, Meta, Uber and Microsoft. Worth saying that the three did not do the same thing. Uber capped spend, Meta shut down an employee-built internal leaderboard that ranked people by token usage, and Microsoft, according to reporting on an internal memo from executive Jay Parikh, told staff that "tokenmaxxing is not what we are optimizing for".
For a startup he flips it, and the reason is worth separating into two parts. The first is that you can actually watch the spend. A twenty-person company is not going to discover a runaway bill across five thousand engineers, so the control problem is tractable in a way it is not at scale.
The second part is competitive, and it is the part founders underweight: "as a startup you need every advantage you can get right now." Put next to the incumbent retreat, that produces a genuinely asymmetric situation. If the large company across the market from you has just capped its engineers at a fixed monthly budget, then your willingness to spend stops being only a cost decision and becomes a difference in the rate at which the two of you can try things. Mignano's own version of this is specific about where it applies: he wants the frontier model for coding, because, as he puts it, "I want every advantage I can get against Salesforce."
The same logic shows up on the team side. Mignano expects engineering organizations to end up with "somewhat smaller teams of higher caliber and higher quality engineers" as lower-level tasks get delegated to agents, which is the same shift we traced through revenue per employee and through the broader move to an AI-native company. A small team of strong people with an unmetered frontier budget is a specific competitive machine, and it is one a company with fifty thousand employees cannot assemble by policy.
It is worth being precise about how far Mignano takes this, because he does not take it all the way. His enthusiasm is for coding, and he reaches for a cheaper model for simple tasks like summarization or operations work. We covered the cost mechanics of that split in detail in AI inference cost, including his forecast that a flatter capability curve would push buyers toward routing and open models. In the same interview he argues the other way for the present, and holds both positions without contradiction, because one is conditional on the capability curve flattening and the other is about the situation in front of you now.
What a real tokenmaxxing experiment cost
The most instructive account of tokenmaxxing taken to its limit comes from someone whose job is measuring models. Rayan Krishnan is co-founder and CEO of Vals AI, an independent evaluation company founded in 2024 that builds private, held-out benchmarks for enterprises, and which a16z announced it was backing in 2026. Speaking on an a16z panel, he described running the experiment on his own team: "I wanted to do a token maxing experiment." He got the team unlimited access to some of the coding tools for a month.
The usage numbers are the part that makes the abstract concrete. He describes engineers spending between 1 to 2 billion tokens a day, and says "peak day was one engineer spending six billion." When he did the arithmetic afterward, the month came to roughly $1.5 million worth of tokens. The access itself was comped for the experiment, so that figure is what it would have cost rather than a bill that got paid, which makes it a cleaner measurement than most: it is the true appetite of a team with no budget in the way.
Then the number that reframes the whole debate. That month of tokens was, in his words, "10x more we were spending in tokens than employee salary for that month." Not a split, not a large line item. Ten times payroll. His reaction was operational: "we cannot continue with this mode of operation for the next month."
What he did next is the part worth copying, because it was not a cap. His team looked through the traces and their own GitHub history, built an internal benchmark from their real code base, and used it to work out which tools were worth which tasks. They found that Cognition's Devin was notably token efficient and adopted it more widely, and found that subscription pricing beat token-based pricing in some places. He describes the goal of all that work as learning how to "effectively token max without spending $1.5 million per month."
The operating mode they landed on kept the access open. Everyone still gets everything. What changed is that the team now gets an automatic recommendation for where to begin a session on any given ticket, which, as he puts it, "should titrate the actual usage depending on the intelligence required for that task." Routing by task difficulty rather than by permission slip.
One finding from that work should change how you read model price lists. Krishnan says that in a lot of cases "sonnet is more expensive than opus because it is so token hungry". The cheaper-per-token model lost on total cost because it consumed more tokens to finish the job. Per-token price is not cost, a point we unpack at length in AI inference cost, and it is the single most common way a team convinces itself it is being frugal while spending more.
Krishnan's broader warning is the one to keep: the direction of travel is one where "token spend may start to eclipse salary spend", and a line item that large has to justify its return far more keenly than it did when it was a rounding error.
The incumbent retreat is real and it is not your playbook
The 2026 pullback is well documented, and reading it correctly matters, because the headlines make it sound like a verdict on the practice.
Uber burned through its entire 2026 budget for AI coding tools in about four months and then capped internal use of the software. What its president and COO Andrew Macdonald said about it is the more useful part. Asked whether the spending was producing more shipped product, he told Fortune that "That link is not there yet", adding that it is very hard to draw a line from token usage to, in his words, "25% more useful consumer features", and that without that line the trade against headcount gets harder to justify. At Meta, Adam Mosseri has said per-engineer token budgets could be capped. And Salesforce, the company Mignano uses as his example, is on the other side of the same trade: Marc Benioff said Salesforce expects to spend about $300 million on Anthropic tokens in 2026, most of it on coding, while engineering hiring stays paused.
Stebbings pushes Mignano on exactly what that implies, framing the question as what share of developer salary a company's token spend represents and what happens to the labs' revenue if that share keeps climbing. It is the right question, and Krishnan's measured 10x is the most concrete answer anyone has put on it.
But notice what the retreat actually is. Uber concluded something much narrower than that agentic coding was worthless: an uncapped budget across a very large engineering organization is not a budget. Those are different findings, and only the first would be a reason for you to spend less.
Uber's own next chapter settles it. In August its CTO Praveen Neppalli Naga said the company was coming to the end of the so-called tokenmaxxing era, and the evidence he gave was not a smaller bill. It was four times as many employees on frontier AI tools as at the start of the year, at a lower cost per token, reached through prompt caching, changed default models, evaluating newer models for efficiency, and giving engineers visibility into their own usage and hourly cost. The company the headlines cast as the retreat ran the instrumentation play and ended up with more people on the best models, not fewer.
The counter-example worth studying is the one that got the bill down without getting the usage down. Coinbase reported cutting its AI spend by roughly half while token usage kept climbing, and it did that with five levers working together rather than with caps. We break those down in open source LLM, but the structural point belongs here: only one of the five was a change of model. The rest were routing, caching, leaner context, and per-engineer visibility, deliberately without usage limits. That is the version of discipline that is compatible with tokenmaxxing, and it is available to you at any size.
Where the tokens should actually go
If the goal is maximum spend on the things that pay and minimum elsewhere, you need a view on which tasks are which. Two of the people with the best data on that question land in the same place from opposite directions.
Mignano's estimate is that "80% of non-coding tasks in the enterprise can be done with models that are not at the frontier", while coding is where he would stay on the frontier. Jeffrey Morgan, co-founder and CEO of Ollama, gave Y Combinator's Lightcone a version with the volumes and the money separated, which is the more useful shape: "The super majority of tokens and this is our take it will be open models within a business. Call it 80 90%. That doesn't mean 89% of the the budget will go to open models." Note the denominators before you put his number next to Mignano's: Morgan's 80 to 90 is a share of token volume, where Mignano's 80 is a share of tasks. Morgan expects open models to carry most of that volume while taking maybe 10 to 20% of the budget, with the hardest work left to the frontier labs.
Put those together and the target state is a barbell rather than a single spending posture. Most of your token volume runs on cheap or open models where the work is well specified and the stakes are low, and a small share of your volume runs on the best model available with no hesitation, on the work that decides whether your product is good.
The Lightcone conversation lands on the line worth keeping, and it describes the opposite of how a capped organization behaves: "You're not thinking about taking away token access from your team. You're giving more and more access." (The speaker turns are merged at that point in the transcript, so credit it to the conversation rather than to either voice.)
Coding sits on the frontier end of that barbell for a reason that is easy to miss and important to state, because it is the reason the whole practice works where it works. Code has a cheap, automatic verifier. Tests run, builds pass or fail, the diff either does the thing or it does not. So when you spend more on a coding task, a machine checks whether the extra spend produced anything, and the loop closes without a human in it. That is also why the discipline in AI testing is upstream of the spending question rather than separate from it: the strength of your verifier sets the ceiling on how much you can usefully spend.
Where there is no verifier, more tokens buy more output and no more certainty. That is the actual dividing line, and it explains the pattern better than the coding-versus-everything-else shorthand does. A support reply that a customer either accepts or escalates has a verifier. A strategy memo nobody scores does not.
When tokenmaxxing is just burn
Four conditions separate the two, and they are all things you can check this week.
There is a verifier, or there is not. If nothing automatically evaluates the output, extra spend is not buying quality, it is buying volume. Build the check before you raise the budget. The strongest form is an internal benchmark built from your own work, which is exactly what Krishnan's team did after their experiment.
The spend is on the critical path, or it is not. Frontier prices are worth paying where a customer can tell the difference and where the work decides the quarter. Mignano's split, frontier for coding and cheaper models for summarization and operations, is a rough version of this test.
Someone can see it per person and per task. Not to cap it. To know what it bought. Coinbase's per-engineer visibility without caps is the model here. Visibility is what lets you keep the budget open honestly, because the alternative to knowing is eventually a cap imposed in a panic. The thing to avoid is a uniform per-person allowance, not a runtime ceiling: budgeting per class of work and enforcing it where the agent runs, as AI agent infrastructure sets out, is instrumentation rather than rationing.
You are measuring cost per finished unit of work, not per token. The token-hungry cheap model is the trap, and it catches careful teams more often than careless ones, because it looks like prudence.
If all four hold, spending more is the rational move and hesitating is the expensive one. If none of them hold, no budget is safe, and the company that caps you will be your own finance team in about four months.
Budget visibility is also a security control
One thing rarely shows up in the spending conversation and should. An agent with an unmetered budget is also an agent with an unmetered blast radius.
The instrumentation that tells you what your tokens bought is mostly the same instrumentation that tells you what your agents did: which keys were used, which tools were called, which repositories and systems were touched, by which identity, on whose behalf. A team that cannot attribute spend per person and per task usually cannot attribute actions per person and per task either, and that second gap is the one that turns into an incident.
So treat the visibility work as dual-purpose rather than as finance overhead. Per-identity attribution, per-task traces and retained logs are a cost tool on Monday and a forensic tool on the day you need one. It is also the practical reason to prefer visibility over caps: a cap tells you nothing about what happened, while a trace does, and only one of those two survives contact with an incident review. If you are building the wider stack around this, internal AI infrastructure for startups covers where these controls sit.
Mignano's own framing of the era is a reasonable place to leave it: "this era of building AI products is in many ways about being first and about moving really really fast." Speed bought with tokens is real speed. It is just worth knowing which of your tokens are buying any.
What to do this week
- Pull last month's token spend and divide it by your payroll for the same month. You now have Krishnan's ratio for your own company, and a number to argue about with real stakes.
- Pick your single highest-value workflow and ask whether it has an automatic verifier. If it does not, building one is this week's work, not raising the budget.
- Build a small benchmark from your own history. Take ten tasks your team actually completed, replay them against two or three model and tool combinations, and record cost per completed task rather than cost per token.
- Turn on per-person and per-task spend visibility, and explicitly do not add caps. Tell the team the numbers are for learning, not rationing, then hold to that.
- Split your workload into the barbell. Name the work where you will always use the best model available, move the well-specified remainder to cheaper or open models, and write down which is which so the default is a decision rather than a habit.
- Check that your spend traces carry identity and tool-call detail, so the same data answers a finance question and a security question.
If you want the wider system this fits into, from choosing models to wiring agents into your operations, that is what AI for startups is about, and it is the backbone of the AI Operating System for Startups course.
Sources
- Why Now is the Time for the App Layer | Why Startups Should be TokenMaxxing, Mike Mignano on 20VC with Harry Stebbings, the interview this article distills. Profile: Union Square Ventures; background on Anchor and the Spotify acquisition: Wikipedia; his move to USV reported by Bloomberg.
- Inside the Race to Measure Frontier Intelligence, a16z, for Rayan Krishnan's account of the token maxing experiment, the 10x salary ratio and the token-hungry model finding. Company background: Vals AI, TechCrunch and a16z's investment announcement.
- Open Models Change The Economics of AI, Y Combinator's Lightcone with Jeffrey Morgan, for the split between share of tokens and share of budget. Profiles: Jeffrey Morgan and Ollama.
- Uber's budget, its president and COO's comments and the subsequent cap: Fortune, Business Insider and The Washington Times.
- Uber CTO Praveen Neppalli Naga on the end of the tokenmaxxing era and what replaced it: Business Insider and Fortune.
- Meta's internal token leaderboard and Microsoft's internal memo: LeadDev.
- Meta's position on per-engineer token budgets: TechCrunch.
- Salesforce's token spend and hiring stance: Business Insider and The Next Web.
Frequently asked questions
Is tokenmaxxing dead?
The corporate version of it is being dismantled and the startup version is not. What died in 2026 was treating raw token volume as a productivity score with no budget attached to it. Uber burned its entire 2026 budget for AI coding tools in roughly four months and then capped internal use, with its president and COO Andrew Macdonald saying the link between token usage and shipped product is "not there yet", and Meta's Adam Mosseri has said per-engineer token budgets could be coming. That is a headcount problem rather than a verdict on the practice: a company with 50,000 employees cannot hand every one of them an unlimited budget, and a company with 20 can. The more accurate reading is that tokenmaxxing got instrumented rather than cancelled, and Uber itself is the proof. In August its CTO Praveen Neppalli Naga said the company was coming to the end of the tokenmaxxing era, and the evidence he gave was four times as many employees on frontier AI tools as at the start of the year at a lower cost per token, reached through prompt caching, changed default models and per-engineer cost visibility. Vals AI co-founder and CEO Rayan Krishnan, whose own month of unlimited access produced roughly $1.5 million worth of tokens, did not answer it with caps either. He built measurement instead, and describes the goal as being able to "effectively token max without spending $1.5 million per month".
Does tokenmaxxing work for tasks other than coding?
It pays off far less clearly outside coding, and the people closest to the spend say so. Mike Mignano, a general partner at Union Square Ventures, would push a startup hard toward frontier models for coding and reach for something cheaper elsewhere, naming summarization and operations work as the cases where the frontier is not worth it. His estimate for the enterprise is blunt: "80% of non-coding tasks in the enterprise can be done with models that are not at the frontier." Coding is the outlier for a specific reason rather than a cultural one. A coding task has a cheap, automatic correctness check, so more spend converts into verified work rather than into more plausible text. Where you have no such check, extra tokens buy you more output and no more certainty, which is spend without a mechanism behind it.
Should a big company tokenmax the way a startup does?
No, and the reason is arithmetic rather than courage. Mignano's version of the split is that the fundamentals break at scale: "there's too many employees at a Salesforce or I don't know a Microsoft such that every employee can just have an unlimited token spend budget." A startup can hold the opposite position because its organization is small enough to watch, and because it is the party that needs an edge. What a large company can copy from the startup posture is the part that is not about unlimited budgets. Coinbase reported cutting its AI bill by about half while token usage kept rising, and it did that with cheaper default models, routing, caching, leaner context, and per-engineer spend visibility, explicitly without usage caps. Visibility without a uniform per-person allowance is the transferable move, which is a different thing from having no runtime ceiling at all: budgeting per class of work and enforcing it where the agent runs is instrumentation, not rationing.
How do you tell tokenmaxxing from burn?
Ask whether the spend has a verifier and an owner attached to it. Spend is an investment when something automatically checks the output (a test suite, a build, a customer acceptance), when the task is on your critical path, and when someone can see what was produced per dollar. It is burn when the tokens are going into work nobody checks, when the model is being chosen by habit rather than by fit, and when nobody can answer what last month bought. The most useful signal is that token efficiency and per-token price are different things. Krishnan's team found that in many cases a cheaper-per-token model was the more expensive choice because it was, in his words, "so token hungry". If you cannot see cost per completed unit of work, you are not tokenmaxxing, you are just spending.
Build your AI Operating System
A practical course to grow with AI, build internal tools, and operate safely. Join the waitlist and you'll be first in when the course opens.