AI Capex: What the Buildout Means for Startups
Cicero Campelo, CISSP
October 5, 2026 · 13 min read
Part of our guide to AI for startups.

Table of contents
- What AI capex is
- A bubble and a real buildout are not mutually exclusive
- The supply side: paybacks inside a year
- The demand side is where founders misjudge the gap
- What a supply constrained decade does to your cost curve
- Capacity is a procurement problem with a lead time
- Own your intelligence, or rent it at someone else's price
- The risk that is not in anyone's model
- What to do this week
- Sources
- Frequently asked questions
Ask whether AI capex is a bubble and you get an argument about stock prices. Ask what the capex cycle is doing to your cost of goods, your capacity lead times and your margins, and you get a plan.
Those are different questions, and founders keep getting handed the first one. The useful version of the discussion happened on a16z's own show, where David George, the general partner who runs the firm's growth practice, spent an hour with Gavin Baker, founder and chief investment officer of Atreides Management, working through the buildout from both sides. Baker spent 18 years at Fidelity, including eight running its OTC Portfolio, before starting Atreides in 2019, so he has watched a few of these cycles price themselves.
Their conclusion is uncomfortable for both camps: the bubble people and the boom people are describing the same thing, and neither answer tells you what to do on Monday.
What AI capex is
AI capex is the capital spent on the physical capacity that produces AI: data centers, accelerators, networking, memory, power generation and the land underneath. It is the spending that turns into the tokens you buy.
The reason it dominates the conversation is scale. The pricing unit stopped being the GPU and became the megawatt of power, and the round numbers the two investors work with on air are roughly 50 billion dollars of capital per gigawatt, of which about 35 billion is the compute hardware itself. Those are illustrative figures in a conversation, not audited ones, but the order of magnitude is the point: the input to your product is now priced like heavy industry.
That matters to you for one reason. When your supplier's cost structure looks like a utility buildout, your cost curve stops being a software cost curve.
A bubble and a real buildout are not mutually exclusive
Both things can be true at once, and Baker's framing is the clearest short version of why:
"Every time you've had a real profound new technology, you get a bubble because the markets get really excited and they get ahead of themselves. Things get overvalued. That overvaluation leads to an overbuild."
Read that carefully, because the causality runs in an order most founders skip. The bubble is in the valuations. The overbuild is the consequence. Neither claim says the underlying demand is fake, and that distinction is the whole game for someone building on top. Railroads were a genuine bubble and also genuinely how freight moved for the next century.
The second half of Baker's position is empirical rather than historical. Through July and August he actively hunted for a bearish case and could not assemble one, using a question he puts to everyone he meets: "can you tell me one quantitative data point in your business that's getting worse? Just one."
By his account nobody could, which is a useful discipline to borrow. It separates a sentiment argument from a data argument. If every operating metric you can observe is accelerating while asset prices wobble, you are looking at a repricing, not a demand collapse. Those require completely different responses, and only one of them should change your roadmap.
One financing detail is worth holding onto, because it is the thing that would actually turn a repricing into a contraction. Baker notes that a majority of the buildout so far is still funded out of operating cash flow rather than debt, and that debt-funded capacity is the dangerous kind: it demands return on a much shorter clock than a three-year payoff. If you want one indicator to watch that is more informative than any valuation headline, it is the share of new capacity being financed with debt.
The supply side: paybacks inside a year
The reason capex keeps coming regardless of sentiment is that the returns on it are, right now, extraordinary. Baker works through the public disclosures and gets to a nine to ten month payback. George's summary of where that leaves the supply side is blunter: "there's a ton of data points out there that paybacks are within a year."
Three things make that possible, and all three are durable for a while:
- Upfront customer commitments. Capacity gets pre-sold, so a meaningful share of the capital comes back before the facility is fully monetized.
- Financeability. The independent builds are the other case: on the round numbers above, a 50 billion dollar data center needs roughly a 15 billion dollar equity check and the rest can be borrowed, at low spreads, from institutional lenders. Baker is dismissive of the circular-financing worry on exactly this ground, pointing out that the people underwriting these deals are the credit desks at firms like Blackstone, KKR and Apollo.
- Extending useful life. Older accelerators keep earning, because demand for anything that computes has not let up. That pushes the asset's lifetime out and the return up.
Here is the founder-relevant inversion. Sub-one-year paybacks mean supply will keep getting added aggressively. It also means that capacity is being allocated to whoever can commit earliest and largest, which is not you. An investment case that good for the supplier is a seller's market for the buyer.
The demand side is where founders misjudge the gap
This is the part of the episode most worth your time, because it reframes a number almost everyone gets wrong. Baker puts the number of genuinely heavy, paying AI users at around 30 million. George takes the under and argues it is probably below 10 million, then sets that against the addressable population:
"There's one and a half billion knowledge workers." His conclusion: "it feels like we're nowhere on the demand side and we're massively supply constrained."
The diffusion data backs it up and is more actionable than the headline. On a later a16z episode, George gave the spread directly from the dataset his team has seen: the median company in the United States spends about 12 dollars per employee per month on AI, while the top 1 percent of that dataset spends about 7,000 dollars per employee per month. That is not a gap in enthusiasm. It is a gap of nearly three orders of magnitude in how much work has actually been moved onto models.
Measured against payroll rather than per seat, George puts old-economy companies that are doing a decent job at roughly 1 percent of human compensation spent on tokens, and genuinely AI-native companies at high single digits to above 10 percent. Baker's own firm is a data point on the steep end: "our internal token consumption has gone up 100x from the month of March" through August.
Three practical consequences:
- You now have a benchmark. Token spend as a share of fully loaded compensation is a number you can compute this week. Under 1 percent and you are running an old-economy cost structure with an AI logo. The threshold that matters is not an absolute dollar figure, it is whether the ratio is moving.
- Your buyers are mostly at 12 dollars per employee. If you sell AI software to the median company, you are not competing with a sophisticated incumbent workflow. You are competing with almost nothing, which is a much better problem than it sounds and a much slower sale than you would like.
- The heavy users are heavy. Within a company, the top engineers consume an order of magnitude more than the median. Pricing per seat against a population distributed that unevenly will either cap your revenue or destroy your margin, and which one it is depends on who signs up. The four tests that separate productive token spend from burn are in tokenmaxxing.
What a supply constrained decade does to your cost curve
Most financial models built in the last two years contain an assumption nobody argues with: cost per token falls forever. The capex discussion is a direct challenge to it.
The worry voiced in the episode is not oversupply but the opposite. George's view of the macro picture is that capacity gets underbuilt through 2028, with forecast builds likely to slip on permitting and local politics. On the earlier a16z episode, George put the lead time in concrete terms: he said he felt confident the market was not in a bubble at that point, and that it is massively supply constrained, with data center capacity at scale unavailable until late 2028 or early 2029.
If that holds, the conversation's logic follows uncomfortably: in a market where demand outruns supply for years, the price of access to frontier intelligence can rise rather than fall.
That is not a prediction I would build a company on, and the two of them do not treat it as a base case either. The useful move is narrower: stop treating falling token prices as a load-bearing assumption. Model a flat frontier price. Then ask what breaks.
For most products the answer is that nothing breaks if you can move work between models, and everything breaks if you cannot. Which is an argument for building the routing seam now, while it is cheap, rather than when your margin needs it. The levers that actually move the bill, as opposed to the per-token price on a rate card, are in what moves AI inference cost.
The direction of the per-token price is also the weakest part of this thesis to lean on. Model efficiency improvements have repeatedly beaten capacity constraints, and a genuine algorithmic breakthrough that shrinks models is the one development that would break the supply-constrained story. Plan for flat, not for rising.
Capacity is a procurement problem with a lead time
The clearest illustration of what a supply crunch does to an early-stage company came from a different source entirely. Explaining why Y Combinator set up a dedicated GPU cluster with Together AI, the YC side of that conversation put the problem in numbers: for a lot of companies, the "upfront they would need to pay in order to secure the capacity for their next two years of compute was greater than their current cash balance."
That sentence should change how you plan. The constraint is not the hourly rate. It is that the vendor wants a 24-month commitment, paid substantially up front, from a company whose demand curve 24 months out is unknowable. YC's answer was to aggregate its portfolio's demand so an individual startup could make a few-month bet instead of a two-year one.
Three things to take from it:
- Start the conversation early. If something launching in six months needs dedicated capacity, that procurement starts now. Lead time, not price, is the binding constraint.
- Aggregate if you can. An accelerator, an existing cloud commitment or a group of portfolio companies can buy terms a single seed-stage company cannot.
- Price the optionality. A shorter commitment at a worse rate is often the right trade when your own demand forecast is a guess. What to check before signing one is covered in the GPU cluster decision.
Own your intelligence, or rent it at someone else's price
The strategic conclusion Baker draws from all of this is the one that generalizes best beyond investing. Open-weight models have closed much of the gap to the frontier, and post-training on proprietary data is now a practical option rather than a research project. His framing:
"if intelligence is like a super important input into your business, you want to own and control your intelligence, its capabilities, its cost."
Two honest caveats before anyone acts on that. Owning a model is a real and recurring cost: the base model has to be swapped out as better ones appear, the fine-tuning has to be redone, and the serving infrastructure is yours to run. For most startups, most of the time, renting frontier capability is still correct, and the cheaper first step toward differentiation is retrieval rather than training. Where that line sits is the subject of RAG versus fine-tuning, and which tasks should stop calling a frontier model at all is covered in where open source LLMs belong.
The second caveat is the one I would push hardest as a security matter. Baker flags that handing your accumulated enterprise context to a frontier lab may be hazardous to your business, and he is right that it is a strategic exposure and not only a compliance checkbox. The context embedded in your data is frequently the most defensible asset you own. Before that data leaves your boundary, read the retention and training terms you are actually on rather than the ones on the marketing page, confirm whether zero-retention applies to your tier and to every endpoint you call, and keep the data classification decision with someone accountable for it. That review is cheap now and expensive to retrofit.
The risk that is not in anyone's model
The episode's sharpest warning is about what a prolonged shortage does to access rather than to price. If capacity stays scarce, the consequence could be "real compute inequality where big companies and wealthy people can afford compute." In a rationed market the smallest buyers are the ones squeezed first, which is my own reading rather than his.
That is the asymmetry to plan around, and it is the opposite of the one founders usually worry about. The risk to a startup from the capex cycle is not that the bubble pops and compute gets cheap. A glut would be good for you. The risk is that it stays tight, that frontier capacity is allocated by who can pre-commit at a scale you cannot, and that your product's unit economics are set by a queue you are at the back of.
Which brings the whole thing back to a single decision. Treat compute as a strategic input with a procurement function, a budget line and an owner, the way you treat payroll. Companies that do that will be in the queue early. Companies that treat it as a variable cost on a rate card will find out what it costs when they have no alternative. The wider frame for running a company this way is our pillar on AI for startups.
What to do this week
- Compute your token ratio. Total model spend divided by fully loaded compensation, for last month. Write the number down. If it is under 1 percent, decide deliberately whether that is a choice or a drift.
- Re-run your model with a flat token price. Not rising, just flat for 24 months. If gross margin breaks, you have a routing problem to solve, not a forecast to revise.
- Find your heaviest 5 percent of users and price them. Per-seat pricing against an order-of-magnitude usage spread is a margin accident waiting for a big customer.
- Ask your inference vendor two questions. What is the lead time for dedicated capacity at 10 times our current volume, and what is the shortest commitment term you will sell. The answers tell you more about the market than any capex headline.
- Audit one data path. Pick the single highest-value proprietary dataset you send to a model provider and confirm, in the contract rather than the docs, what happens to it.
- Pick your one quantitative data point. Put Baker's question to your own team: which number in our business is getting worse. If you cannot name one, you are not measuring enough of them.
The course is the structured version of this: AI Operating System for Startups covers how to budget compute, instrument the spend and build the seams that keep a model swap cheap.
Sources
- Why AI Demand Is Outrunning Compute Supply (a16z), the conversation between David George and Gavin Baker that this article distills.
- The New Rule for Picking AI Winners (a16z), for George on the bubble question and on data center capacity being unavailable at scale until late 2028 or early 2029.
- Why Investors Are Rethinking Everything for the AI Era (a16z), for the diffusion spread between the median company and the top 1 percent of AI spenders.
- The First Dedicated YC GPU Cluster, with Together AI (Y Combinator's Root Access), for the two-year compute commitment problem and YC's aggregation answer.
- Profiles: David George at Andreessen Horowitz and Gavin Baker at Atreides Management, for current roles and background.
Frequently asked questions
What is AI capex?
AI capex is the capital spent building the physical capacity that produces AI: data centers, accelerators, networking, memory, power generation and land. The pricing unit has shifted from the GPU to the megawatt of power, and on a16z's podcast David George and Gavin Baker work with round figures of roughly 50 billion dollars of capital per gigawatt, of which about 35 billion is the compute hardware itself. Those are illustrative numbers in a conversation rather than audited ones, but the order of magnitude is what matters to anyone buying tokens: the input to your product is now priced like a utility buildout rather than like software, which is why the capex cycle shows up in your cost of goods.
Is AI capex a bubble?
It can be a bubble in valuations and a real buildout at the same time, and those are separate claims. Gavin Baker's framing is that every genuinely profound technology produces a bubble because markets get ahead of themselves, and that overvaluation then leads to an overbuild. Neither half of that says the underlying demand is fake, which is the distinction that matters to anyone building on top. His empirical test cuts the other way from the bubble narrative: through July and August he asked operators to name one quantitative data point in their business that was getting worse and could not find one. David George said separately that he felt confident the market was not in a bubble at that point, while being less confident about three years out. The indicator that would genuinely change the picture is the share of new capacity funded with debt rather than operating cash flow, because debt demands a much shorter return clock.
Will AI compute stay supply constrained?
The investors in this conversation expect scarcity rather than glut through roughly 2028, which is the opposite of the consensus worry about oversupply. David George has said data center capacity at scale is unavailable until late 2028 or early 2029, and that forecast builds are likely to slip further on permitting and local opposition. The demand side supports it: George puts genuinely heavy paying users below 10 million, revising down Baker's 30 million, against his own figure of roughly 1.5 billion knowledge workers, and the median US company spends only about 12 dollars per employee per month on AI versus about 7,000 dollars for the top 1 percent of the dataset his team has seen. The main thing that would break the constraint is an algorithmic breakthrough that makes capable models much smaller.
What does the AI capex cycle mean for startups?
Three practical things. First, stop treating a falling cost per token as a load-bearing assumption and re-run your gross margin with a flat frontier price for 24 months, because what breaks under that test is usually the absence of a routing layer rather than the forecast. Second, treat capacity as procurement with a lead time rather than a variable cost: vendors want long commitments paid up front, and Y Combinator built a shared GPU cluster precisely because the upfront cost of securing two years of compute exceeded many startups' entire cash balance. Third, the real risk from the cycle is not that the bubble pops, since cheap compute would help you. It is that scarcity persists and frontier capacity gets allocated to whoever can pre-commit at a scale a seed-stage company cannot.
Build your AI Operating System
A practical course to grow with AI, build internal tools, and operate safely. Join the waitlist and you'll be first in when the course opens.