Skip to content
CampeloLabs
← Blog

LLM SEO: Write for the Excerpt, Not the Click

Cicero Campelo

Cicero Campelo, CISSP
October 4, 2026 · 16 min read

Part of our guide to AI for startups.

A founder watching a single paragraph lift off a published web page and travel into an AI agent's context window
Table of contents

Every page you publish now has two readers. One is a person. The other is a model that will read the page, lift one paragraph out of it, and hand that paragraph to someone who never loads your site. LLM SEO is the work of making sure that paragraph is yours.

That framing is not a marketing line. It is why Parag Agrawal, the former Twitter CEO now building Parallel Web Systems, picked the company's name. Talking to Sequoia, he described visualizing a parallel web built for AIs, and then the consequence: "when you publish, you're now thinking of, as we all are now, of two audiences." His own team already treats the second one as primary. "It's their agents reading our docs and code in our SDKs," he said of Parallel's customers. "It's not humans fumbling around docs pages for the most part."

Most advice filed under LLM SEO takes that observation and bolts it onto the old playbook: same tactics, new acronym. The interview is more useful than that, because Agrawal is building the retrieval layer itself and will say out loud which parts of the old machinery he considers broken. Two of them are load-bearing for everything SEO has done for twenty years.

What LLM SEO actually is

LLM SEO is the practice of making your content retrievable, quotable and correctly attributed when a language model answers a question, rather than when a person browses a results page. It is sometimes called answer engine optimization or generative engine optimization. The names differ; the job is the same.

The useful way to think about it is not as a new channel bolted beside Google. It is the same web, read by a different kind of reader with a different budget. A person skims, scrolls, tolerates a slow page, and clicks a second result when the first disappoints. An agent does none of that. It pulls text, scores it, and either uses your paragraph or uses someone else's.

That difference is small in description and large in practice, because the two things SEO optimizes hardest, the page and the click, are both the wrong unit now.

The unit of retrieval is the excerpt, not the page

The unit an answer engine retrieves is a passage of roughly 1,000 tokens, not your page. Ask Agrawal what a query looks like inside his system and the answer is startlingly concrete: "every query is essentially give me 1,000 tokens from a trillion web pages on the web, and make sure that they're the right 1,000 tokens."

A thousand tokens is roughly 750 words. Not a page. Not a ranked list of ten pages. One excerpt, or a handful of them, assembled into a model's context window. Agrawal describes the pipeline as boiling tens to hundreds of billions of URLs or documents down to thousands or tens of thousands, then to specific excerpts and paragraphs, with bigger models running at each stage.

He is explicit about what the system is reaching for: an excerpt from the most authoritative place on the web, "trying to bring it to the agent's context window."

Here is why that reframes the work. A page that ranks well is a page that wins a comparison between pages. An excerpt that gets retrieved is a passage that answers one question completely, on its own, without the rest of the page around it. Those are not the same artifact. A well-optimized 3,000-word guide with the answer distributed across an intro, a table and a conclusion is a good page and a bad excerpt.

Agrawal also points out that the old economy produced a specific kind of page as a rational response to human laziness. Asked how agents handle the affiliate-slop pages that rank for "best product for X," he walked through the example of a public company's headline revenue number: the authoritative figure sits deep inside a slow-loading SEC filing, so a faster page that puts it above the fold wins the human. "It is worth putting it on a page that loads fast, where this information is above the fold." His verdict on that whole genre is pointed: "you can call it slop, pre-AI human slop, or you can call it catering to a lazy human and being successful at SEO."

In Agrawal's design the agent does not need that intermediary: it takes the excerpt from the authoritative source directly. Which means the strategy of being a faster, cleaner middleman between a reader and a primary document is the single most exposed position on the web right now.

Clicks stopped being the feedback signal

Clicks stopped being the feedback signal because the customer is now an agent, and Agrawal states it as company doctrine: "human click data is a bug."

Sit with that. He names click data specifically, but the whole behavioral layer search engines built on top of it, click-through rate, dwell time, pogo-sticking, bounce, rests on the same assumption: that what a human did next tells you whether a result was good. Parallel's position is that an agent doing work with search should rely on agent feedback instead, and that large models have made it cheap enough to have experts create ratings data directly, so nobody has to mine click logs for a quality signal.

If retrieval quality gets judged by whether the excerpt actually helped the agent complete its task, then the levers move. Headline curiosity gaps, clickbait framing, interstitials, anything that optimizes the decision to click, is optimizing a signal that is being deliberately discarded. What survives is narrower and harder: is the passage accurate, is it specific, is it self-contained, and does it hold up when a model checks it against three other sources in the same context window.

That last point deserves emphasis because it inverts a familiar incentive. Vague, hedged, broadly applicable copy is good SEO insurance and terrible retrieval material. Ten pages saying roughly the same safe thing are interchangeable, and a ranking system that needs the right 1,000 tokens will take whichever one is most concrete. Specificity is the moat.

The query coming at you changes too. "Humans rely on keyword search," Agrawal said, before describing how different the agent side looks: "Fewer typos, better specified queries, perhaps longer queries, less for the search engine to guess what the agent might want." Writing to a three-word keyword stub was always a compromise with how humans type. The compromise is expiring.

The traffic math is already visible

More than half of internet traffic is no longer human, and human visits to software and IT services pages have fallen by as much as 40% in under a year. None of this is a forecast you have to take on faith, and you should not take it from a search startup alone. Cloudflare sits underneath a large share of the web and is reporting what crosses its own network. It is not a disinterested party either, since the same post launches its crawler-control and pay-per-use products, but its traffic numbers are first-party measurements rather than forecasts. Its September 2026 report is titled, almost word for word, "The Internet has a second audience."

The numbers in it are the ones worth putting in front of a board:

  • "This year, for the first time, more than half of Internet traffic wasn't human."
  • Daily requests from AI agents on Cloudflare's network grew by more than 1,700% over the preceding year.
  • The share of crawler requests declared as being for AI training went from 22% in spring 2025 to 52% by June 2026.
  • In some of the most heavily crawled categories, including retail, computer software, IT services and financial services, human traffic fell by as much as 40% in under a year.

Cloudflare's framing of the break is the clearest one-sentence version of the problem: for thirty years, being found and getting paid were the same thing. Answer engines read the page and summarize it, which separates the two.

A human-traffic decline of as much as 40% in categories that include computer software and IT services is not a distant risk for a founder publishing docs and content. That is the category you are in.

Skip llms.txt, at least for now

Publishing an llms.txt file is the most common recommendation in LLM SEO writing and the weakest-evidenced one, so it should not be first on your list. It is a markdown index of your site aimed at models, proposed at llmstxt.org, and it takes an afternoon, which is exactly why it ends up at the top of every checklist.

The evidence does not support putting it first. Ahrefs analyzed 137,000 sites and published the result in the title: 97% of llms.txt files never get read. What it measured is that 97% of them got no requests at all in May 2026. PPC Land pairs that with Originality.ai's separate tracking study, which counted llms.txt adoption growing 8.8 times in twelve months, from 4,088 files in June 2025 to 36,120 by May 2026. Google has not committed to using it for Search, and neither OpenAI nor Anthropic documents reading third-party llms.txt files, though both publish their own for their developer docs. The largest single share of the requests those files do get, 21.7%, comes from SEO audit tools, ahead of every AI crawler in the study.

That is not an argument that the idea is wrong. It is an argument about sequencing. A file that almost nothing fetches cannot be the first thing you do, and treating it as the deliverable is how a team convinces itself it has addressed LLM SEO without changing a single page a model actually reads.

The more concrete version of the same instinct is to serve machine-readable versions of the pages agents actually request. Resend, a Y Combinator company, does both: it publishes an llms.txt, serves its pricing page as real markdown at resend.com/pricing.md, and keeps an AI onboarding page in its docs. Its founder Zeno Rocha said on LinkedIn in May 2025 that ChatGPT had become a top-three source of traffic for the site. He did not say which surface did the work, and nobody outside Resend can. In fairness, we publish an llms.txt on this site too. It is not what moved anything. We covered how that reshapes dev-tool go-to-market in why developer relations is a product decision.

What actually moves the needle

Seven changes move LLM SEO outcomes, and the first, making every section answer its own heading, is worth more than the other six combined. Each follows from how retrieval actually works rather than from folklore.

  1. Make every section a complete answer. If a model lifts one H2 and its first two paragraphs, that fragment should answer the question in the heading without the rest of the page. State the answer in the first sentence, then support it. This is the single highest-leverage change, and it is a writing change, not a technical one.
  2. Put the number, the name and the date in the sentence. Not in a chart, not in a table two sections away, not implied. An excerpt carries no page around it, so a claim that depends on context the reader can see is a claim the model cannot use.
  3. Be the authoritative source for something, or quote one explicitly. Agrawal's system is reaching for the most authoritative place on the web. Original data, your own benchmarks, your pricing, your changelog, the results you can observe and nobody else can, these are retrievable. Rewrites of other people's posts are what the excerpt pipeline is designed to route around.
  4. Serve a clean machine-readable version of your high-intent pages. Pricing, docs, API reference, changelog. Markdown at a predictable URL, correct content type, no JavaScript required to see the text. Start with the pages a buying question would land on.
  5. Let the agents in, deliberately. Robots rules written for 2019 scrapers now block the retrieval systems that would cite you. Decide per crawler, separating training crawlers from inference-time retrieval, rather than issuing a blanket no. The harder half of that problem, telling a customer's agent apart from an attacker's, is its own discipline: see bot detection when bots are your customers.
  6. Write to the full question. Agrawal describes agent queries as better specified, with fewer typos, and perhaps longer. Headings phrased as the complete question a founder would ask outperform headings phrased as keyword stubs, and they survive being read with no page around them.
  7. Measure citation, not rank. Position in a results page is increasingly the wrong instrument. Ask the major assistants your category's real questions on a schedule, record who gets named, and track whether that changes. Referral data will undercount agent-driven demand badly, because the model reads and the human arrives later by typing your name.

Some of this overlaps with designing a product agents can use, which is a related but distinct job covered in agent experience when your user is an agent. The difference is the surface: that one is your API and your auth, this one is your published text.

Who gets paid when the agent answers

The uncomfortable part of LLM SEO is that the business model underneath it has not been rebuilt yet. Agrawal is blunt that the existing option, a fixed-fee licensing deal with a model lab, is available only to the head of the web and is broken even there: inference volume grows several times a year while the contracts do not, so every publisher expects its share to fall at renewal.

Parallel's proposed answer is to price each source's contribution using Shapley values, a game-theoretic method for splitting the gains of a collaboration, estimated rather than computed exactly because the full computation would cost more than the agent run itself. His timeline: "we're 12 to 24 months from this math being able to give meaningful dollars for a very wide range of content owners on the web."

Treat that as a vendor's roadmap, not a fact. It is a prediction by the person selling the mechanism, and it is worth knowing mostly because it tells you what to prepare for rather than what to count on. The planning question for a founder is narrower and answerable today: if agent reads never pay, is your content still worth publishing? For most startups it is, because the content was never an ad-supported business. It was a way to be the named default in your category, and in an answer engine, being named is the whole prize. That dynamic, and how a challenger breaks an established default, is the subject of how defaults became the market in the AI agent economy.

The strongest case this is overblown

The strongest counter-argument is that AI answering questions has been positive sum for search rather than zero sum, and it comes from someone with the data to know.

Logan Kilpatrick, who works on Google AI Studio and the Gemini API at Google DeepMind, argues the zero-sum reading has not held up so far. Everyone assumed AI answering questions would be negative sum for search, he told Sequoia, "and actually what ended up happening is it's been incredibly positive sum for search." His evidence is behavioral: "people are searching more, people are doing more," with agent queries arriving as a new market on top rather than as a substitute for human ones. He bounds it himself, calling the next one to two years somewhat clear and three to five years much less so.

He is not disinterested either. Google has the strongest commercial reason of anyone to say the pie grew. But his point stands on its own, and Cloudflare's data is consistent with it: total requests on its network nearly doubled. The human web did not shrink so much as get diluted.

The honest synthesis is that the volume is going up and the share reaching you as a visit is going down. Both parties are describing the same curve from different ends. For a founder that resolves cleanly enough: publish for retrieval, stop valuing the work by sessions, and do not let either an optimistic or an apocalyptic reading of the traffic charts decide your content budget.

One disclosure, since this piece leans on a Sequoia interview: Sequoia led Parallel's $100 million Series B in April 2026, which valued the company at $2 billion, and one of the two interviewers, Sequoia partner Andrew Reed, joined Parallel's board in the same round.

What to do this week

  1. Pick the five pages that would answer a buying question in your category. For each, rewrite the opening sentence of every section so it answers that section's heading on its own.
  2. Ask ChatGPT, Claude and Gemini the ten questions a prospect actually asks before buying. Write down who gets cited. That list is your baseline, and nothing else you do this quarter is measurable without it.
  3. Publish a markdown version of your pricing page at a predictable URL with the correct content type. Confirm it returns text/markdown and not HTML.
  4. Audit your robots.txt and WAF rules for crawlers you are blocking by accident. Separate training crawlers from inference-time retrieval and decide each on purpose.
  5. Find the one number, benchmark or dataset only you can publish, and publish it as a standalone sentence with the figure, the method and the date in it.
  6. Skip llms.txt until the five items above are done.

Getting this right is part of a larger shift in how a startup operates when models sit in the middle of every workflow, which is the subject of our pillar on running an AI-first startup. If you want the operating system around it, the course is AI Operating System for Startups.

Sources

Frequently asked questions

What is LLM SEO?

LLM SEO is the practice of making content retrievable, quotable and correctly attributed when a language model answers a question, rather than when a person browses a results page. It is also called answer engine optimization or generative engine optimization. The mechanical difference from traditional SEO is the unit being retrieved. Parag Agrawal, founder of the agentic search company Parallel Web Systems, describes a query to his system as asking for roughly 1,000 tokens out of a trillion web pages, narrowed down to specific excerpts and paragraphs rather than to a ranked list of pages. So the thing being optimized is a passage that answers one question completely on its own, with no page around it, rather than a document that wins a comparison against other documents.

Is LLM SEO different from traditional SEO?

LLM SEO shares the fundamentals with traditional SEO, being crawlable, accurate and genuinely useful, but two load-bearing parts of the old playbook do not carry over. First, the retrieval unit is the excerpt rather than the page, so an answer distributed across an introduction, a table and a conclusion is a good page and a bad excerpt. Second, the behavioral signals are different. Parallel's stated position is that "human click data is a bug" and that agent search should rely on agent feedback instead, which means tactics aimed at winning the click, including curiosity-gap headlines and clickbait framing, optimize a signal that is deliberately being discarded. Agrawal also describes agent queries as better specified and carrying fewer typos, and perhaps longer, so writing to short keyword stubs matters less than answering the full question.

Does llms.txt help with LLM SEO?

The published evidence says llms.txt does not help much yet, so it should not be the first thing a team does. Ahrefs analyzed 137,000 sites and reported that 97 percent of llms.txt files received no requests at all in the month it measured, May 2026, and that the largest single share of the requests that did arrive, 21.7 percent, came from SEO audit tools rather than from any AI crawler. Google has not committed to using the file for Search, and neither OpenAI nor Anthropic documents reading third-party llms.txt files, though both publish their own for their developer documentation. A lower-risk version of the idea is serving machine-readable versions of the pages agents actually request: Resend serves its pricing page as real markdown at resend.com/pricing.md and keeps an AI onboarding page in its docs, alongside an llms.txt of its own.

How do you measure LLM SEO?

Measure citation rather than rank, because referral analytics will undercount agent-driven demand badly: the model reads the page and the human often arrives later by typing your brand name directly, with no referrer attached. The practical baseline is to ask the major assistants the ten questions a prospect actually asks before buying, on a fixed schedule, and record which sources get named each time. Track the change in that list rather than position in a results page. Cloudflare's September 2026 data shows why the old instruments mislead: more than half of internet traffic is no longer human, and in heavily crawled categories such as software and IT services, human traffic has fallen by as much as 40 percent in under a year while total requests rose.

Build your AI Operating System

A practical course to grow with AI, build internal tools, and operate safely. Join the waitlist and you'll be first in when the course opens.