Plumbing
"Write self-contained sections" is half the advice. Perplexity published the other half.
By Arnav Mukherjee, founder of TofuBofu · October 1, 2026
TL;DR
- Stop writing isolated islands. On the primary answer route, 1,067 of 2,099 benchmark queries need two or more supporting passages.
- Name the subject inside the passage. Perplexity says a chunk inheriting meaning from a heading may stop matching.
- Put the proof beside the claim, not in a later section. 735 queries need exactly two evidence groups.
- Stop writing one 4,000-word page that covers everything. One pooled vector "cannot faithfully capture every facet".
- Proofread against the paper's twelve named dependencies. Three of them carry your price, promise and proof.
We've published the standard advice ourselves. In how an AI engine indexes your content we told you to make passages self-contained and avoid references that only work with the paragraph above. It's the most repeated instruction in this category, and I've yet to see it carry a source. Ours didn't either.
On 30 September Perplexity Research and turbopuffer published the mechanism underneath it, with a model and a benchmark. It largely confirms the advice and it corrects one part, and the corrected part is the one that changes what you'd do on Monday. The target isn't an island. It's a passage that stands up alone, with its proof within reach.
What this source is, and what it is not
- What was published. A contextual embedding model,
pplx-embed-v2-context-9b-preview, and context-bench, a new retrieval benchmark built and privately held by turbopuffer. Preview weights are on Hugging Face. - The benchmark runs 2,099 queries over 38,894 documents, cut into 2,458,072 sentence chunks, across 21 domains including corporate filings, clinical research, legal contracts and software documentation. The target documents are long on purpose: median around 6,100 tokens, and 1,061 of the 1,197 distinct primary targets run to at least 200 sentences.
- Why it is citable here. It is a search company publishing how retrieval of this kind works, with the method and the benchmark design in the open. That is the sourcing bar we hold ourselves to: read what the engine vendors publish, not what the industry says about them.
- What it is not: it is not documentation of how perplexity.ai picks which site to cite in a live answer, and it says nothing about ChatGPT, Gemini, Google AI Mode or Copilot. Nobody should read a ranking factor out of it, us included.
- One more limit worth naming. Perplexity's model is evaluated on a benchmark held by its co-author, submitted blind, and the paper says so. We're quoting the mechanism and the benchmark design, not the leaderboard.
What does the engine actually index?
Not your page. A retrieval system splits a long document into smaller pieces that can be indexed and searched on their own, then matches the question against each piece. The piece is what competes for a citation, which means your page is judged in parts and the parts aren't equally strong.
Perplexity names the cost of doing it that way: "Chunking is convenient in practice, but it removes the context in which each chunk originally appeared, creating a trade-off between the granularity of information compression and the context used during encoding."
Then the sentence that should be pinned above every B2B content calendar in the country:
"A chunk may refer to an entity introduced earlier, inherit its meaning from a section heading, or rely on definitions located elsewhere in the document. Once extracted, it may no longer contain enough information to be matched reliably to a query."
Read that as a description of your pricing page. The number sits in a table whose currency is in the header. The caveat lives two sections down. The paragraph says "the platform" because the H2 above it said your product name. Every one of those is a passage that reads perfectly to a human and arrives at the index missing the thing that would have made it findable.
Which part of the standard advice was wrong?
The part that treats a passage as a self-sufficient unit. Retrieval research has mostly assumed one relevant chunk per question, the "gold passage". Perplexity's whole argument is that one passage usually isn't enough: "It may contain the answer but lack the supporting context needed to understand or verify it, leaving it ambiguous in isolation."
Their benchmark quantifies how often that happens, and this is the number I'd act on. An answer route is one valid combination of an answer passage plus the evidence needed to interpret or verify it. Across the 2,099 queries:
| Evidence groups needed | Queries |
|---|---|
| One | 964 |
| Two | 735 |
| Three or more | 332 |
| None | 68 |
As a sentence, because a table can't be quoted: on context-bench, along the primary target-document answer route, 1,067 of 2,099 queries need two or more groups of supporting evidence alongside the answer passage, and 964 are satisfied by a single one. The paper adds that alternative valid answer routes can require different evidence groups. So a page built as a stack of unrelated islands isn't optimised, it's stripped. The answer has nothing standing next to it that lets a reader or an agent check it.
And before anyone reaches for the opposite fix: one enormous page doesn't solve it either. "Encoding the entire document as a single vector does not resolve this, since one pooled representation cannot faithfully capture every facet of a long document." That is the mechanism by which a long pillar page covering a whole category can lose a citation to a short page covering one question.
Which of your pages is actually getting read? A free TofuBofu scan puts your real buying questions to all five engines, shows you the verbatim answers, and lists the sources each engine pulled instead of you.
Run a free scanTwelve ways a passage breaks, and what each one looks like on a B2B site
The benchmark tests twelve contextual capabilities, each one a dependency that a passage loses when it's extracted. The paper gives a worked example of each. Here are all twelve, with what each failure typically looks like on a B2B services or software site. The dependency names and definitions are theirs. The B2B translation is ours.
- Entity identity, where the passage says "the company" and the introduction said who. On your site: every case study that calls the client "the customer" and every feature page that says "the platform".
- Period or version, where instructions apply to one release and the release is named elsewhere. On your site: a pricing page updated in place, with no date in the sentence carrying the price.
- Definitions and aliases, where a figure is reported against a label defined earlier. On your site: internal tier names, "Professional" and "Growth", used in a passage that never says what they include.
- Table structure, where a number takes its meaning from a column header and a unit. On your site: the comparison table, where the row is lifted without the header that gave its numbers meaning.
- Conditions and exceptions, where a section introduction restricted the coverage below it. On your site: the SLA whose exclusions sit under a separate heading, so the retrievable passage promises more than you do.
- References and pronouns, as in "she took office in 2021" with the name in the paragraph above. On your site: "they reduced ticket volume by 40%", where "they" was named two paragraphs back.
- List order and position, where an item means what it means because of where it sits. On your site: numbered implementation steps that don't restate what they're steps of.
- Section scope, where a passage belonging to one section reads as a general statement. On your site: a limitation written under "Small teams" that an engine can quote as though it applied to everyone.
- Speaker attribution, where a quotation draws its authority from whoever said it. On your site: a testimonial separated from the name and job title that made it worth printing.
- Changes over time, where a current value replaced an earlier one on the same page. On your site: a changelog or a policy page where the superseded figure is as retrievable as the live one.
- Causal relationships, as in "that fault caused the shutdown" with the fault identified elsewhere. On your site: a results claim whose cause sits in a different paragraph, so the number arrives with no reason attached.
- Figurative or cross-language meaning, where a glossary fixes the sense of a term. On your site: category jargon you coined, and any page serving more than one language or market.
Run that down your three most commercially important pages. Start with table structure, conditions and exceptions, and references and pronouns, because our read is that those three cost the most on a B2B site: they are the passages carrying your price, your promise and your proof.
What should you change this week?
All of this is editing, not engineering. That's the useful part: it's the cheapest AEO work available, and it needs nobody's developer.
- Name the subject in the passage. Don't let a heading carry it. If a paragraph says "the platform", write your product's name instead, even where it reads as repetition to a human.
- Put the qualifier inside the sentence holding the number. Currency, unit, billing period, region, date. A price without its period is a passage that can be quoted against you.
- Move the proof next to the claim. This is the correction to the old advice. 1,067 of 2,099 queries needed two or more evidence groups, so the claim and the thing that verifies it should be neighbours.
- Restate the condition rather than cross-referencing it. "Excludes hardware" belongs in the passage about coverage, as well as under Exclusions.
- Split the pillar page by question. One page per question a buyer actually asks beats one page covering the category, and the pooled-vector sentence above is the reason why.
- Then check whether it worked. Not by reading the page, by asking the engines. A passage you believe is self-contained and an engine's actual answer are two different pieces of evidence.
SEO is still the floor under all of it. A page that can't be crawled isn't eligible to be chunked in the first place. But none of the twelve dependencies above is a ranking problem, and you can fix every one of them on a page that already ranks and still isn't getting quoted. That's the layer AEO is, and it's the layer you can win while your Google position doesn't move at all.
Frequently asked questions
Why doesn't AI retrieve my best page?
Often because the passage that holds your answer stops making sense once it is cut out of the page. Retrieval systems split a document into chunks and match the question against each chunk separately, so a passage that leans on a heading, a definition, a pronoun or a table header elsewhere on the page can lose the thing that made it findable. Perplexity Research states the mechanism directly: 'A chunk may refer to an entity introduced earlier, inherit its meaning from a section heading, or rely on definitions located elsewhere in the document. Once extracted, it may no longer contain enough information to be matched reliably to a query.' That is a property of your writing, not of your rankings.
Is 'make every section self-contained' still the right advice?
It is the right instinct and it is incomplete, and the correction is useful. Perplexity and turbopuffer's benchmark reports that, along the primary target-document answer route, 1,067 of its 2,099 queries need two or more groups of supporting evidence rather than one passage: 964 queries need one group, 735 need two, 332 need three or more and 68 need none. The paper adds that alternative valid answer routes can require different evidence groups. So the target is not an island. It is a passage that names its own subject, plus nearby passages that confirm it. Write the answer so it stands up alone, and keep the evidence that verifies it close by instead of three screens away.
What is a chunk, and why does it decide whether I get cited?
A chunk is the unit a retrieval system actually indexes: a sentence, a paragraph or a fixed window of your page rather than the page itself. The engine converts each chunk and the question into numeric representations and matches them, then builds its answer from the closest chunks. So the thing competing for a citation is a passage, not a domain. Perplexity's own framing of the cost is that 'chunking is convenient in practice, but it removes the context in which each chunk originally appeared ...' Your page is judged in pieces, and the pieces are not all equally strong.
Can't the engine just read my whole page?
Compressing a long document into one representation loses the detail that answers a specific question. Perplexity states it plainly: 'Encoding the entire document as a single vector does not resolve this, since one pooled representation cannot faithfully capture every facet of a long document.' This is why a 4,000-word pillar page that covers everything can lose to a short page that covers one thing. The long page is not penalised for length. It is diluted, because no single summary of it is a close match for any one question.
Does this paper tell me how Perplexity ranks my website?
No, and anyone telling you it does is overselling it. This is research on an embedding model and a retrieval benchmark, published by a company that builds search. It is not documentation of how perplexity.ai chooses which site to cite in a live answer, and it says nothing at all about ChatGPT, Gemini, Google AI Mode or Copilot. What it gives you is the mechanism of contextual chunk retrieval, described by people who build one, with a benchmark attached. Treat it as the best available account of why passages fail, and not as a ranking factor list.
What are the twelve ways a passage loses its meaning?
The benchmark tests twelve contextual capabilities, each one a dependency that breaks when a passage is extracted: entity identity, period or version, definitions and aliases, table structure, conditions and exceptions, references and pronouns, list order and position, section scope, speaker attribution, changes over time, causal relationships, and figurative or cross-language meaning. Read them as a proofreading list. Three are worth checking first on a B2B site, and this is our read rather than the paper's: table structure, the pricing table whose currency lives in the header; conditions and exceptions, the feature whose caveat sits in a different section; and references and pronouns, the paragraph that says 'the platform' and never says which. Those three carry your price, your promise and your proof.
What should I change on my pages this week?
Take the three pages you most want cited and read each section as if it were the only thing anyone would ever see of your site. Name the subject in the passage rather than leaning on the heading above it. Repeat the qualifier instead of referring to it. Put the unit, the currency, the region and the date inside the sentence carrying the number. Move the proof that verifies a claim next to the claim rather than into a separate section. None of that is a technical change and none of it needs a developer, which is why it is the cheapest AEO work available to a small team.
Sources and further reading
- Perplexity Research and turbopuffer, "Contextual embedding beyond the gold passage", 30 September 2026. Source of every quotation and every figure on this page, including the twelve capabilities and the evidence-group counts. Read at source 1 October 2026.
- Lewis et al., "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks", arXiv:2005.11401. The original description of retrieving passages to condition a generated answer, which is the architecture all of this sits inside.
- Google Search Central, "Optimizing your website for generative AI features on Google Search". Worth reading alongside, because it is the other vendor that has published guidance in its own words, and it is more conservative than the industry is about what you need to change.