Search Central LiveDeep Dive Europe 2026

Knowledge base v2.13.0 · Community edition · data through 2 October 2026

Topic · Indexing

Main content and page sections

Google extracts every element of a page to tell header, navigation and main content apart; its slides marked the main content very important and the header and navigation not so important, and its canonicalization guide says it determines each page's primary content, or centerpiece, when indexing. Gary Illyes said words are weighted by where they sit: footer text is unlikely to contribute much to ranking, and moving a term into the main content is the simplest way to make it count; Google's Search Essentials supports this only in general terms, advising that search words go in prominent places such as the title and main heading, and describes no weighting. Page structure also drives soft 404 detection, extending Day 1's soft 404 discussion: a BERT-like model trained on page layout ignores navigation and footer, so an error message alone in the main content makes a soft 404 while one in the navigation need not. Pages that translate only the menu and footer around the same main content are clustered as duplicates, as Google's hreflang guide also says. A Google software engineer working on data ingestion said extraction sometimes misjudges the main content or picks up extraneous data such as related products' prices, which structured data helps prevent, and Google said a 'Crawled – currently not indexed' page on a site of even quality may use a template that hides where its content is. A second recording captured Gary Illyes's part on main content in Day 2: main content is any part of a page that directly helps it achieve its purpose, not only text but images, videos or a tool, user-generated content on a UGC site, a comment section, content in tabs, and all headings and the visible title, matching Google's Search Quality Rater Guidelines, which add that tabs and comments can be main or supplementary content depending on the page's purpose. He said the main content is what Google considers when ranking a page, and that whatever a site puts in its navigation or header tells Google it does not particularly care about that content. Both recordings have the token metadata marking a word as in the header, the main content, bold or a heading (a third item, heard as title in one recording and italics in the other, is left out). Author’s view: a logical heading hierarchy is not a Google Search requirement, since Google's SEO Starter Guide says out-of-order headings do not matter to Search; the case for it is accessibility and agents that read the accessibility tree.

What to do

  • Give every template one clearly delimited main content area, and put the terms a page should rank for in its title, headings and opening paragraphs, not in footers, sidebars or tag lists.
  • Never leave a 200 page whose main content is only an error message; return 404 or 410 for missing content and 503 for outages.
  • Translate the whole page for each language version, template and main content alike; translating only the menu and footer makes the versions duplicates.
  • On product pages with related-product carousels, make the Product markup describe only the main item and its price.
  • If one template's pages sit in 'Crawled – currently not indexed' while others are indexed, check whether its layout makes the content hard to locate.

Day 1: Crawling 3

Said on stage 1

StageConsistent with docsD1-C277

A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.

Speaker Carlos OrtegaIn Day 1, 13:10 · Lightning session A: Automation and AIEvidence transcript

Used byrequirement DEV-HTM-04

  • Extended by D1-C314 Day 1: A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do…

What Google's documentation says 1

Analysis by the author 1

AnalysisD1-C314

A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do not matter to Search; the case for it is accessibility, where web.dev advises against skipping levels, and agents that read the accessibility tree.

Author Ibrahim AnjroAnnotates Day 1, 13:10 · Lightning session A: Automation and AI

  • Extends D1-C277 Day 1: A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section…

Day 2: Indexing 33

Shown on screen 5

SlideConsistent with docsD2-C028

Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

Speaker Cherry PrommawinIn Day 2, 10:25 · How is HTML interpretedEvidence 2 slide photos, transcript

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • Repeated by D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
  • Extended by D2-C868 Day 2: Gary Illyes said the main content is what Google considers when ranking a page.
  • Extended by D2-C474 Day 2: Google's systems sometimes fail to determine a page's main content correctly, and structured data helps…
SlideConsistent with docsD2-C309

A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence video, transcript

Used byrequirement DEV-HTM-01

  • Repeats D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
  • Extended by D2-C861 Day 2: Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not…
SlideConsistent with docsD2-C314

On Google's example blog page, the post title and opening sentence counted as important because they sit in the main content, in front of the user, while the site tagline, the 'Categories' sidebar and category links such as 'Hugo (7)' counted as less important supplementary text.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byrequirement DEV-HTM-02

SlideConfirmed by docsD2-C366

Pages whose boilerplate, such as menu and footer, is translated while the main content is not are near matches: the main reason to visit is the same, so Google clusters them as duplicates.

“When main content is the same, pages may be clustered.”

Wording checked against the slide or recording

Speaker John MuellerIn Day 2, 11:55 · Handling web duplicationEvidence slide photo, transcript

Used byrequirement DEV-INT-07

Said on stage 20

StageConsistent with docsD2-C311

Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

  • Extended by D3-C170 Day 3: Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content…
StageConsistent with docsD2-C315

To make a word count for ranking a page, Gary Illyes said the simplest step is to move it into the main content, because where text sits on a page already contributes quite a bit to ranking.

“where you position text on a page will already contribute quite a bit to ranking that page”

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byrequirement DEV-HTM-02

StageNot in docsD2-C323

When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byrequirement DEV-HTM-07

  • Repeated by D2-C720 Day 2: Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the…
StageNot in docsD2-C338

Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.

“This is basically an LLM thing, something like BERT, that is specifically trained to understand page structure”

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

Things
  • Extends D1-C042 Day 1: BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at…
StageNot in docsD2-C339

For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence slide photo, transcript

Used byrequirement DEV-ERR-03

  • Extends D1-C078 Day 1: For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a…
StageConsistent with docsD2-C861

Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
  • Extends D2-C309 Day 2: A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so…
StageConfirmed by docsD2-C862

Gary Illyes defined a page's main content as any part of the page that directly helps the page achieve its purpose, what it was built for.

“Main content is any part of the page that directly helps the page achieve its purpose”

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
StageConfirmed by docsD2-C866

Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

Used byrequirement DEV-REN-02

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
  • Extends D2-C202 Day 2: Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden…
StageConsistent with docsD2-C868

Gary Illyes said the main content is what Google considers when ranking a page.

“It's the main content that we consider for ranking.”

Speaker Gary IllyesIn Day 2, 11:30 · Understanding what's on a pageEvidence transcript

Used byrequirement DEV-HTM-01

  • Extends D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…
StageNot in docsD2-C474

Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.

Speaker Ryan LeveringIn Day 2, 13:35 · What is Structured Data and why we need it on the internet.Evidence transcript

Used byrequirement DEV-SDA-07

  • Extends D2-C028 Day 2: Google extracts every element of a page so that it can tell the header, the navigation and the main content…

What Google's documentation says 4

DocsSourceD2-C310

Google's canonicalization guide says that when Google indexes a page it determines the page's primary content, which it also calls the centerpiece, and clusters pages whose primary content is the same or very similar.

“When Google indexes a page, it determines the primary content (or centerpiece) of each page.”

Publisher Google Search CentralAnnotates Day 2, 11:30 · Understanding what's on a page

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

DocsSourceD2-C870

Google's Search Quality Rater Guidelines define main content as any part of the page that directly helps it achieve its purpose, including text, images, videos, page features such as calculators and content created by users, and they count the title at the top of the page as part of it.

“Main Content is any part of the page that directly helps the page achieve its purpose.”

Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)Annotates Day 2, 11:30 · Understanding what's on a page

Used byrequirement DEV-HTM-01glossary term Main content (centerpiece)

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
DocsSourceD2-C871

Google's Search Quality Rater Guidelines say navigation links are a common type of supplementary content, and that content behind tabs and user reviews or comments may count as main content on some pages and as supplementary content on others, depending on the page's purpose.

Publisher Google Search Quality Rater Guidelines (PDF, 11 September 2025)Annotates Day 2, 11:30 · Understanding what's on a page

Used byrequirement DEV-HTM-01

  • General Guidelines Google Search Quality Rater Guidelines (PDF, 11 September 2025) · checked 3 October 2026
DocsSourceD2-C589

Google's hreflang guide says pages that translate only the template (navigation, footer) around main content in one language still count as duplicates, because localized versions are duplicates only if the main content stays untranslated; it recommends hreflang for such pages.

Publisher Google Search CentralAnnotates Day 2, 14:15 · Focusing on Internationalisation and Localisation

Used byrequirement DEV-INT-07

Analysis by the author 4

Across days and sessions 10

  1. Analysis D1-C314 Day 1 · Lightning session A: Automation and AI

    A logical heading hierarchy is not a Google Search requirement, since Google says out-of-order headings do not matter to Search; the case for it is accessibility, where web.dev advises against skipping levels, and agents that read the accessibility tree.

    extends
    Stage D1-C277 Day 1 · Lightning session A: Automation and AI

    A community speaker advised keeping headings inside the main content in a logical hierarchy (title, section subtitles, subsections) so agents can follow its structure, avoiding several H1 elements and skipped levels such as an H3 followed directly by an H5.

  2. Stage D2-C338 Day 2 · Understanding what's on a page

    Google detects soft 404s with a language model, described as something like BERT, that is trained to understand the structure and layout of a page as well as its language, instead of reading the page as one flat wall of text.

    extends
    Slide D1-C042 Day 1 · How Search works and where's AI?

    BERT is used in indexing to understand each word in the context of the whole sentence rather than one word at a time.

  3. Stage D2-C339 Day 2 · Understanding what's on a page

    For soft 404 detection, the position of error text decides: an error in a less important part such as the navigation need not make a page a soft 404, but an error message alone in the main content, like 'Error establishing a database connection', does, much as a human would judge it.

    extends
    Analysis D1-C078 Day 1 · How crawling errors affect Search

    For temporary blocks return 503 or 429, never a 200 page with a captcha or error message, which becomes a soft 404 and can replace real content in the index. Verify Googlebot by reverse DNS, not by user agent.

  4. Stage D2-C474 Day 2 · What is Structured Data and why we need it on the internet.

    Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.

    extends
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  5. Stage D2-C861 Day 2 · Understanding what's on a page

    Gary Illyes said that whatever a site puts in its navigation or header tells Google the site does not particularly care about that content: it may help users do something on the side, but it is not what the page wants them to do, read or take away.

    extends
    Slide D2-C309 Day 2 · Understanding what's on a page

    A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

  6. Stage D2-C866 Day 2 · Understanding what's on a page

    Content inside tabs, for example separate tabs for a product description and a manufacturer description, might be part of a page's main content.

    extends
    Slide D2-C202 Day 2 · Lightning session D: Rendering and JavaScript

    Tab and accordion content should be in the DOM from the start and only hidden with CSS or the hidden attribute; Google indexes such hidden content.

  7. Stage D2-C868 Day 2 · Understanding what's on a page

    Gary Illyes said the main content is what Google considers when ranking a page.

    extends
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  8. Stage D3-C170 Day 3 · How Google thinks about Quality

    Google's quality talk pointed to page 21 of the Search Quality Rater Guidelines for its definition of content quality by effort, originality, talent or skill and accuracy, noting that the document is updated from time to time.

    extends
    Stage D2-C311 Day 2 · Understanding what's on a page

    Gary Illyes pointed to Google's Search Quality Rater Guidelines as the detailed source on how Google thinks about the main content of a page.

  9. Slide D2-C309 Day 2 · Understanding what's on a page

    A Google slide labelled the parts of an example blog page, marking the header and the navigation as not so important and the main content as very important.

    repeats
    Slide D2-C028 Day 2 · How is HTML interpreted

    Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.

  10. Stage D2-C720 Day 2 · How does the index look like?

    Google's Search index stores the tokens produced by tokenizing each page, as they are, together with the metadata attached to the tokens during tokenization.

    repeats
    Stage D2-C323 Day 2 · Understanding what's on a page

    When tokenizing for Search, Google attaches metadata to each token for use in ranking, such as whether the word appeared in the header or in the main content (tagged 'centerpiece' on the slide), in bold or in a heading; Gary Illyes said he was not showing all of the metadata.

Built on these claims 8

Developer requirements 8

Sources 8