Google gave four reasons why structured data is still valuable even though models can extract information from pages: precision, extra content, efficiency and focus.
Speaker Ryan LeveringEvidence slide photo, transcript
Knowledge base v2.13.0 · Community edition · data through 2 October 2026
Day 2 · Thursday 1 October 2026 · 13:35
Named in the previous speaker's hand-over ('Ryan is going to talk about how we are actually dealing with structured data'), heard clearly in a second attendee recording; he said he has been a software engineer at Google for about 15 years, focused on data ingestion.
Google gave four reasons why structured data is still valuable even though models can extract information from pages: precision, extra content, efficiency and focus.
Speaker Ryan LeveringEvidence slide photo, transcript
Structured data gives the high precision that complex schemas such as sale pricing need, with higher accuracy than large-scale extraction by large language models (LLMs).
“Structured data provides the high precision needed for complex schema (sale pricing), achieving higher accuracy than large-scale LLM extraction.”
Wording checked against the slide or recording
Speaker Ryan LeveringEvidence slide photo, transcript
Structured data often carries non-visible metadata that the page text lacks, such as full ISO dates or stable identifiers for user-generated content.
“It often contains non-visible metadata, such as full ISO dates or stable identifiers for UGC, that is not present in the page text.”
Wording checked against the slide or recording
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirements DEV-SDA-04, DEV-SDA-05
Parsing structured data is significantly cheaper and more efficient than relying on LLMs for every complex extraction task.
Speaker Ryan LeveringEvidence slide photo, transcript
Structured data explicitly points to the pertinent data on a page, which reduces noise and stops Google's systems from pulling in extraneous information such as prices of related products.
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirement DEV-SDA-07
A Google case study shown on stage (published May 2018) reported that Rakuten Recipes got 2.7 times more traffic from search engines and a 1.5 times increase in session duration after implementing recipe structured data.
Speaker Ryan LeveringEvidence slide photo, transcript
Google recommends using the Search gallery in its developer documentation to find the structured data features that suit a site; the gallery shows each feature and how Google uses the markup.
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirement DEV-SDA-03
Google recommends testing and previewing structured data in the Rich Results Test, which shows the rich result features it detected and whether the markup is valid.
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirement DEV-SDA-11glossary term Rich Results Test
Google announced server-side structured data validation as coming soon: it will publish downloadable validation rules in SHACL on each structured data feature guide, with other kinds of checks to follow.
Speaker Ryan LeveringEvidence slide photo, transcript
Used byrequirement DEV-SDA-12glossary term SHACL
Google's planned validation workflow has five steps: download the rules from the feature guide, generate the JSON or embedded microdata or RDFa, run the rules against the generated server-side markup as a first check, deploy and test in the Rich Results Test, and monitor ongoing performance in Search Console.
Speaker Ryan LeveringEvidence slide photo
Used byrequirement DEV-SDA-12
Google's example SHACL shape for Event requires a name, treats a missing description only as a warning, and accepts the image either as an ImageObject or as a URL.
Speaker Ryan LeveringEvidence slide photo, transcript
Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
Speaker Ryan LeveringEvidence transcript
AI Overviews and AI Mode launched as fairly text-heavy answers with little image content and few tables, and they have become more structured over time because that is what users want.
Speaker Ryan LeveringEvidence transcript
Structured data turns the loosely structured web into structured information that powers visual search features such as review stars and recipe filters (for example by preparation time).
Speaker Ryan LeveringEvidence transcript
Used byglossary term Rich results
Describing a page's content in a structured way lets Google understand and interpret that content more accurately.
Speaker Ryan LeveringEvidence transcript
Used byglossary term Structured data
Structured data can bring more qualified traffic, because the site can be shown in new, more interesting ways to more people, who are then more likely to click.
Speaker Ryan LeveringEvidence transcript
Inside Google, views on structured data split into two camps: one says it is useless because models can generate it or read the page directly, the other says it is the future of machines talking to each other through MCP servers and new standards; the speaker said the truth is in the middle.
Speaker Ryan LeveringEvidence transcript
In the speaker's own tests, even the latest LLMs asked to generate schema.org markup for a page often invent properties that do not exist, get deeply nested schemas such as complex pricing models wrong, and duplicate content across several fields.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-11
Automatic LLM extraction of structured information makes a great demo but is not yet good enough for extraction that aims at something like 99.9% accuracy, the speaker said.
Speaker Ryan LeveringEvidence transcript
Google's own extraction systems focus heavily on what is visible on a page, which improves precision but means they can miss content or interpret it incorrectly.
Speaker Ryan LeveringEvidence transcript
Gemini in Chrome relies heavily on the screenshot it takes of a page.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-HTM-04
Markup can state a date as a full ISO 8601 value, which disambiguates a visible date whose local time zone Google might not detect correctly; this matters for event extraction and wherever the date must be exactly right.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-05
Even Google cannot afford to run complex AI models on every page in its index, and distilling them into cheaper models makes them less precise.
Speaker Ryan LeveringEvidence transcript
Rule-based parsing of markup is nearly free by comparison with AI models, so Google will always prefer extracting information from structured data over model-based extraction.
“So we're always going to prefer that particular approach.”
Speaker Ryan LeveringEvidence transcript
Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-07
The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning and quality work, the same data goes to both the classic results page and the AI features.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-AIF-02
As far as the speaker knows, in most of Google's main AI uses a page's schema.org markup is not turned into text and put directly into the model's context; the data is first sorted out, checked for quality and indexed before it is passed on as grounding context.
Speaker Ryan LeveringEvidence transcript
The speaker said there are concerns that feeding raw schema.org markup straight into an AI model's context would be a big abuse vector, one reason Google does not do so.
Speaker Ryan LeveringEvidence transcript
Google's developer site has a number of case studies that show traffic and session-duration gains from structured data.
Speaker Ryan LeveringEvidence transcript
The speaker called the Rakuten study old and said the world has changed a lot, but expects the link between structured data and more interactive, visually appealing results to hold for the foreseeable future.
Speaker Ryan LeveringEvidence transcript
Schema.org is a common vocabulary that several major search engines started together so that site owners can mark up pages and every consumer interprets the markup the same way; the speaker put its start 15 to 20 years ago (Google, Bing and Yahoo! announced it in June 2011).
Speaker Ryan LeveringEvidence transcript
Used byglossary term Schema.org
Schema.org is an open public collaboration, and the speaker invited people who enjoy ontologies and data to join it.
Speaker Ryan LeveringEvidence transcript
Marking something up with only a very generic schema.org type says little about it and is very hard for Google to use.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Google uses schema.org as its markup vocabulary but mostly consumes only the subsets that its developer documentation declares, for the features it builds.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03glossary term Schema.org
Google will never penalise a site just for having more structured data on its pages than Google uses; extra markup does not hurt.
“we will never penalize you for having more structured data on your pages”
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Describing every semantic detail of a page in markup is probably not worth the effort; focus on the structured data that Google or other consumers actually use.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Google's structured data types fall roughly into two groups: broad page-level types such as breadcrumbs and article, and vertical-specific types such as recipes, events and products.
Speaker Ryan LeveringEvidence transcript
Pick only the structured data types that are relevant to a page instead of adding every type Google is interested in.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Most sites now run on content management systems whose plug-ins generate the markup, so check that the plug-ins in your stack produce good markup, for example a calendar plug-in that outputs event markup.
Speaker Ryan LeveringEvidence transcript
Google accepts three structured data syntaxes, JSON-LD, microdata and RDFa, which are all valid and are extracted at the very start into the same pipelines, so they are interpreted identically downstream.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-01glossary term JSON-LD
Google recommends JSON-LD because it is one contiguous block that is easier to author, so people make fewer mistakes with it.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-01glossary term JSON-LD
Microdata has an advantage when payload size matters: embedded in the existing HTML, it avoids duplicating the page content in a separate JSON-LD block.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-01
Test structured data manually first and only then put it into the site's templates.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-11
Structured data that is not relevant to the page's content can be treated as abusive: Google's filters make it ineffective, and egregious cases can lead to a manual action.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-02
Use unique identifiers in structured data, for example the homepage URL in an organisation's url property, so Google can tell which specific entity is meant rather than reading just a name string.
Speaker Ryan LeveringEvidence transcript
Used byrequirements DEV-SDA-04, DEV-SDA-09
Several plug-ins emitting the same markup type is one of the most common structured data problems: the duplicates can make an event details page look like a list of events and change how Google interprets the page.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-06
Google's shopping structured data launches of the previous year (2025) added support for merchant loyalty programs and shipping policies, letting merchants define a policy at organisation level and specify details at product level.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-09
Google's structured data speaker said that in the months before the event Google added support for validity dates on sale prices in product structured data, so merchants no longer need to rush to remove a sale price when the sale ends for fear it shows wrongly in snippets.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-08
More shopping structured data news was left for a Day 3 talk by a Google colleague, Alex.
Speaker Ryan LeveringEvidence transcript
Schema.org, in which Google is a major participant, began publishing usage statistics for all its types and properties in 2026, showing in buckets how many domains use each one; the data is also in schema.org's GitHub repository.
Speaker Ryan LeveringEvidence transcript
Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a chicken-and-egg problem that the schema.org usage statistics are meant to help break.
Speaker Ryan LeveringEvidence transcript
Schema.org added support for RDF lists and sets, which give a way to express ordered values, because RDF triples are not ordered by nature.
Speaker Ryan LeveringEvidence transcript
The SHACL rules are meant to run inside a site's content generation, so markup is sanity-checked before it is published and does not silently regress later, a breakage site owners might otherwise discover only through a Search Console report.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-12glossary term SHACL
The SHACL rules will not replace Search Console as the canonical place for structured data reports, because some checks use Google's internal libraries and cannot be expressed in SHACL, but they will catch problems such as a missing required field.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-12
Google plans a SHACL rule set for each of its structured data feature types and will release the rule sets gradually once they have been checked.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-12
Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its recap.
Speaker Ryan LeveringEvidence transcript
Used byglossary terms Rich results, Structured data
Google keeps adding structured data features and recommendations but also removes them: in the previous year (2025) it removed several features that brought little benefit and were little used.
Speaker Ryan LeveringEvidence transcript
Used byrequirement DEV-SDA-03
Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
Publisher Google Search Central
Used byrequirement DEV-SDA-11
Google's structured data documentation says not to mark up content that is not visible to readers of the page, and not to add structured data about information that users cannot see even if it is accurate.
“don't add structured data about information that is not visible to the user, even if the information is accurate”
Publisher Google Search Central
Used byrequirement DEV-SDA-02
Google's guide to optimizing for generative AI features lists 'overfocusing on structured data' among the things site owners don't need to do: structured data is not required for generative AI search and no special schema.org markup is needed, though it remains worth using because it helps pages become eligible for rich results.
Publisher Google Search Central
Used byrequirement DEV-AIF-02
Google's general structured data guidelines say a structured data manual action removes a page's eligibility for rich results but does not affect how the page ranks in web search.
Publisher Google Search Central
Used byrequirement DEV-SDA-02
Google's documentation updates log calls the July 2026 sale price change a clarification: a new Sale duration section of the merchant listing guide explains validFrom with validThrough or priceValidUntil, aligned with Merchant Center's sale_price_effective_date attribute.
Publisher Google Search Central
Used byrequirement DEV-SDA-08
Google's merchant listing guide says Product rich results only support pages that focus on a single product, or on several variants of the same product.
Publisher Google Search Central
Used byrequirement DEV-SDA-07
Google's merchant listing guide warns that a listing may not display if its priceValidUntil property indicates a past date.
Publisher Google Search Central
Used byrequirement DEV-SDA-08
Google's Event structured data guide says to give a date without a time, such as 2019-08-15, when the start hour is not known, and to include the UTC or GMT offset whenever a time is given.
Publisher Google Search Central
Used byrequirement DEV-SDA-05
The 'non-visible metadata' argument is not a licence to mark up hidden content, since Google's structured data guidelines still require markup to describe what users can see; use markup for the machine-precise form of facts the page shows, such as a full ISO date with time zone for a visible event date or a homepage url for a named organisation.
Author Ibrahim Anjro
Used byrequirements DEV-SDA-02, DEV-SDA-04, DEV-SDA-05
Do not drop markup on the assumption that AI reads the page anyway: by Google's own account LLM extraction is not precise enough for prices and nested offers and too costly to run on every page, so explicit markup remains the dependable route for those facts.
Author Ibrahim Anjro
On product pages with related-product or recently-viewed carousels, make sure the Product markup describes only the main item and its price; Google's own example of what markup prevents was pulling a price from related products.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-07
There is no separate markup for AI features: by Google's account the same processed markup feeds classic results and AI answers, and raw schema.org is generally not passed into model context, so invest in the types Google documents for its features rather than in extra markup written for AI.
Author Ibrahim Anjro
Used byrequirement DEV-AIF-02
Default to JSON-LD and switch to microdata only where page weight is critical, since Google interprets all three syntaxes identically and JSON-LD is the one people get wrong least often.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-01
Audit CMS templates for duplicate markup by running one page per template through the Rich Results Test and checking whether an SEO plug-in and a theme or events plug-in emit the same type twice; 'more markup never hurts' covers relevant, non-duplicated markup only.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-06
For time-limited sales, give the sale price its validity dates in the product markup (Google's merchant listing documentation describes validFrom, validThrough and priceValidUntil) instead of editing markup by hand when the sale ends, and keep the dates aligned with the Merchant Center feed.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-08
Markup for schema.org types that Google does not use yet is a bet on future features; check a type's usage bucket on schema.org before investing, since adoption is what Google says it watches before building a feature.
Author Ibrahim Anjro
When the SHACL rules appear, wire them into the build or CMS publishing step as an automated test, so a template change that drops a required property fails before deployment instead of surfacing weeks later in Search Console.
Author Ibrahim Anjro
Used byrequirement DEV-SDA-12
Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
Ecosystem principle 2, the SERP will evolve: besides organic links and ads, it will hold other elements such as videos and cards.
Search results have moved from ten blue links to feature-rich, media-rich results because users wanted more.
Google said that even the classic ten-blue-links layout was settled only after millions of experiments, as part of using as much data as possible for product decisions.
AI Overviews and AI Mode launched as fairly text-heavy answers with little image content and few tables, and they have become more structured over time because that is what users want.
Ecosystem principle 2, the SERP will evolve: besides organic links and ads, it will hold other elements such as videos and cards.
Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
In a community demo, a fix loop gave a large LLM (Claude Opus) the current markup, the errors Google's test reported and Google's documentation, had it write new JSON-LD and re-ran the test, with at most three attempts; the demo's page passed on the second.
Gemini in Chrome relies heavily on the screenshot it takes of a page.
Google's guide for generative AI features says browser agents may read a site through screenshots, the DOM structure and the accessibility tree, and recommends semantic HTML because it helps users such as screen reader users navigate a page.
Google's systems sometimes fail to determine a page's main content correctly, and structured data helps because site owners tend to mark up what is actually important rather than boilerplate or ads.
Google extracts every element of a page so that it can tell the header, the navigation and the main content apart, and a slide labelled the main content as very important.
The structured data Google processes is not fed very differently to AI Overviews and AI Mode: after cleaning and quality work, the same data goes to both the classic results page and the AI features.
AI Mode and AI Overviews use the same crawling and the same index as Search. At serving they add grounding on the Search index and query fan-out.
As far as the speaker knows, in most of Google's main AI uses a page's schema.org markup is not turned into text and put directly into the model's context; the data is first sorted out, checked for quality and indexed before it is passed on as grounding context.
Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in Search, and that such files neither harm nor help visibility because Google Search ignores them.
Google's guide to optimizing for generative AI features lists 'overfocusing on structured data' among the things site owners don't need to do: structured data is not required for generative AI search and no special schema.org markup is needed, though it remains worth using because it helps pages become eligible for rich results.
Google's guide says new machine-readable files, AI text files or special markup are not needed to appear in Search, and that such files neither harm nor help visibility because Google Search ignores them.
Google often needs to see a markup type adopted on websites before it invests in a feature that uses it, a chicken-and-egg problem that the schema.org usage statistics are meant to help break.
Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its recap.
In a community speaker's case, an LLM asked to write alt text for a product image of a veterinary anxiety medicine for cats described only what it could see, a cat and a veterinarian, and dropped the stress, anxiety and medical-treatment intent of the page.
Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
A community speaker warned that articles still recommend rewriting alt text with AI, called AI a tool, and said that adopting such new AI implementations without applying existing SEO knowledge can damage ranking performance.
Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
Rich results differ from other search features because Google builds them from extra data that site owners provide, usually structured data and usually in JSON-LD format.
Structured data makes pages eligible to appear as rich results, Google's structured data talk said in its recap.
Review structured data lets a site specify how its users rated something; the review snippet shows an average star rating and often the number of reviews of a product, service or piece of content.
Structured data turns the loosely structured web into structured information that powers visual search features such as review stars and recipe filters (for example by preparation time).
Google added six Merchant Center feed attributes for AI shopping experiences: question and answer, documents, related products, item group title and variant option for variants, and popularity rank.
More shopping structured data news was left for a Day 3 talk by a Google colleague, Alex.
Google Shopping called rich data with light structure the AI sweet spot: a little structure that helps the AI system understand the data, not deep, highly nested, complex structures.
Describing every semantic detail of a page in markup is probably not worth the effort; focus on the structured data that Google or other consumers actually use.
In product markup, an offer's shipping and return information can point through a JSON-LD identifier (@id) to shipping and return data defined elsewhere.
Google's shopping structured data launches of the previous year (2025) added support for merchant loyalty programs and shipping policies, letting merchants define a policy at organisation level and specify details at product level.
Google works to make sure that what merchants can express in its shopping feeds can also be expressed in schema.org, adding vocabulary where schema.org lacks it, for example for product details.
Google's structured data speaker said that in the months before the event Google added support for validity dates on sale prices in product structured data, so merchants no longer need to rush to remove a sale price when the sale ends for fear it shows wrongly in snippets.
Google may never use structured data from a site it does not trust: once it sees markup it does not trust, it does not touch it.
Structured data that is not relevant to the page's content can be treated as abusive: Google's filters make it ineffective, and egregious cases can lead to a manual action.
Hallucinations can happen with any AI model, and with current training methods there is no way to get rid of them.
Google's guidance on generative AI content warns that AI output can contain hallucinations and says AI-generated metadata, including structured data and image alt text, should be fact-checked before publishing and the markup validated.
Google's structured data feature guide lists the kinds of structured data Google supports with a search feature and what each can do to a site's search results.
Google recommends using the Search gallery in its developer documentation to find the structured data features that suit a site; the gallery shows each feature and how Google uses the markup.
Web markup is an efficient and unambiguous way for sites to share product data with Google, Google Shopping said, repeating the Day 2 structured data talk.
Parsing structured data is significantly cheaper and more efficient than relying on LLMs for every complex extraction task.