Structured Data, llms.txt and the Technical Basics of AI Search
Let AI crawlers in through robots.txt, mark up pages with schema.org, serve content as HTML and keep sitemaps clean. Add llms.txt, but expect little from it.
The technical basics of AI search are mostly the basics of good SEO, done carefully. Let the right AI crawlers into your site through robots.txt, serve your important content as real HTML rather than something that only appears after JavaScript runs, describe your business and products with schema.org structured data, and keep an accurate sitemap. An llms.txt file is cheap to add, but it is a proposed convention that the major AI tools have not confirmed they use, so do not expect much from it. None of this gets you recommended on its own. It removes the reasons an AI tool might fail to read or understand you.
Start with robots.txt
Your robots.txt file sits at the root of your domain and tells crawlers what they may fetch. AI companies now run several different crawlers, and they do different jobs. Blocking the wrong one can quietly remove you from answers.
The ones worth knowing:
- OAI-SearchBot is what OpenAI uses to find pages for search results inside ChatGPT. If you block it, your pages are less likely to be shown or linked in ChatGPT search answers.
- GPTBot is OpenAI’s crawler for collecting training data. Blocking it is a choice about training, separate from search.
- ChatGPT-User fetches a page when a user’s request causes ChatGPT to visit it.
- PerplexityBot indexes pages for Perplexity’s answers.
- Google-Extended is not a separate crawler. It is a control token. It tells Google whether your content may be used for Gemini model training and some Gemini uses. It does not affect Google Search, and it does not keep you out of AI Overviews, which are part of Search and governed by Googlebot.
- ClaudeBot and Claude-SearchBot are Anthropic’s crawlers for training and for search respectively.
Crawler names and their exact roles do change, so check each company’s own documentation before you edit the file.
For most brands that want to be recommended, the sensible default is to allow the search crawlers. A simple robots.txt that does that looks like this:
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Googlebot
Allow: /
User-agent: *
Disallow: /cart
Disallow: /checkout
Disallow: /account
Sitemap: https://www.example.com/sitemap.xml
If you prefer not to have your content used for training, you can add a block for GPTBot and a block for Google-Extended with “Disallow: /”. That is a business decision. Just make sure you are not blocking a search crawler by accident.
Two common mistakes. First, a firewall or bot protection service blocking AI crawlers even though robots.txt allows them. Check your CDN settings, not only the file. Second, a staging rule like “Disallow: /” left in place after a site launch. It happens more than anyone admits.
Make sure the content is in the HTML
Many AI crawlers do not run JavaScript the way a browser does, or do it less reliably than Googlebot. If your product descriptions, reviews, prices or FAQ answers only load after scripts run, a crawler may see an almost empty page.
A quick test: open a product page, view the page source (not the inspector), and search for a sentence from your product description. If it is not there, the crawler may not see it either. The fix is server-side rendering or static generation, which most modern ecommerce platforms and frameworks support. Shopify themes generally ship product content in the HTML already. Heavily customised headless builds are where problems tend to show up.
Also check that your key facts are written as text. Shipping policy in an image, sizing in a PDF, and ingredients in a carousel graphic are all hard for a model to read.
Use schema.org structured data
Structured data is a block of code, usually JSON-LD, that labels what a page is about. Search engines use it to understand pages, and it is reasonable to expect AI tools that draw on search indexes to benefit from that clarity. No AI company has promised it changes their answers, so treat it as good hygiene rather than a lever.
The types that matter most for an ecommerce brand:
- Organization on your home page. Your official name, logo, URL, contact details and “sameAs” links to your real profiles on other sites. This helps tie your brand together across the web.
- Product on product pages, with offers, availability, brand and, where they are genuine, aggregate ratings.
- FAQPage on pages with real questions and answers. Google has narrowed where it shows FAQ rich results, but the markup still describes the page clearly.
- Article or BlogPosting on blog content, with the author and date.
- BreadcrumbList to show how pages relate.
- LocalBusiness if you also have a physical store.
The rule with structured data is simple. It must match what is visible on the page. Marking up reviews you do not show, or ratings you made up, breaks Google’s guidelines and can get your rich results removed. Validate your markup with Google’s Rich Results Test and the Schema Markup Validator.
Keep your sitemap honest
A sitemap lists the pages you want found. Keep it to live, canonical pages that return a normal response. Remove redirected, deleted and noindexed URLs. Include a last modified date that changes only when the page really changes. Submit it in Google Search Console and Bing Webmaster Tools.
Bing matters more than people think here. Some AI tools have drawn on Bing’s index, so make sure Bing can crawl you and that your site is verified there.
What llms.txt is, and what it is not
llms.txt is a proposed convention. The idea is a plain text file at the root of your site, written in Markdown, that gives language models a short summary of who you are and links to your most useful pages. A minimal version:
# Example Brand
> Example Brand makes organic dog treats for dogs with sensitive stomachs.
## Key pages
- [Shop all treats](https://www.example.com/collections/all)
- [Ingredients and sourcing](https://www.example.com/pages/ingredients)
- [Shipping and returns](https://www.example.com/pages/shipping)
Here is the honest part. The major AI tools have not confirmed that they read llms.txt when building answers. Some developer tools use it for documentation. For a brand trying to appear in ChatGPT or AI Overviews, there is no public evidence it changes anything.
So add one if you like. It takes half an hour and costs nothing to host. Then spend your real effort elsewhere. Be wary of anyone selling llms.txt as the main event.
Speed, errors and the boring stuff
A few more checks that help every crawler:
- Pages load quickly and return clean status codes.
- One canonical URL per product, with variants handled consistently.
- No important pages blocked by “noindex” by mistake.
- Clear titles and headings that say what the page is.
- An About page that states plainly what you sell, who you are and where you are based.
What the technical work can and cannot do
Technical work makes you readable. It does not make you recommended. Whether ChatGPT or Gemini names you depends far more on whether the wider web talks about you clearly and consistently, which we cover in why consistent mentions matter for AI answers. For how Google picks the links inside its summaries, see how Google AI Overviews choose sources.
If you would rather have someone handle the technical side, our page on AI Overviews optimization explains what that work involves, and our ranking of AI SEO agencies compares who does it.
The short version
- Allow OAI-SearchBot, PerplexityBot and Googlebot. Decide on GPTBot and Google-Extended separately.
- Check your firewall as well as robots.txt.
- Put key content in the HTML, as text.
- Add Organization, Product, FAQPage and Article markup that matches the page.
- Keep sitemaps clean and verify with Bing too.
- llms.txt is cheap and unproven. Add it, then move on.