In Part 1, we covered the sitemap, JavaScript, and CSS — the foundation that determines whether a page can be reached and rendered at all. That layer gets a system to your content. It doesn't tell that system what any of it means, or whether it's even allowed to use it. That's the job of the three pieces in this part: robots.txt, schema markup, and llms.txt.
robots.txt: Traffic Control for Humans, Crawlers, and Now, AI Models
robots.txt used to be a simple, mostly-ignored file: block a few admin folders, point to the sitemap, done. That's no longer true. robots.txt is now doing real strategic work across three different audiences at once:
- Traditional search crawlers (Googlebot, Bingbot) — standard indexing behavior, largely unchanged.
- SEO-tool and scraper bots (Ahrefs, Semrush, MJ12, and similar) — many site owners now explicitly block these to preserve crawl budget and bandwidth for bots that actually feed customer-facing discovery, rather than third-party rank-tracking tools.
- AI crawlers (GPTBot, ClaudeBot, PerplexityBot, Google-Extended, Amazonbot, and a fast-growing list of others) — this is the newest and most consequential category. A robots.txt file that blanket-blocks all bots, written before AI crawlers existed, can accidentally shut a business out of ChatGPT, Perplexity, Gemini, and every other AI-driven discovery channel without anyone realizing it happened.
For GEO in particular, robots.txt is now a direct visibility lever: if you want your business recommended inside AI-generated answers, the crawlers that build those answers need explicit permission to read your site. Reviewing robots.txt with AI crawlers specifically in mind — rather than treating it as a "set once and forget" file — is one of the highest-leverage, lowest-effort changes available for GEO visibility today.
Schema Markup: The Rosetta Stone for AI Systems
If JavaScript and CSS determine whether content can be seen, schema markup determines whether it can be understood correctly. Schema is structured data — a machine-readable layer added to a page that explicitly labels what things are: this is a business, this is its address, this is a review, this is a frequently asked question, this is the answer to it.
Without schema, a crawler is left inferring meaning from unstructured text and page layout — a process prone to error. With schema, there's no ambiguity. This is the single highest-leverage technical investment for AEO and GEO specifically, because both disciplines depend on a system extracting a precise fact and presenting it with confidence:
- FAQPage schema maps question-and-answer content almost directly onto how AI systems retrieve and surface answers.
- LocalBusiness / Organization schema gives AI systems a verified, structured version of your name, address, hours, and services — the exact facts assistants get asked about constantly ("is this place open now," "what does this business do").
- Review / AggregateRating schema feeds trust signals that both traditional search snippets and AI recommendation systems draw on when comparing multiple candidates for a query.
- BreadcrumbList and Article/BlogPosting schema help systems understand site structure and content hierarchy, which affects both traditional rich results and how confidently an AI system attributes a piece of information to a specific, credible page.
Schema doesn't change what a human sees on the page. It changes how confidently and accurately a machine — search engine or AI model — can describe what's on that page to someone else. That distinction is the entire game in AEO and GEO.
llms.txt: The New Front Door for Generative Engines
llms.txt is the newest addition to this stack, and it exists specifically for GEO. Where robots.txt tells crawlers what they're allowed to access, llms.txt goes a step further: it's a concise, curated summary of a site — written specifically for large language models — that points to the most important pages, explains what the business does, and gives an AI system a fast, reliable orientation instead of forcing it to reconstruct that understanding from scattered crawled pages.
Think of it as the difference between handing someone a stack of loose documents and handing them a one-page executive summary with a table of contents. Both technically contain the same information, but one gets processed faster, more accurately, and with far less risk of an AI system missing or misrepresenting something important.
Adoption of llms.txt is still early, which is exactly why it matters right now — the sites present in AI training and retrieval pipelines with clean, current llms.txt files have a real head start on being the sources those systems default to, in the same way early structured-data adopters had an advantage in rich search results years before it became standard practice.
The pattern across all six of these — sitemap, JS, CSS, robots.txt, schema, and llms.txt — is the same. None of them change what a human visitor experiences. Every one of them changes what a machine can reliably extract, trust, and act on. That's precisely why they get overlooked: there's no visual before-and-after to point to, just a difference in whether your business shows up in the places people are increasingly starting their search.
This has become core technical infrastructure work for us across every industry we serve — rebuilding sitemaps to be self-updating, auditing JavaScript-dependent content, adding schema layer by layer, and writing llms.txt files from scratch for clients who've never heard the term but whose competitors are quietly already ahead on it.
Coming Up in Part 3
Sitemaps, JavaScript, CSS, robots.txt, schema, and llms.txt cover the six pillars most people picture when they think "technical SEO." Part 3 looks at a seventh piece that's easy to overlook — cookies and consent — and then pulls all seven together into a single technical visibility checklist.
Not sure if your robots.txt, schema, or llms.txt are helping or hurting your AI visibility?
We'll review what your bot rules actually allow, audit your schema coverage, and build or fix your llms.txt file — so nothing is quietly locking AI systems out.
Get Started