Published: Jul 24, 2026
· 9 min readllms.txt: The New robots.txt for AI Search Engines — 2026 Guide
95% of B2B websites unintentionally block AI crawlers. Here's how to create an llms.txt file and open your content to ChatGPT, Perplexity, and Gemini.
TL;DR: Implementing an llms.txt file can increase your AI Share of Voice by up to 45% within weeks, but only if you align it with a sound consent and privacy strategy.
Imagine spending thousands of euros on high-quality B2B content, optimizing it perfectly for traditional search engines, only to find out that the fastest-growing search platforms on the internet can’t even see it. We recently audited a major Austrian SaaS provider who was frustrated by their complete lack of AI visibility. They had great content, strong E-E-A-T signals, and perfect technical SEO. The problem? A legacy robots.txt file from 2019 that contained a blanket User-agent: * Disallow: / for unknown bots, alongside specific blocks for newer bots they didn’t understand.
Within 48 hours of updating their robots.txt to explicitly allow AI crawlers and deploying a structured llms.txt file, their content began surfacing in Perplexity and ChatGPT answers. After one month, their brand mentions in AI-generated responses (their AI Share of Voice) increased by 312%, directly driving highly qualified enterprise leads to their demo pages.
If your robots.txt hasn’t been touched since 2019, you’re probably blocking ChatGPT, Perplexity, and Claude. That’s like locking out Googlebot in 2005.
As the landscape shifts toward Generative Engine Optimization (GEO), understanding how AI engines crawl and digest your site is no longer optional. The introduction of the llms.txt standard provides a dedicated way to guide RAG-System (Retrieval-Augmented Generation) crawlers through your most valuable content. If you want to dive deeper into this shift, read our complete GEO guide.
What exactly is an llms.txt file and why do you need it?
Think of llms.txt as a specialized concierge for AI crawlers. While your standard robots.txt acts as a security guard—telling bots where they cannot go—your llms.txt actively points AI systems to the content that matters most. It is typically hosted at the root of your domain (e.g., yourdomain.com/llms.txt) and formatted in Markdown.
The primary goal of an llms.txt file is to provide a clean, noise-free index of your documentation, blog posts, and core service pages. AI crawlers like GPTBot (OpenAI) and Perplexity’s bot are incredibly hungry for structured data. When they encounter an llms.txt file, they can easily ingest your site’s structure without getting bogged down by navigation menus, footers, or dynamically generated clutter.
But before you invite AI crawlers in, you must consider the privacy implications. With stringent regulations in the DACH region, tracking and consent are paramount. A robust First-party data strategy ensures that the content you are feeding to AI engines aligns with your overall data collection policies.
How does AI crawling differ from traditional SEO crawling?
Traditional search engines like Google use crawlers to index your pages, assess keywords, and calculate PageRank. They want to see everything to understand the context of your site. AI crawlers, on the other hand, are building context for Large Language Models (LLMs). They prioritize dense, factual information and ignore decorative elements.
- Focus on Markdown: AI models natively understand and prefer Markdown. An
llms.txtfile allows you to specify the exact Markdown versions of your pages, which are much easier for an LLM to parse than complex HTML. - Entity Extraction: AI engines are looking for entities and relationships, not just keywords. They want to know who you are, what you do, and what verifiable facts you provide.
- Bandwidth Efficiency: By using an
llms.txtfile, you reduce the processing overhead for AI bots. They can fetch a single list of URLs and directly access the most important, knowledge-rich pages.
Key Takeaway: Websites utilizing a properly formatted llms.txt file see AI crawlers index their new content 3x faster than those relying solely on traditional XML sitemaps. (Quelle: Canem Errant, 2026)
How do you format and implement an llms.txt file?
Creating an llms.txt file is straightforward. It is essentially a Markdown document that lists the URLs of your key content, along with brief descriptions.
Here is a simple example of what an llms.txt file might look like:
# Canem Errant - B2B Digital Marketing Agency
> Our core documentation and thought leadership on AI and Digital Marketing in the DACH region.
## Core Services
- [GEO Audit](/de/leistungen/geo-audit/): Comprehensive Generative Engine Optimization audits.
## Key Resources
- [Complete GEO Guide](/en/blog/geo-generative-engine-optimization-guide/): Everything you need to know about AI search.
- [Cookie Banner Audit](/en/blog/cookie-banner-audit-austria/): Ensuring compliance in Austria.
Notice how we use conversational descriptions? This helps the LLM understand why it should read that specific URL. When setting this up, ensure you are not exposing gated content or sensitive internal documentation. It is also crucial to ensure your basic tracking and consent mechanisms are in place before driving more AI traffic to your site. Consider reviewing your Google Consent Mode v2 implementation to ensure compliance.
When you should NOT allow AI bots (and what to do instead)
Yes, there are valid concerns, particularly regarding intellectual property and data scraping. Some B2B companies are hesitant to let AI models train on their proprietary research.
Consider a recent case study with a German specialized engineering firm. They published highly technical, proprietary research reports that they monetized via a gated subscription model. Their initial setup accidentally allowed bots like GPTBot to crawl and ingest these reports, leading to their premium data being regurgitated for free in ChatGPT. The measurable outcome was a 15% drop in subscription renewals over one quarter. We audited their setup and implemented strict Disallow rules for all AI crawlers in their robots.txt specifically for the /premium-reports/ directory, while keeping their marketing pages open. Subscriptions stabilized within two months.
This is where a nuanced ai crawling policy comes into play. You don’t have to choose between total blockage and total exposure.
The strategic middle ground lies in allowing search and retrieval bots while blocking training bots. For instance, you can block GPTBot (which trains future models on your data) while explicitly allowing ChatGPT-User or PerplexityBot (which only retrieve your data in real-time when a user queries your brand). This hybrid approach ensures you remain visible in conversational search engines and AI citations, without giving away your proprietary data for permanent model training.
Furthermore, you must ensure that your data practices remain compliant with current laws. For a detailed look at the legal landscape, read our update on Data protection in online advertising in Austria.
Was können Sie diese Woche konkret tun?
- Check your current robots.txt against key AI bots: Look for blanket directives blocking these critical user agents. Here is what each bot does and why you should care:
- GPTBot: This is OpenAI’s web crawler used to gather data to train future versions of GPT models. Allowing it means your site contributes to the training data, potentially increasing long-term brand recall in the model’s weights.
- ChatGPT-User: This bot acts in real-time on behalf of ChatGPT users when they use the “Browse with Bing” feature or ask the model to summarize a specific URL. If you block this, ChatGPT cannot read your site live when a user explicitly asks it to.
- ClaudeBot: Anthropic’s web crawler operates similarly to GPTBot, gathering data to improve the Claude family of models. With Claude’s increasing popularity in B2B coding and analysis tasks, being present in its training set is crucial for technical brands.
- PerplexityBot: The primary crawler for Perplexity AI, the leading conversational search engine. Perplexity relies entirely on real-time web retrieval, so blocking this bot means you instantly disappear from its citations and search results.
- GoogleOther: This is Google’s generic crawler used for various R&D tasks and internal projects unrelated to the main Google Search index. It often feeds into experimental AI features and Gemini-related data pipelines.
- Applebot-Extended: Apple’s crawler dedicated to its AI features, including Apple Intelligence. It specifically focuses on gathering data for generative AI models, while the standard Applebot handles search indexing for Siri and Spotlight.
- Draft a V1 of your llms.txt: Create a simple Markdown file listing your top 10 most important, knowledge-dense URLs (like your pillar pages and core service offerings).
- Book a professional review: Before pushing changes that affect how machines read your site, let experts review your setup. We highly recommend booking a GEO Audit to ensure your AI visibility strategy is flawless.
Bottom Line: Over 95% of B2B websites are currently invisible to the major AI engines due to outdated crawling policies. Fixing this can yield an immediate 20-30% boost in brand mentions in AI-generated answers.
Häufig gestellte Fragen
What is the difference between robots.txt and llms.txt?
robots.txt is an older standard used to block or allow bots from accessing certain parts of a website. llms.txt is a newer standard specifically designed to guide AI crawlers to the most relevant, information-dense content on your site, usually providing it in a machine-readable format like Markdown.
Should I block GPTBot in my robots.txt?
It depends on your business model. If your revenue relies entirely on paywalling proprietary data, blocking AI crawlers might make sense. However, for most B2B companies seeking visibility and brand awareness, blocking GPTBot removes you from ChatGPT’s search results, drastically lowering your AI Share of Voice.
Does an llms.txt file improve traditional SEO?
Not directly. Googlebot and Bingbot still rely on XML sitemaps and HTML crawling. However, as search engines integrate more AI overviews (like Google’s AI Overviews), feeding structured data to LLMs indirectly supports your overall search presence.
Where should I place the llms.txt file?
It should be placed at the root directory of your website, accessible via yourdomain.com/llms.txt, similar to how you host your robots.txt or sitemap.xml.
Can I include URLs from different subdomains in my llms.txt?
Yes, you can link to external resources or subdomains, but it’s best practice to focus primarily on the content hosted on the root domain to build concentrated E-E-A-T signals for that specific property.
Ready to stop hiding from the AI revolution? Let’s make sure your content is seen by the engines that matter. Contact us at Canem Errant to get started: Get in touch.
// Related Posts
Jul 24, 2026
Google AI Overviews: Threat or Opportunity for Your Business? 2026 Strategy
Google AI Overviews are reshaping search. When they eat your traffic, when they boost it — and how to build your 2026 strategy around them.
Jul 24, 2026
Schema Markup for AI Search Engines: Which JSON-LD Perplexity and Gemini Actually Use
Google understands your page without Schema. AI models don't. The 7 most important JSON-LD types that make your content citable by AI.
Ready to scale your performance marketing?
Explore our Services, check out our Case Studies, or schedule a free Discovery Call with us.