You published that guide. You checked every claim before it went live, cited your sources, and it ranked. Then six months later, a chatbot answers the exact question your article covers, using your structure, in something close to your words, and the reader never leaves the chat window. No visit. No name mentioned. No click back to your site.
That gut reaction, the "wait, that's mine" feeling, is fair. Bots crawl your site every day, and plenty of what they take never comes back as so much as a mention.
But some of the bots in your logs are the reason your brand shows up in an AI answer at all, so blocking every one of them backfires. The skill worth building here is telling them apart: shut the door on the ones that only take, and leave it open for the ones that can send you a mention, a citation, or a customer.
So who's taking it?
That comes down to which bot visited and what it did with your page. Is AI Stealing My Website Content? covers that argument in full: the difference between a bot that mentions your brand and one that folds your words into an answer with nobody credited. This piece skips the debate and goes straight to what you control: access.
If you want the deeper case for why "stealing" undersells the problem, and why "fair use" oversimplifies it too, that's where to go. Short version here: the content is yours, the access to it is one-sided by default, and the fix starts in a plain text file on your server, and you can write it this afternoon.
What happens when an AI crawler visits my site?
Every AI crawler identifies itself by name, both in your server logs and in the request it sends: GPTBot, ClaudeBot, PerplexityBot, and a growing list of others. Some only train future models. Some fetch a page in real time because a customer asked a live question. Your robots.txt file can allow or block each one separately, by name.
That distinction matters more than most site owners realize. OpenAI runs at least three separate crawlers under different names. GPTBot pulls pages for training and has nothing to do with what ChatGPT tells a user right now. OAI-SearchBot is the one that decides whether your site can surface in ChatGPT's search results. ChatGPT-User shows up when a person's request sends ChatGPT to your page, and OpenAI notes that robots.txt rules may not apply to it, since a person asked. Block the wrong one of these three and you can disappear from live ChatGPT answers while GPTBot keeps crawling anyway. Or block GPTBot and change nothing about whether you show up in an answer today.
Anthropic runs the same three-way split. ClaudeBot collects pages for training. Claude-SearchBot crawls to improve the results Claude returns when it searches the web. Claude-User fetches a page when a person asks Claude a question that sends it there. Anthropic's crawler docs Perplexity runs PerplexityBot to index pages for its answer engine, and Perplexity-User when a person's question sends it to fetch a page live. Perplexity says that one generally ignores robots.txt, because a person asked for the page. Google's version is the strangest of all: Google-Extended is a flag inside robots.txt, and you'll never see it in your server logs. It tells Google whether pages it already crawled for Search can be used to train Gemini and to ground Gemini's answers. Block it and your Google Search rankings stay where they are. Apple's Applebot-Extended works the same way for Apple Intelligence.
One more thing worth knowing before you touch robots.txt: a bot can put whatever it wants in its user-agent string. The major operators publish IP ranges (OpenAI, Anthropic, Perplexity, Google) you can check server logs against, because anyone can fake a header and nobody can fake the IP address the request came from.
How do I block the bots I don't want?
Robots.txt takes a separate rule block for every crawler you want to name, and any bot without its own block falls back to whatever you set under User-agent: *. Naming each AI bot individually is what keeps you from locking out search traffic by accident.
A block aimed at OpenAI's training crawler alone looks like this:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /That keeps GPTBot out of the training pipeline while leaving ChatGPT's live retrieval and browsing untouched.
Widen it out to the rest of the field. Here's one possible setup, for a site that's comfortable with the big AI vendors training on its pages and draws the line at open datasets anyone can download:
User-agent: ClaudeBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Google-Extended
Allow: /
User-agent: Applebot-Extended
Allow: /
User-agent: CCBot
Disallow: /CCBot feeds the Common Crawl dataset, which a long list of open models train on with no relationship back to you at all. Blocking it costs you nothing you were getting anyway.
What you don't want is a bare User-agent: * / Disallow: / sitting at the top of the file with nothing else underneath it. That blocks the search crawlers you need along with everything else. It usually gets there by accident: a staging site's robots.txt rides along to production on launch day, and nobody looks at the file again.
Can I let AI see part of my site and block the rest?
Yes. Robots.txt works at the path level as well as the whole-site level, so you can block a bot from specific directories while leaving everything else open. llms.txt works in the opposite direction: it only lists what you want surfaced, so anything you leave out is already excluded by omission.
Say your blog and product pages are what you want an AI system reading, but your customer portal, internal wiki, and gated reports aren't. Block the paths you want private and leave the rest open:
User-agent: ClaudeBot
Disallow: /account/
Disallow: /internal/
Disallow: /reports/gated/
Allow: /This works because robots.txt applies whichever rule most specifically matches a given path, regardless of line order in the file. A Disallow: /account/ overrides a blanket Allow: / for that same path. Repeat the path rules under every bot's own User-agent block. There's no shortcut that applies one path list to every crawler at once.
llms.txt has no block directive at all. It's a curated list of what you want an AI system to prioritize, and leaving a page off the list excludes it by omission. A SaaS company might list its docs, pricing, and blog in llms.txt while leaving its careers page and legal terms off. Robots.txt still governs whether a bot can reach those pages in the first place. Leaving them off the list is a separate signal: these aren't the pages the company wants an AI answer built from.
One limit worth knowing before you lean on path rules alone: robots.txt runs on the honor system. TollBit's State of the Bots report for the second half of 2025 found that roughly 30% of AI bot scrapes ignored robots.txt, and that ChatGPT-User fetched disallowed content 42% of the time. If a page holds something you can't risk having scraped, put it behind a login.
Which bots should I block, and which ones will hurt me if I do?
Block the crawlers that only train future models and never send a mention back, things like CCBot, or any vendor's training-only bot once you've decided you don't want your content in that dataset. Leave open anything tagged as retrieval, search, or "user," because those are the bots deciding what an AI answer shows a customer today.
The instinct to block "AI" wholesale is understandable, and it's also the version of this that backfires hardest. Picture a site that copies a "block all AI" list off a forum and pastes it in one pass: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot. It feels like the door is locked. A few months later, its AI Visibility Report shows a citation rate near zero on ChatGPT and Perplexity. The training bots were never the ones putting it in front of customers. The retrieval bots were, and those got shut out along with everything else.
Think of it as a guest list. Training crawlers are the ones photographing your house for a catalog you never agreed to be in. Retrieval crawlers are the ones walking a customer through your door right now, because that customer already asked to see the house. Block the photographer if that's the line you want to draw. Block the second kind and you've turned away someone who was already on their way in.
How do I actively share what I want AI to find?
llms.txt is a plain-text file at your site root that summarizes what you do and points to your best pages, built for AI systems the way robots.txt was built for crawlers. It's worth publishing, with modest expectations: no major AI vendor has formally committed to reading it, and most published files sit unread.
The honest state of llms.txt in 2026: adoption is growing fast off a small base, and reading isn't following. Rankability's tracker found an llms.txt on 8.7% of the top 1,000 sites in June 2026, and 9.3% by its September scan. Originality.ai's year-long tracker found the total count of published llms.txt files grew 8.8x in twelve months, from around 4,100 in June 2025 to over 36,000 by May 2026. Ahrefs pulled May 2026 server logs from 137,000 domains and found that 97% of llms.txt files got zero requests that month. Where requests did land, GPTBot made up about 4.5% of them and ClaudeBot under 1%. Google's John Mueller has publicly compared llms.txt to the keywords meta tag, the one search engines stopped reading years ago.
Publish one anyway. A file that costs an afternoon and might help is worth the afternoon. It works as a curation layer for the engines willing to read it. Whether they can reach your content in the first place is the crawler rules' job, covered above.
The parts of "sharing" that carry more weight right now: keep your highest-value pages crawlable without a JavaScript wall in front of them, keep a current XML sitemap, and add structured data (Article, Product, and FAQ markup, each only on pages that match it) so a crawler doesn't have to guess what it's looking at.
How do I know if any of this is working?
Robots.txt and llms.txt control access. Neither one tells you what happened after access was granted. That's a separate question: whether the engines that could read your content are mentioning you when they answer, using your words with someone else's name attached, or leaving your name out entirely.
That gap is what an AI Visibility Report is built to check. It measures the two things access settings can't: your mention rate, how often an engine says your brand name in the answer, and your citation rate, how often it links your site as the source. The gap between them is the ghost citation: your page did the work, and your name never made it into the sentence. Get every bot setting right, get crawled by everything on this list, and a site can still land a low GEO Citation Score if nothing it publishes gives an engine a reason to say the name. Mentions vs. Citations breaks down that gap in more depth, and which lever to fix first walks through what to do once you know which side of it you're on.
If you'd rather check by hand before running a full report, how to find ghost citations by hand covers the manual process.
Getting the access settings right is the first half of the job. A free scan is how you check the second half: it reads your homepage and asks one real buyer question on ChatGPT, Google AI Overviews, Gemini and Perplexity, in about a minute.
Run a free scan