One in Seven Practices Cannot Be Read
We tested 12 AI crawlers against 56 private practice therapy websites. Eight refuse the crawlers behind ChatGPT, Claude or Perplexity, and seven of those eight never wrote that rule: a server or a CDN did. The bigger number, two thirds of the sample, we threw away, and this explains why.
The number we threw away, and why
The first pass said 37 of 56 sites, two thirds of the sample, block at least one AI crawler. That headline was available, it was technically true, and it was useless.
Thirty five of those 37 cases were the same crawler: a bulk collector that many hosts block by default because it requests aggressively. Blocking it is a decision about server load. It is not a decision about whether a person asking an assistant for a therapist will ever see your practice.
So we split the twelve crawlers into two groups and asked two different questions. The first group is the assistants people actually use: the crawlers behind ChatGPT, Claude, Perplexity and Google's AI surfaces. Blocking those has a consequence you can describe in one sentence. The second group collects at scale for training. Blocking those is defensible and says nothing about visibility.
Everything below counts the first group only.
The two layers, because one of them is invisible
Two places decide whether a crawler can read your site, and only one of them is public.
The first is robots.txt, a text file at the root of your domain that names which crawlers may request which paths. Anyone can read it, including you, right now. It is where every audit stops.
The second is the response itself. A site can hand a page to a browser and refuse the identical page to a named crawler, with a 403, at the edge, before any of your website's own code runs. That block appears nowhere in robots.txt. It arrives from a CDN setting, a security plugin, or a managed rule somebody enabled once and never revisited.
We checked both layers for each of the 12 crawlers: the longest matching rule in robots.txt, and then a real request carrying that crawler's user agent.
Two controls, because without them the test accuses innocent sites
The first control is a generic bot. Before recording a block, we requested the same page with a bot user agent that has nothing to do with AI. A site that refuses every non browser is not making a decision about AI, it is making a decision about bots, and it cannot count. Two sites failed this control, so we dropped them from the sample entirely.
The second control is patience. We treated every 429 as our own rate limit rather than as their policy: we waited, then asked again more slowly. Five sites produced a 429 at some point, and we counted none of them as blocking.
That is why the denominator below is 56 and not 58.
What we found
| Assistant crawler | Sites that refuse it | Of |
|---|---|---|
| ClaudeBot | 8 | 56 |
| GPTBot | 7 | 56 |
| OAI-SearchBot | 6 | 56 |
| ChatGPT-User | 6 | 56 |
| PerplexityBot | 6 | 56 |
| Claude-SearchBot | 5 | 56 |
| Perplexity-User | 5 | 56 |
| Google-Extended | 1 | 56 |
Eight distinct practices out of 56 close the door to at least one assistant crawler. That is roughly one in seven.
The shape of those eight matters more than the count. Five of them refuse seven of the eight assistant crawlers at once, which is not what a considered policy looks like. A practice that decided to keep ChatGPT out would write one line. Refusing almost every assistant simultaneously, with nothing in robots.txt, is the signature of a single switch somewhere in front of the website.
Seven of the eight never declared it. Only one site has the rule written in its own robots.txt, where the owner could find it. In the other seven the block is invisible from inside the website.
The contrast nobody expects
Fourteen of the 56 practices publish an llms.txt, the newer file meant to help assistants read a site. That is almost twice as many sites doing the optional new thing as sites accidentally locked out of the assistants entirely.
Read that in order. Fourteen practices went looking for the current best practice and added a file. Eight have a door closed that they did not close and cannot see. The work went to the part that is visible in a blog post, not to the part that decides whether a crawler gets a page at all.
One more measurement, because it corrects something we expected to find: not a single site in the sample uses the noai meta tag. Whatever is happening here, almost nobody chose it deliberately.
How to check your own, in four minutes
None of this needs a tool or a consultant.
First, open your robots.txt
Type your domain followed by /robots.txt in a browser. If you see rules naming GPTBot, ClaudeBot, PerplexityBot or Google-Extended with Disallow, that is a declared block. If the file is empty or missing, that layer is clear, which is not the same as open.
Second, look at what sits in front of your site
Most invisible blocks come from a CDN or a security layer, not from your website builder. If your domain runs through a service that filters traffic, find its bot settings. Rules that promise to stop scrapers or AI bots do exactly that, including the assistants you want.
Third, ask an assistant for your own page
Give ChatGPT or Claude the URL of your practice page and ask it to tell you what the page says. If it cannot fetch it while the page loads fine in your browser, you have the second layer problem, and you just proved it in thirty seconds.
Fourth, ask your host the one useful question
The question is not whether they block AI. It is: does anything in front of my site return 403 to a named crawler? A support agent can answer that in a sentence.
What this does not prove, and it matters
Being readable is not being cited. Forty eight of these 56 practices can be fetched perfectly well by every assistant crawler we tested, and most of them still will not come up when someone asks for a therapist in their city. Access is the floor, not the outcome. A crawler that can read a page nobody would ever quote has changed nothing.
What the measurement does prove is narrower and worth knowing: for about one in seven practices, the work of being chosen has not started, because the page cannot be read at all. And in seven of those eight cases, nobody made that decision. It was made for them, by something they are paying for.
If you want the rest of the picture, the part that decides whether an assistant quotes you once it can read you, that is what our method measures, and how to make a page worth citing is the practical version. The file everyone is adding is covered in what to actually put in llms.txt.
We did not name any of the 56 practices. They did not ask to be measured, and seven of the eight blocked sites are not doing anything wrong: they are the ones with the least visibility into their own setup. Naming them would punish the people this piece is meant to help.
Questions we get about this
Does blocking an AI crawler hurt my Google ranking?
No. These are separate crawlers from the one that indexes your site for classic search results. Blocking GPTBot does nothing to your position in Google's blue links. What it changes is whether an assistant can read your page when someone asks it a question.
My robots.txt is empty, so I am fine, right?
Not necessarily. In seven of the eight blocked practices we found, robots.txt was clean and the refusal came from a layer in front of the website. An empty robots.txt means you did not write a block. It does not mean nobody else did.
Should I add an llms.txt file?
It does not hurt, and fourteen of these 56 practices already have one. Just be aware of the order: a file that tells an assistant what your site contains is useless if the assistant cannot fetch your pages in the first place. Check the door before you add the sign.
Is being readable enough to get recommended?
No, and this is the honest limit of everything above. Forty eight of these 56 practices are perfectly readable and most still will not be named when someone asks for a therapist in their city. Access is the floor, not the outcome.
Why did you not name the practices you tested?
They did not ask to be measured, and seven of the eight blocked sites are not doing anything wrong. They are the ones with the least visibility into their own setup, and naming them would punish exactly the people the piece is meant to help.
If you want to know which of the two problems you have, the door or the citation, book a call and we will check both while you are on it.

