We Measured 54 Therapy Practice Websites, and Access Was Never the Problem

We measured 54 therapy practice websites in ten American cities. Machines can reach almost all of them without trouble. Machines can quote almost none of them. Access turned out to be the part everybody already has, and the part nobody thinks about is the one that is missing.

Why we measured therapists, and why we are telling you first

Contento sells to this niche. There is a page on this site written for therapists, and if the numbers below make private practices look like they need help, we are the ones offering it. You should read the rest with that in mind, which is exactly why the method comes before the results and the sample list is published alongside them.

We picked therapy practices for a duller reason too. They are small, independent, owner run, and they buy their websites once. That makes them a fair test of what a competent small business site looks like when nobody on the team thinks about machines at all.

How the sample was built

One search, repeated identically in ten metros: Denver, Seattle, Austin, Atlanta, Chicago, Boston, Phoenix, Portland, Nashville and Minneapolis. From the first page of results we kept the sites of practices that see clients. We dropped directories, marketplaces, national chains, office rental businesses and hospital systems, and our notes name every exclusion rather than summarising them.

That produced 58 candidate sites. Four could not be read from where we measure: two returned 403 to every agent we tried, including a browser, and two never answered. We report those four as unreadable and keep them out of every average. Scoring a site zero because our own tooling could not see it is the cheapest way to publish a false number, and we would rather have a smaller sample than a wrong one.

So the sample is 54 sites, and 50 of them were measured across eight pages each.

One bias matters more than the others, and it runs in a direction we cannot correct. We sampled sites that already appear in a public search, so the first of the three scores is inflated by construction. What this sample can say honestly is how the practices that already win the search perform on everything that happens after somebody arrives.

The three scores

Our method page describes the instrument. It reads the HTML a site actually serves, with no session, no paid tools and no access to anybody's analytics, and it scores twelve signals grouped into three questions. Six of the twelve can be measured from outside. The other six cannot, so the instrument marks them not measured rather than guessing, and each score covers only the weight it really saw.

ScoreThe question it asksAverageMedianRange
FoundDo they turn up when somebody with the problem looks?73.774.045 to 92
TrustedDoes what they find hold up?73.270.042 to 95
ChosenIs there enough to decide on?68.068.035 to 100

Three averages in the sixties and seventies look like a competent field having a competent day. The averages are the least interesting thing here. What the six underlying signals show is a field that is uniformly excellent at one thing and uniformly poor at another, with very little in between.

Access was never the problem

The single highest signal in the whole study is retrievability, at 97.0 out of 100. Fifty of the 54 sites serve a sitemap. Fifty two serve more than three hundred words of real text without running any JavaScript. Fifty three declare a canonical URL consistently across the pages we read.

Then there is the finding that made us rewrite our own tool. Ten sites appeared to block crawlers, which would have been the headline. It was wrong. Our first pass searched for a blocking rule anywhere in the file without checking which agent the rule applied to. Read properly, group by group, the number of sites in this sample that block every crawler is zero, and the number that block the crawlers behind the major assistants is one.

Ten sites do block some named agent. Eight of them block only a Huawei crawler that feeds no assistant at all, and a ninth adds Common Crawl and ByteDance. That leaves one practice in 54 that closed the door on ChatGPT, Claude and Gemini. Thirty of the 54 name AI agents in their robots file, and twenty of those name thirteen of them and block none. That is a hosting platform shipping a panel with the switch left off. It is a default, not a decision.

Fifty three of 54 therapy practices are wide open to the crawlers that feed ChatGPT, Claude, Gemini and Perplexity. Whatever is keeping these practices out of an assistant's answer, it is not a locked door.

The lowest signal is whether a page can be quoted

Against a retrievability of 97.0, quotability sits at 50.5. It is the widest gap in the study and the whole finding lives inside it.

Quotability does not measure whether anybody quotes you. It measures whether a page has the shape of something quotable: a heading that states a question somebody actually asks, an answer in the first sentence under it rather than the fourth paragraph, a list or a table where a list or table belongs, a visible date, a claim specific enough to survive being lifted out of its page. We wrote about what makes a page citable before we had a number for it. The number is worse than we expected.

These sites are, overwhelmingly, well written brochures. A model can read every word and still find nothing it can hand to somebody as an answer.

What a stranger cannot find out about your practice

The second half of the story is in the two decision signals. Here is what the sample serves, counted over all 54 sites.

What the page tells a strangerSites that say itSites that do not
What the service covers459
A price, of any kind3618
A named human being in the site data2529
How the work actually proceeds3024
Who this is not a good fit for2925
A contact form present in the served HTML2826
Any external profile linked as site data2529
An about page a crawler can reach2331
How long anything takes1737

Read the bottom half of that table again. More than half of these sites do not link a single external profile in their structured data, which means nothing outside the site confirms that the practice exists. More than half name no person at all in their site data, in a field where the entire purchase is a person. That last one carries a caveat we would rather state than bury: it asks whether a person is named in the structured data of the eight pages we read, and one practice in our own sample published three therapist biographies that carry no such data. Read it as a floor, not a census. Half publish no contact form that exists before JavaScript runs, so an assistant asked how to reach them has to answer from an address it may not find.

We report that last one with some humility. We measured the same defect on our own booking page five days ago and it is still there.

Three things our instrument got wrong before it got them right

Every number above survived a round of us being wrong, and the corrections are more useful than the results.

We rate limited the sites and blamed them. Three practices came back unreadable. They were not. Fetching eight pages in a burst tripped their rate limiting, and the error our own load produced was about to be published as a fact about them.

We turned eleven sites into one page sites. Our sitemap reader did not follow a sitemap index and choked on a common way of wrapping URLs in XML. Eleven practices looked like they served a single crawlable page. One of them publishes 187. Had we not checked, the headline would have been that a fifth of this field is invisible past the front door, and it would have been our bug.

We read a rule without reading who it applied to. That is the ten sites blocking everything that turned out to be zero, described above. We found it a fourth time on the day this went out, and the fourth time is the one worth copying. The correction had been written into a new script, and the older script that actually produces these averages still had the original reading, so ten practices were still being docked for a rule that was never aimed at them. Retrievability was about to publish at 92.6 when it is 97.0. Fixing a bug is not the same as fixing it everywhere the number comes from. The lesson is one line: before reporting a defect, rule out that the instrument is the defect.

If this is your practice, the order matters

Nothing here calls for a rebuild. The expensive half of the problem, being reachable, is already solved for 53 of these 54 practices, and it is the half that usually costs money.

What is missing is cheap and specific. Put a price or a range on a page. Name the clinicians in the site data, not only in a photo caption. Say who you are not the right therapist for, which is the single most persuasive sentence most practice sites do not contain. Link the profiles that already exist elsewhere so something outside your own domain confirms you are real. Then write one page that answers a question in its first sentence instead of its fourth paragraph.

That is roughly a day of work, and it moves the two scores that are actually low.

Questions

Does a low score mean the practice is doing badly?
No. Every site here belongs to a working practice with clients, and most of them read well to a human being. The instrument measures what a machine can retrieve, verify and act on, which is a narrow question and not a judgement of the therapy or the design.

Why publish averages instead of naming each site?
Because a ranking of named third parties who never agreed to be measured is a different publication with different obligations. The method, the sample frame and the exclusions are public so anybody can rerun this, including on their own domain.

Is this sample representative of therapy practices in the United States?
No, and we do not offer it as one. It is a convenience sample of 54 sites drawn from the first page of one repeated search in ten metros. It describes practices that already surface in search, which is a narrower and more flattering group than the field as a whole.

Six of your twelve signals were not measured. Does that make the scores soft?
It makes them partial, and we would rather say so. Question coverage, index presence, evidence for claims, outside corroboration and recommendation share cannot be read off served HTML, so each score covers only the weight we actually observed, and the report prints the missing weight beside it rather than filling it in.

Would blocking AI crawlers protect a practice?
One site in this sample does it, and that is a legitimate choice for reasons that have nothing to do with visibility. It is worth separating the two decisions. Blocking is about whether you want your words used for training and answers. Being quotable is about what happens when somebody has already decided to ask about you.

The short version

Fifty three of 54 therapy practices are perfectly readable by the machines that now answer questions about therapists, and roughly half of them serve no price, no named clinician, no reachable about page and nothing outside their own domain that confirms they exist. The door is open. There is very little inside worth carrying back.

We are running this instrument over other niches next, and we will publish those the same way, with the method and the mistakes attached. If you want to know what your own site returns before we get to your field, that is a reasonable thing to bring to a call. What the work covers and what it costs is on the services page, with more detail on the FAQ page. The call is thirty minutes.

Previous
Previous

How to Build an AI Visibility Index for Your Own Niche

Next
Next

What to Put in an llms.txt File, and What to Leave Out