Why a Perfect Consistency Score Is a Warning Sign

Consistency is the signal almost nobody fails. Fifty one of the 54 therapy websites we measured scored 85 or better, and 17 scored a perfect 100. Then we split the cohort by that score and it ran backwards: the flawless sites were lower on all five of the other signals we measured, and lower on the index overall.

What we actually measure when we say consistency

The word covers a lot of ground in marketing writing, so here is the narrow version we score. Consistency asks whether a site contradicts itself in the places a machine reads. Not whether the brand feels coherent. Whether the served HTML says two different things.

Three checks, and they start at 100 and come down from there. We deduct 15 points if any page title drops the brand name that the rest of the titles carry. We deduct 25 if the word placeholder is sitting in text a visitor can read, which is the fingerprint of a template that was filled in halfway. And we count the distinct prices published across the sampled pages, which we report as data and do not penalise at all.

That last decision was a correction to our own tooling. An earlier version subtracted 20 points from any site publishing five or more price points, and that punishes a practice with seven services for having seven services. The real contradiction is the same item carrying two different prices, and this method cannot tell those apart from the outside, so it does not claim to. The method piece covers the other four times our instrument produced a confident wrong number.

Fifty one sites out of 54 pass it

Here is the whole distribution. Seventeen sites at 100, thirty four at 85, one at 75, two at 60. Nothing below 60 in the entire cohort.

What loses pointsDeductionSites affected
A page title that drops the brand the other titles carry1536 of 54
The word placeholder visible in served text253 of 54
Several different prices published across pages0, reported as data34 publish at least one price

Compare that to the rest of the index. Citability averages 50.5 across the same 54 sites. Verifiable identity, 57.8. Decision served, 60.5. Consistency sits at 88.6, which makes it the second highest of the six signals and the highest one that anybody has to work at, since retrievability at 97 is mostly the hosting platform doing its job.

A signal that 94 percent of a niche passes cannot separate anyone inside that niche. That was the finding we expected to write up, and it would have been a short and slightly boring piece. Then we split the cohort.

Then we split the cohort, and the score ran backwards

We grouped the 54 sites by their consistency score and averaged everything else. We expected the two groups to look roughly the same, because a signal nobody fails should not predict anything in either direction.

Signal17 sites scoring 10034 sites scoring 85
Retrievability92.998.8
Citability38.456.9
Verifiable identity48.262.5
Decision served57.162.9
Friction to contact68.277.5
Overall index score67.573.9

Every row goes the wrong way. The sites that scored perfectly on consistency came out 6.4 points lower on the index than the sites that lost 15 points to a stray page title, and the gap is widest on citability, at 18.5 points.

Our first move was to suspect the measurement rather than the finding, which is the habit that piece eight is mostly about. Three of the 17 perfect scorers had fewer than eight pages read, so we dropped them and ran it again on the 50 sites where the full sample came back. The gap survived: 63.2 against 71.7 on the five other signals. Smaller, still there, still pointing the same way.

Say the size of this honestly. It is 17 sites against 34, in one cohort, in one niche, with no significance test run on it. That is a pattern worth explaining, not a law worth quoting.

Why a thin site cannot contradict itself

The mechanism falls out of the deduction, once you look at what it takes to trigger it.

Losing those 15 points requires having a page whose title says something the other titles do not. A site with eight distinct pages, each about a specific problem, each titled for that problem, will eventually have a title that carries the topic and not the practice name. A site with four pages titled Home, About, Services and Contact, each with the brand pasted on the end, cannot possibly trip the check. There is nothing there to be inconsistent with.

Citability is the row that makes this concrete, because it counts the opposite thing: headings, questions, lists, dated content, answers near the top of the page. A site built out of specific pages scores well there and is exposed to the title deduction. A site built out of four template slots scores 100 on consistency and 38 on citability, and those two numbers are the same fact described from two directions.

So the perfect score is not measuring discipline. In this cohort it is mostly measuring how little the site says. That is a hypothesis that fits the data rather than something we proved, and it is the one we would test first on the next niche.

What this instrument cannot see

The uncomfortable part is that consistency is also the signal where our measurement is furthest from the thing it is named after.

Real inconsistency is a fee that says 150 on the services page and 180 on the FAQ. It is an intake process described one way in the about page and another way in the booking flow. It is a therapist listed as accepting a payer that the contact page says is not accepted. None of that is visible to a script reading eight pages, because catching it requires knowing which two statements are supposed to match.

What we can see are two proxies, brand in the title and template leftovers, plus a count of prices we deliberately do not score. Calling that combination consistency is convenient shorthand. It is also the weakest name to score mapping in the whole index, and we would rather write that down than let the 88.6 sit there looking like good news.

Twenty sites publish no price at all

One number from this signal is worth pulling out on its own, even though it costs nobody a point. Twenty of the 54 sites publish no price anywhere in the sampled pages. No fee, no range, no sliding scale.

Those sites cannot be caught contradicting themselves on price, because they never say anything to contradict. Their consistency score is clean by absence. Their decision served score, which is the signal that asks whether a visitor can work out fit and cost without contacting anyone, is where that silence gets paid for, and the cohort average there is 60.5.

This is the pattern the whole piece keeps circling. Publishing less is the cheapest way to look consistent, and every signal that measures usefulness moves in the other direction.

What to do with a signal that runs backwards

Three things, and only the first one is work.

Check the template leftovers, because that one is real and it is fast. Three sites out of 54 are serving the word placeholder to visitors, and there is no argument for leaving it there. Search your own site for it, along with lorem ipsum and any stock copy your platform shipped with. It takes ten minutes and it is the only part of this signal we would ask anyone to fix on a deadline.

Do not chase the other 15 points. If your consistency score is 85 because a page title carries a topic rather than your practice name, that title is very likely doing its job. Rewriting eight titles to include the brand will lift one signal out of six by 15 points and will not touch citability, identity or whether a visitor can decide. Spending an afternoon there is spending it on the scoreboard.

And treat a perfect score as a prompt rather than a result. If you scored 100, the question is not what you did right. It is whether there is enough on the site to be inconsistent about. That question is answered by the citability and decision served numbers sitting next to it, which is the argument for scoring the three axes separately instead of blending them, and it is the whole reason the method is built the way it is.

Questions

Does a perfect consistency score mean my site is bad?
No. It means the score on its own tells you almost nothing, so read it next to citability and decision served. In this cohort the sites at 100 averaged 38.4 on citability, and that is the number that carries the information.

Should I put my brand name in every page title?
Only if it fits. Our check looks for the brand somewhere in the title rather than at the end, and dropping it on one page costs 15 points on one signal of six. A title that describes what the page is about is worth more than the 15 points.

Why do you report prices without scoring them?
Because publishing seven prices is what a practice with seven services looks like, not a contradiction. The contradiction we would want to catch is the same service priced twice, and reading eight pages from the outside cannot distinguish the two, so we report the count and leave the judgement to the reader.

Is this finding strong enough to act on?
Act on the template leftovers. Treat the inversion as a reason not to spend time on the remaining 15 points, which is a decision about where not to work rather than a claim about what will happen if you do. It is 17 sites against 34 in a single niche with no significance test.

How would I check this on my own site?
Read the served HTML of your eight most important pages, not the rendered page. Count how many carry your brand in the title, search all eight for placeholder text, and list every price you find. The method piece has the sampling and exclusion rules that make the result repeatable.

The part that transfers

A signal everybody passes is not a standard, it is a floor, and a floor tells you who fell through rather than who is doing well. The useful move is to check which of your signals separate the field and which of them only separate the broken from everyone else. In this cohort, retrievability at 97 and consistency at 88.6 are floors. Citability at 50.5 is where the field actually spreads out.

Run the same split on your own index and expect at least one signal to run backwards. Ours did, on the axis we would have been happiest to report, which is roughly how these things go. The full cohort results are in the index itself, and we should say plainly that we sell into the niche we measured.

If you want a read on what a machine can currently retrieve, trust and quote from your own site, that is a reasonable thing to bring to a call. What the practice does and what gets reported is on the services page, with more detail on the FAQ page. The call is thirty minutes.

Previous
Previous

How to Build an AI Visibility Index for Your Own Niche

Next
Next

We Measured 54 Therapy Practice Websites, and Access Was Never the Problem