Does a distinctiveness score prove anything about a generated page
The question comes up every time a spam update lands, and it arrives in the same shape: someone has a pool of generated pages, a checker that reports them as highly distinct from one another, and a hunch that the number is not answering the question they actually have. They are correct to be doubtful, and the reason is easier to demonstrate than to argue about, because both parts of it can be checked on the same group of pages in around a second.
I generate location pages for one service, one page per city, from a skeleton with plenty of synonym slots. My checker says the pages barely overlap, which everyone tells me is what you want. But the search engine's own policy names synonymising as an example of the thing it acts against, so which one am I supposed to believe, my score or their page?
Both, since they are not measuring the same thing. The policy language is about what a batch of pages is for and how little it adds for the reader; the method is clearly not the violation, and the same sentence says so whether a individual or a machine wrote it. A similarity score, meanwhile, measures variation density over short stretches of words. A pool can move that score as far as it goes and still express one thing in many costumes, which is the phenomenon the policy example describes rather than the opposite of it. The first thing to establish is not how distinct the pages look but where their differences came from.
There is a test that reveals where the differences came from, and it requires one pass. Replace every value that arrived from data, the town name included, with a placeholder, and check the pool again. If what remains is many pages that are still distinct, the template manufactured them and the reader receives one page every time. If what remains collapses to a single page repeated, the data was holding the difference, and the pages are worth as much as that data behind them. That second part the test cannot resolve for you, and no measurement will.
Why the score moves the wrong way when the pages contain facts
Construct the same pool twice. Once with synonyms and nothing else varying, and once with no synonyms at all but a handful of real values tied per page, the sort of thing a listing would hold. The synonymised pool scores as hardly overlapping. The pool carrying values scores worse, because facts repeat across pages in the same sentence patterns and short word runs survive. By the score, the pages built to say something particular look like the questionable ones, and the pages stating every town the same thing in different adjectives look clean. That inversion is the whole answer to the original question: the score cannot distinguish a pool that says something from one that merely appears varied, so a good number is not evidence of anything the policy cares about and a poor one is not an accusation. It measures the surface, which is exactly the part a synonymiser is built to move.
The write-up, with both pools measured and the file that replicates them: spintax deliverability
Run the deletion pass before the publication decision rather than after, because it resolves a question the dashboard never will: what does a reader find on the second page that they did not see on the first. If the answer is a different arrangement of the same sentence, more variation will not fix it, and the number that says otherwise is measuring the wrong surface.