Because on a massive, generic topic — what a screwdriver is, what a GPS tracker is — you're competing against knowledge the model already carries from training, largely from sources like Wikipedia. On a specific topic in your niche, on the other hand, you're in the only zone where what you publish actually decides the answer: a University of Toronto study measured that the match between what an AI answers and a careful, pairwise comparison drops from 91% on popular topics to 56% on specific ones.
The gap isn't random, and it isn't a flaw in the model. It's proof that, faced with a topic it already knows by heart, the AI barely needs what it finds when it searches — it uses that evidence to confirm what it already had stored from training. Faced with a topic it doesn't know by heart, it has nothing to confirm against: it depends entirely on what it actually finds.
This has a direct consequence for any small business competing on a massive topic. If your content explains something the AI already knows — what a screwdriver is, what a GPS tracker is — you're competing against knowledge the model already carries, largely from sources like Wikipedia, which dominates precisely the most general topics. If your content answers something specific, you're in the only zone where what you publish actually decides the answer.
Because the model already carries a knowledge hierarchy built during training, and what it finds when it searches only confirms it — it doesn't build it. The study proved this by putting the model's answers through three stress tests: shuffling the order in which the retrieved evidence is presented, restricting the model to only that evidence without leaning on what it already knew, and swapping entity names for others. For questions about popular entities — the study's own example is "best SUVs to buy in 2025" — none of the three tests moved the result much: the final ranking changed by an average of just 2.3 positions overall, even after fully reshuffling the available evidence.
The authors' own interpretation is direct: for popular topics, the evidence the model finds when searching works as confirmation, not discovery. The model already has an internal hierarchy built during training, and what it reads while searching barely adjusts it.
For niche entities — the study's own example is "best family law firms in Toronto" — the behavior flips. The same three stress tests moved the ranking almost twice as much (4.15 positions on average) and changed the final result far more sharply when entities were swapped. The model, in the study's own words, enters a "knowledge-seeking mode": since it has no prior hierarchy built for that specific topic, it constructs the answer from what it finds instead of confirming it.
The study doesn't propose a business criterion — it measures how a model behaves across different query types. From that reading, Nostos built its own Topical specificity criterion, with its own methodology: 40% real content specificity, 30% technical markup and 30% verifiable data density. It isn't a formula the paper recommends; it's the operational translation of a mechanism the paper does demonstrate.
For a neighborhood hardware store, this changes what's worth writing entirely. An article about "what is a drill" lands in the zone where the model already has the answer solved from memory — no matter how well the text is written, it's competing against a knowledge hierarchy that already exists. An article about "which drill bit to use on ceramic tile without cracking it" lands in the zone where, according to the mechanism measured in the study, the model needs real evidence to build its answer — and that evidence can be the hardware store's own page, if it's specific and well written.
Nostos measures its own site with the same engine it uses to analyze its clients. On Topical specificity, nostosgeo.app scored 85/100 (measured 07/02/2026) — well above the 43/100 average Nostos found across 39 Argentine e-commerce domains in its own sector study.
The gap isn't a coincidence: nostosgeo.app was built writing about the specifics of its own subject — criteria, papers, how AI engines actually work — instead of competing on generic definitions any model already resolves on its own.
The mechanism the study measured gives a concrete operational rule:
Competing against Wikipedia on a generic topic is a race the model's own mechanism already decided long before the first article gets written. Competing on the specific is the only race where a site's content actually decides who wins.
Get your GEO score across the eight criteria backed by academic research, including your current Topical specificity.
Run the analysis at nostosgeo.app