Keyword cannibalization: proving a conflict without counting pages
Two pages on one query prove nothing. Across 49 published detection methods, not one separates an engine that is hesitating from an engine that has decided. Here is the protocol that does, by hand, and what our own thresholds are worth.
The NessFlow team (Product engineering, NessFlow) · · 14 min read
A real product screen, rendered on a fictional demo dataset: the figures shown belong to no client.
On 26 August 2026 we pulled the results for ten cannibalization queries by hand, five on google.fr and five on google.com. All ten carried a generated overview. Here are two of them, produced the same day, by the same engine.
On keyword cannibalization myth:
Keyword cannibalization is partly a myth because having multiple pages rank for the same term is not an automatic penalty.
On deux pages même mot clé, in French:
Baisse de classement : Google hésite entre les deux pages et pénalise les deux.
[Ranking drop: Google hesitates between the two pages and penalises both.]
Not one of the ten overviews cites a Google source. They cite SEO blogs, agency sites, one video and one forum thread. And the second claim, the penalty one, appears in no Google documentation and in no statement by any Google spokesperson, ever.
That is the subject in one picture. Fifteen years of discussion, hundreds of published pages, and the most widely read answer now contradicts itself from one query to the next.
This article does not replay that argument. It asks a measurement question: what proves that cannibalization is happening, and what is only a suspicion?
What the engines actually say
Two statements matter, because they are primary, dated, and they say different things.
Google, John Mueller, 20 September 2025, answering a public question about detecting cannibalization:
Search Console shows data for when pages were actually shown, it's not a theoretical measurement. Assuming you're looking for pages ranking for the same query, you'd see that only if they were actually shown. (IMO it's not really "cannibalization" if it's theoretical.)
And in the same thread:
If you have 3 different pages appearing in the same search result, that doesn't seem problematic to me just because it's "more than 1".
Microsoft, Bing Webmaster blog, 19 December 2025, signed by two principal product managers:
Competition in search is tough enough without unintentionally competing against your own content.
When several URLs contain the same content, signals such as clicks, links, impressions, and engagement are often diluted.
If your signals are unclear or inconsistent, the version that ranks may not be the one you intended.
These two positions do not contradict each other. They are about two different things.
Microsoft describes a mechanism: when several URLs cover the same topic, signals split, the engine picks a representative, and that representative may not be the one you wanted. The same post states explicitly that duplicate content does not trigger a penalty on its own.
Google does not dispute the mechanism. Its documentation describes it in other words: deduplication, clustering, canonicalization, signal consolidation. What Mueller rejects is a metric: counting the pages that could rank for a query, and raising an alert because there is more than one.
So the most defensible official position is neither "it does not exist" nor "it is serious". It is: not a penalty, an arbitration. The real risk is not being punished, it is losing control of which page represents your site.
That arbitration has a direct consequence for detection, because it invalidates the most common method.
Counting pages proves nothing
The detection rule almost everyone applies fits in one line: two URLs from the site appear for the same query, therefore cannibalization.
We went through 49 published SEO detection methods, English and French, from 2016 to 2026. What that corpus gives:
- 17 of them publish no numeric threshold at all. None. The method is prose, and the reader guesses.
- 31 say nothing about their own false positives. The other 18 address it qualitatively. Not one source, in either language, publishes a false positive rate, a confidence interval, or a validation study for its threshold.
- No (signal, threshold) pair is reproduced identically by two independent sources, with two exceptions: the two URL threshold, which is a definition rather than a threshold, and a 10 % click share that travels from one source to another through reuse of the same script.
- The most scattered parameter is ranking depth: published values run from top 3 to top 100, a factor of 33 between the bounds.
- The recommended observation window runs from 28 days to 16 months.
- The stated delay before a fix shows an effect runs from two days to six months, a ratio of roughly 1 to 45, and none of these sources notes that it contradicts its neighbour.
One figure comes close to a reliability measurement. A tool vendor sampled the cases its own "two or more URLs on the query" rule surfaced on one domain, and counted how many warranted action: one in eighty.
That number does not say cannibalization is rare. It says the counting criterion does not measure it. A normal, well built site constantly produces situations where several of its pages touch one query without any of them harming another: a category page and a product page, a guide and its edge case, an article and its update. Counting pages counts all of that alongside the real conflicts.
The discriminator nobody instruments
There is a distinction that separates signal from noise, and it is temporal.
Coexistence: both URLs are shown together, on the same results page, on the same day. The engine decided both earned their place. That is not a conflict, it is a double listing, and it is usually what you want.
Alternation: the engine changes its elected page from one reading to the next. One day it is one URL, the following week the other, and neither settles. Here the engine is arbitrating, and arbitrating without conviction. This is the signature Microsoft describes: the version that ranks is not the one you intended, and it keeps changing.
This distinction is identified as the central discriminator of the subject, and no published threshold instruments it. All 49 methods read an aggregate, a Search Console export over thirty or ninety days, or a weekly rank tracking file, in which two URLs may never have been shown on the same day and still sit side by side in the table.
That is exactly Mueller's objection. A monthly aggregate shows two pages on a query. It does not show whether they were there together or in turn. Both situations produce the same row, and they call for opposite decisions.
Hence the rule this article defends, in one sentence: an export does not prove cannibalization, a series can.
What Search Console measures, and what it does not
The protocol below runs on Search Console, so it is worth being precise about what its numbers are worth. Four properties of that tool bear directly on this subject, and none of the methods we reviewed accounts for them.
1. A query's position is your best URL's position, not an average across your URLs. At query grain, the tool keeps the position of the site's highest ranking page. The second URL of a competing pair therefore never appears in that row. A pair whose better page sits steadily at position 4 shows a clean position 4, whatever happens to the other one.
2. Filtering by query removes anonymized queries from the total. Search Console hides queries issued too rarely. They stay in the chart totals, unless a query filter is applied, which is precisely what any cannibalization analysis does. A share computed against a filtered total and a share computed against an unfiltered total are not the same quantity.
The proportion at stake is not marginal. The widest third party measurement we found, across 887,534 properties and 22 billion clicks, puts anonymized queries at 46.77 % of clicks for one month in April 2025. Two caveats, and they matter: that figure is about clicks, and no study, official or third party, measures the share of anonymized impressions. A detector built on impression share therefore works on a denominator that is incomplete by an amount nobody has quantified.
3. The denominator changed nature twice in eighteen months. Impressions from generative answer surfaces entered the "Web" type on 16 June 2025. And impressions were mislogged from 13 May 2025 to 27 April 2026, an error Google acknowledged and corrected, noting that it affected impressions, click through rate and average position, but not clicks. An impression share computed over a window that straddles those dates does not measure the same quantity end to end. None of the 49 methods mentions this.
4. An accepted canonical moves impressions and clicks to the canonical URL in Search Console. This is documented, and it is the most dangerous artefact in the subject: after a consolidation, the surviving page mechanically shows the numbers of both, producing a flattering before and after without a single extra visitor. None of the methods, and none of the remedy cases we read, warns about this artefact. It is on its own enough to manufacture an apparent success.
The protocol
It runs by hand, with a Search Console export and a spreadsheet. Five steps.
Step 1. Pull query and page pairs. In the performance report, select a query, then the Pages tab: you get the URLs actually shown for that query. Through the API, request the query and page dimensions together. Take a window of at least ninety days for volume, but cut it into periods, because the rest depends on it.
Step 2. Drop queries that are too small. A query worth fewer than 30 impressions over the window supports no conclusion: the split between two pages there is noise. We use 30. That threshold is discussed below.
Step 3. Find the impression split. For each remaining query, compute each page's share of impressions. Keep as competing those pages holding at least 15 % of the query's impressions, and keep the query only if at least 2 pages clear that bar. A page at 96 % against a page at 2 % is not a conflict, it is a main page and a marginal one.
Step 4, the one that decides. Separate alternation from coexistence. Cut the window into periods, one week works, and for each period look at which URL was served.
- Both URLs present in the same period, regularly: coexistence. The engine ruled in favour of both. Leave it alone.
- The served URL changes from period to period, and rarely both at once: alternation. That is a conflict.
- One URL dominates and the other appears intermittently without ever settling: an emerging conflict, worth watching before acting.
This step is what separates a diagnosis from a suspicion, and it is the one no published method quantifies.
It is also the only one that requires having kept something. A position, unlike a page, cannot be reconstructed after the fact: the URL an engine served last Tuesday for a given query exists nowhere if nobody wrote it down last Tuesday. Search Console gives you a trace for the windows it retains, at filtered query grain and with the caveats above; a rank tracking file gives it to you day by day, provided somebody recorded it while it was true. The day you notice alternation, the history that proves it is already behind you, or it does not exist.
Step 5. Check that both pages target the same intent. Alternation can also come from an engine hesitating between two different answers to an ambiguous query. Read both pages. If they answer two distinct questions, the problem is not cannibalization, it is query targeting. Our guide on semantic coverage separates what can be counted on a page from what can only be judged, and that is the reading which settles this one.
Our thresholds, and what they are worth
We publish the three our product applies, because a market where seventeen methods out of forty nine publish no number at all does not need an eighteenth.
| Parameter | Our value | What it excludes |
|---|---|---|
| Minimum impressions for the query | 30 over the window | Queries where the split is noise |
| Minimum impression share for a page | 15 % | Marginal pages mistaken for competitors |
| Minimum competing pages | 2 | The definition of the case |
What these thresholds are not.
They are not validated. We do not publish a false positive rate, and on that point we are in exactly the position of the thirty one methods out of forty nine this article just named. Nor have we recalibrated them after the two series breaks described above, and an impression share does not mean the same thing before and after 27 April 2026.
They are not universal either. 30 impressions is reasonable on a site of a few hundred pages and absurd on a site of a million, where it would surface tens of thousands of cases. No published threshold, ours included, is calibrated to site size or query volume.
Take them as a documented starting point, not a reference. The only setting that counts is the one you check against your own data.
Deciding, and what the evidence supports
Once a conflict is established, the SEO literature offers six options: merge with a permanent redirect, set a canonical, differentiate the intent, deindex the weaker page, consolidate through internal links, or do nothing.
Here is the actual state of the evidence behind those options, after going through the published cases.
None has ever been tested with a control group. Zero. Across all six, there is no controlled test of a page merge, an intent differentiation or a deindexation. Nor is there any direct comparison of two remedies on comparable groups: not redirect against canonical, not merge against differentiation.
The most cited risk is quantified nowhere. "Merging loses your long tail" is the standard objection to the most recommended option. We found no primary source publishing the number of distinct queries before and after a merge, except one, and it runs the other way (67 to 106 queries per day, rising).
The failure mechanism that is actually documented lies elsewhere. It is not long tail loss, it is the redirect to a non equivalent page, reclassified as a soft 404 and then deindexed. Four independent cases describe it, including one retailer that went from 40 to 70 clicks a day to zero, deindexed within two months after a replatforming with more than 15,000 badly redirected URLs. The second best established mechanism is that a canonical is a hint and not a directive: several cases document thousands of canonicalized URLs that Google ruled on differently.
And consolidation does not always win. Across eleven merge cases on record, eight publish positive percentages only. The same analyst who documents successful consolidations also published one where a visibility index went from 127 to 113.11, roughly 11 % net loss. The literature keeps only the first kind.
What that state of evidence allows us to write:
- Merge when both pages say the same thing, and redirect to a genuinely equivalent page. Equivalence is the condition, not a detail: it is where the only well documented failure mechanism lives.
- Differentiate when the two pages serve two adjacent but distinct intents. Each takes its angle, and they link to each other.
- Do nothing on stable coexistence. Two pages settled together on a query do not need repairing.
- Deindex as a last resort only. It is the option that loses the content without recovering it elsewhere, and it is also the one with no isolated, quantified case behind it.
Note what this list does not say. It does not say "merge first". The order depends on what the measurement showed, and the default move, when the measurement does not settle it, is to do nothing.
Before acting, keep enough to come back. Nobody publishes a case where two merged URLs were later restored separately with a measurement of recovered traffic. That gap is a warning: keep the absorbed page's content, its query list, its referring domains and its positions as of the merge date. That is what separates fixing from observing.
Measuring the effect without lying to yourself
One last trap, and it is the one that invalidates the most reports.
After a merge, the surviving page mechanically inherits the impressions and clicks of the absorbed page in Search Console, because that is how the tool attributes traffic for a canonicalized URL. The surviving page's chart goes up, always. That rise proves nothing, and it happens even when total site traffic has fallen.
Three precautions are enough:
- Measure at query grain, not page grain. The question is not "did the surviving page go up", it is "did the query gain clicks, and is the served URL now stable".
- Sum both URLs before the merge, and compare that sum to the single URL after. Comparing one page to one page compares one thing to two.
- Wait for the served URL to hold steady across several consecutive readings before concluding. It is the same signal that established the diagnosis, and it is the only one that proves the fix.
That is everything this subject asks for, and it is more than most SEO audits give it: a series rather than an export, a threshold you own rather than a head count, and a default move that is to do nothing.