Back to all posts
tutorials 9 min read

Too Many Thin, Competing Pages? How to Decide Which Ones to Prune

Serap Gündoğdu ·
Too Many Thin, Competing Pages? How to Decide Which Ones to Prune

There is a version of the indexing problem that no one warns you about at the start. You keep publishing, the site grows, and one day you have four blog posts, two category pages, and an old landing page all reaching for the same query. None of them ranks well. They are not fighting Google. They are fighting each other. This is the opposite of “why won’t Google index my page”, and if that is your problem, our guide on pages stuck as “discovered, currently not indexed” is the one you want. This article is for the site that has too many pages, not too few, and needs to decide which ones to cut.

Pruning feels wrong the first time you do it. You wrote those pages. Deleting them looks like throwing away work. But an oversized, overlapping content library is not an asset, it is drag: it splits your ranking signals, wastes crawl attention on pages that will never win, and buries your best work under near-duplicates. The goal is not to have more pages. It is to have the right pages, each one clearly the best answer to a distinct query.

What Keyword Cannibalization Actually Is

Keyword cannibalization is when two or more of your own pages target the same search intent, so search engines cannot tell which one to rank. Instead of one strong page, you have two or three weak ones dividing the clicks, the internal links, and the authority between them. Google may swap which page it shows, rank the wrong one, or hold both back because neither is a clear winner.

It helps to say what it is not. Two pages that happen to share a keyword are not automatically cannibalizing each other. “Best running shoes” and “how to clean running shoes” both contain “running shoes”, but they answer completely different questions, so they can coexist happily. Cannibalization is about shared intent, not shared words. The test is simple: if a searcher typing one query would be equally served by either page, those pages compete. If each page answers a different question, they do not.

Cannibalization, Thin Content, and Duplication Are One Decision

In practice these three problems tangle together, and it is easier to treat them as a single audit than to chase each one separately.

  • Cannibalization is two pages competing for the same intent.
  • Thin content is a page that does not say enough distinct or useful to earn a ranking at all.
  • Duplication is a page whose content substantially repeats another, yours or someone else’s.

A single weak page is often all three at once: a thin article that duplicates the angle of a stronger post and competes with it for the same query. That is why the fix is one workflow, not three. You find the weak and overlapping pages, then you decide, page by page, what each one deserves.

Find the Candidates with a Crawl

Search Console will show you queries where more than one URL gets impressions, which is a useful hint, but it does not show you the structural picture: which pages overlap in topic, which are thin, which are orphaned, and how they link to each other. For that you need to see your site the way a crawler does.

  1. Group pages by target query and title overlap. Crawl the site and sort by title and H1. Clusters of near-identical titles are your first cannibalization suspects. Two posts both titled around “XML sitemap best practices” are almost certainly competing.
  2. Sort by word count to expose thin pages. The pages at the bottom of the distribution, a few hundred words wrapped in a lot of template, are prime prune candidates. They rarely win and they dilute the rest.
  3. Detect near-duplicate content. Look for pages whose body text closely matches another. These split authority even when their titles differ.
  4. Find the orphans. A page with zero or one internal link pointing to it is a page you have already stopped believing in. If it is also thin or duplicative, that is your answer.
  5. Map the internal links between competitors. When two pages compete, see which one your own site links to more. Your internal linking usually already “votes” for the stronger page.

The output you want is a concrete list: these three posts all target the same intent, this one is thin, these two are near-duplicates, this old page is orphaned. Now you can decide.

The Decision: Delete, Redirect, Noindex, Consolidate, or Canonical

This is the part that trips people up, because the five tools look interchangeable and are not. Match the tool to the situation.

Consolidate (then 301 redirect). This is the default for cannibalization. When two or three pages compete for the same intent, merge the best parts into one strong page, then 301 redirect the losers to the winner. You keep the good content, pass the link equity of the old URLs to the survivor, and give Google one obvious answer. Update the internal links so they point at the winner directly rather than through the redirect.

Delete (410 or 301). For a genuinely worthless page with no useful content and no earned links, removing it is fine. If it has any inbound links or lingering traffic, 301 it to the most relevant surviving page so that equity is not lost. If it is pure junk with nothing pointing at it, a 410 (Gone) tells Google to drop it cleanly.

Noindex. Use this for pages that must exist for users but should not compete in search: thin tag archives, thank-you pages, internal search results, filtered listings. Noindex keeps the page live for people while removing it from the index. Do not noindex a page you actually want to rank; that is a common way to accidentally bury good content.

Canonical. Use rel=canonical when near-duplicate pages both need to stay live at their own URLs, such as a product available in two categories, and you simply want to name the primary. It is a signal, not a command, and it is the wrong tool for pages you want gone. Our guide on canonical tags and duplicate content covers the mechanics and the common mistakes.

Leave it alone. Not every overlap needs surgery. If two pages genuinely serve different intents, the fix is to sharpen each one so the difference is obvious, not to merge them. Pruning too aggressively is its own mistake.

A quick way to choose: if the pages compete for one intent, consolidate. If a page is worthless and unlinked, delete. If it serves users but not search, noindex. If it is a legitimate duplicate that must stay, canonical. If it earns its own keep, leave it and differentiate it.

How to Consolidate Two Competing Pages

Consolidation is the most common move and the easiest to do badly, so it is worth a short playbook.

  1. Pick the winner. Usually the page with more links, more traffic, or the stronger URL. Keep that URL.
  2. Merge the content. Move any unique, useful sections from the losing pages into the winner so nothing valuable is lost.
  3. 301 the losers to the winner. Every old URL redirects to the surviving page.
  4. Update internal links. Point your internal links at the new canonical URL directly. Leaving them aimed at redirects wastes a little authority and slows crawling. Our internal linking guide goes deeper.
  5. Refresh the sitemap. Remove the retired URLs so your sitemap lists only live, canonical pages.

Common Mistakes

  • Noindexing when you meant to redirect. A noindexed page still exists and still splits internal links; it just cannot rank. If the page should be gone, 301 it.
  • Redirecting everything to the homepage. A 301 only passes meaningful equity when the target is genuinely relevant. Homepage redirects for unrelated pages read as soft 404s.
  • Pruning by word count alone. A short page that perfectly answers a narrow query is not thin. Judge by value and intent, not just length.
  • Forgetting the internal links after a merge. Consolidation is not finished until the links point at the survivor.
  • Deleting before checking for backlinks. A page with external links is worth redirecting, not erasing.

How Seodisias Helps

The hard part of pruning is not the decision framework, it is getting an honest map of your own site before you decide. Seodisias is a free, cross-platform desktop crawler that walks your entire site the way a search engine does and surfaces exactly the patterns this workflow needs: clusters of near-identical titles competing for the same intent, thin pages sorted to the bottom by word count, near-duplicate bodies, and orphaned pages with no internal links pointing in. Instead of guessing which of your pages compete, you get a concrete list to act on. No account, no URL limit, and your crawl data never leaves your machine.

The Bottom Line

A big content library is not the same as a strong one. When several pages chase the same query, they take turns losing, and your best work pays the price. The move is not to publish more, it is to prune with intent: crawl your site to find the thin, duplicate, and competing pages, then match each one to the right tool, consolidate the competitors, delete the dead weight, noindex what belongs to users but not to search, and canonical the true duplicates. Fewer, sharper pages beat a crowd of near-duplicates every time.

Want to see which of your pages are thin, duplicated, or competing for the same query? Crawl your site free with Seodisias.