Technical SEO FAQ · Indexing, Canonicalization & Duplicate Content
How would you identify index bloat?
Do not define index bloat simply as "too many indexed pages."
Compare:
Indexed URLs
against:
valuable, intentional, indexable URL inventory.
Look for:
- internal search pages;
- parameter URLs;
- duplicate pages;
- thin archive pages;
- obsolete pages;
- generated combinations;
- soft 404-like pages.
Then determine whether these URLs are actually appearing in Google's index and whether they create a meaningful search-management problem.
The goal is not the smallest possible index. It is an index containing the pages that should represent the site in Search.
Related questions
- How do you determine why an important page isn't indexed?
- How do you investigate "Crawled - currently not indexed"?
- How do you investigate "Discovered - currently not indexed"?
- How do you validate whether an indexation fix actually worked?
- How would you manage indexation on a large e-commerce website?
- What happens when canonical and noindex signals conflict?
- What is the difference between crawling, rendering, indexing, ranking and serving?
- What signals influence Google's canonical selection?
- When should a page be canonicalized instead of noindexed?
