Duplicate content is when identical or near-identical content appears on multiple URLs — either within your own site or across different domains. It doesn’t usually trigger a “penalty,” but it does force search engines to choose which version to rank, which can dilute ranking signals and waste crawl budget.
Common internal causes: URL parameters (tracking codes, filters, session IDs) creating multiple versions of the same page, www vs. non-www or HTTP vs. HTTPS both being accessible, printer-friendly page versions, and paginated content that isn’t handled correctly.
Common external causes: Content syndicated to partner sites without proper attribution, or scrapers republishing your content elsewhere.
How to diagnose it: Use Search Console’s Page indexing report to check for “Duplicate, Google chose different canonical than user” — this flags cases where Google picked a different page than you intended. Tools like Screaming Frog can also crawl your site and flag pages with identical or near-identical title tags and content.
How to fix internal duplication: 1. Set canonical tags pointing to the preferred version of each page. 2. Use 301 redirects to consolidate genuinely duplicate URLs into one. 3. Standardize on one version — www or non-www, trailing slash or not — and redirect all variants to it. 4. Use noindex for low-value duplicate pages like print versions.
How to handle external duplication: For syndicated content, ask partners to include a canonical tag pointing back to your original, or a link back to the source article.
Is duplicate content a penalty?
No. Google has said many times that normal duplicate content, such as product variations or printer-friendly pages, does not lead to a penalty. Google simply picks one version to show and filters out the rest. The real costs are softer: ranking signals split across several URLs, Google may pick the "wrong" version to rank, and crawlers spend time on copies instead of new pages. The exception is deliberately copying other sites' content at scale, which falls under Google's spam policies.
Near-duplicate content: the issue most sites miss
Exact copies are easy to spot. The harder problem is near-duplicates: pages that are 80–90% the same. Common examples are:
- City or location pages where only the city name changes ("SEO services in Delhi", "SEO services in Noida").
- Product pages for sizes or colours that share one description.
- Several blog posts that answer the same question with slightly different titles.
For location pages, add genuinely local content: local case examples, addresses, service areas, team members, local FAQs. For similar blog posts, merge them into the strongest one and 301-redirect the others. This also fixes keyword cannibalization.
Which fix to use when
- The duplicate URL has no reason to exist: use a 301 redirect to the main version.
- Users need both URLs (filters, sorting, tracking parameters): keep both and add a canonical tag on the duplicate. See canonical tags explained.
- The page should never appear in search (internal search results, thin tag pages): use noindex, and keep it crawlable so Google can see the tag.
- Two pages cover the same topic: merge the content, keep the better URL, and redirect the other.
A step-by-step duplicate content audit
- In Search Console, open Pages (the Page indexing report) and review the statuses "Duplicate without user-selected canonical" and "Duplicate, Google chose different canonical than user".
- Crawl the site and sort by title and meta description. Identical titles are a quick signal of duplicate or near-duplicate pages.
- Use the crawler's near-duplicate feature to group pages by content similarity.
- Run a site: search with a distinctive sentence from your content in quotes to find copies on other domains.
- Decide on one action per group (redirect, canonical, noindex or rewrite) and record it in a sheet.
- After the fixes, re-crawl and check the Search Console report again after a few weeks.
Preventing duplicate content in the future
- Choose one URL format (HTTPS, with or without www, with or without trailing slash) and enforce it with redirects.
- Add self-referencing canonical tags to every indexable page template.
- Keep tracking parameters (UTM codes) out of internal links. Use them only on external campaign links.
- Before writing a new post, search your own site to check the topic isn't already covered.
Duplicate content rarely causes a sudden crash, but cleaning it up is one of the most reliable ways to concentrate ranking signals on the pages that matter.
