Guide

robots.txt, noindex, sitemap and canonical are not interchangeable

Think of them as crawl control, index control, discovery help, and canonicalization. Similar neighborhood, different jobs.

robots.txt controls crawling, not secrecy

It tells cooperative crawlers where not to fetch. It is not authentication, and blocking a URL does not guarantee the URL disappears from search.

noindex needs to be seen to work

If you block a page from crawling and also expect a page-level noindex to be read, you can create a contradiction.

Canonical points to the preferred version

It is a signal for duplicate or near-duplicate URLs, not a replacement for every redirect.

Sitemaps help discovery

Keep them focused on canonical, indexable URLs you actually want search engines to find.

Check one representative URL from end to end

Ask: should it be crawled, should it be indexed, what is its canonical URL, and should that canonical URL be in the sitemap?

If those answers point in different directions, fix the contradiction before scaling the setting across the site.

Related tools and references

Related topics