Use each mechanism for its own job
robots.txt is a crawler access directive. It can tell a compliant crawler which paths should not be fetched, but it is not authentication or a hard security boundary. llms.txt is an emerging convention for describing useful site content to language-model tooling; it should not be treated as a standardized access-control mechanism. X-Robots-Tag is an HTTP indexing directive that can express noindex, nofollow, snippet limits, and related instructions.
Decide what you are trying to control
- If the question is whether a crawler may fetch a path, start with robots.txt and the crawler vendor documentation.
- If the goal is to provide a concise content map, llms.txt can be used as optional guidance where appropriate.
- If the question is whether a fetched response should be indexed or shown in snippets, use indexing directives such as X-Robots-Tag or meta robots where supported.
Keep vendor-specific rules reviewable
AI crawler names and product policies can change. Keep each user-agent block separate and document why it exists, then re-check the vendor documentation when your policy changes. Avoid assuming that one file controls training, search indexing, and crawling at the same time.