robots.txt · AI crawler reference

robots.txt for AI crawlers: use precise, documented directives

robots.txt is a crawler-access protocol. Use documented user agents and directives, then keep the intended crawler and purpose explicit.

Direct answer

Keep the search question explicit.

Robots.txt can communicate access preferences to documented crawlers, but it is not a universal AI-search switch or a guarantee of crawling, indexing, inclusion, or citation.

Reference guide

Concepts and evidence boundaries.

Name the user agent

Target a documented user agent instead of relying on broad labels or assumptions about every AI system.

Use the protocol for access guidance

Robots.txt tells compliant crawlers which paths they may request; it is not a ranking or source-management control.

Verify directive effects

Test directives against the documented user agent, confirm the file is reachable and valid, and check server logs for the crawler you intended to address. A directive only expresses a preference; verification shows whether compliant crawlers honor it in practice.

Keep outcomes separate

Changes to access preferences do not establish whether a page will be crawled, indexed, included, or cited. Access is a precondition you can inspect; answer behavior is a provider decision you can only observe.

Primary sources

Official guidance reviewed 2026-08-16.

OpenAI SearchBot

OpenAI. Platform guidance can change; recheck the source before acting.

Related guides

Continue with a focused question.

GEO guide

Return to the educational GEO reference hub.

Questions

What this reference does not promise.

Can robots.txt force an AI answer outcome?

No. It is an access-control convention, not an outcome guarantee. Outcomes belong to each provider’s systems and cannot be forced from a directives file.

Have a real research question?

Inspect evidence in Gavix.

Use supported workflows and keep their boundaries visible.

Start free trial