How to get cited by ChatGPT with crawler setup and clean capsules

Learn how to get cited by ChatGPT by configuring OAI-SearchBot, publishing link-free answer capsules, and tracking referrals with UTM source tags.

Bunny Team · · 7 min read

On this page ↓

How to get cited by ChatGPT

Buyers who research software categories in conversational assistants expect direct answers to their queries. Understanding how to get cited by ChatGPT requires managing technical crawler access, structuring on-page copy for extraction, and tracking incoming visits through specific referral tags.

To learn how to get cited by ChatGPT, ensure OAI-SearchBot can access your site through robots.txt and firewalls, then place link-free answer capsules of 120 to 150 characters directly beneath question headings. Back claims with original data and external consensus, and track inbound visits using the utm_source=chatgpt.com referral parameter in your web analytics.

Before you start: verify crawler eligibility and access rules

Your content cannot appear in ChatGPT search summaries if your server blocks OpenAI automated crawlers. OpenAI's crawler documentation, live on 2026-09-24, lists four separate bots that handle distinct functions across its products:

  • OAI-SearchBot surfaces websites in ChatGPT search results. Opting out of OAI-SearchBot prevents your content from appearing in search answers.
  • OAI-AdsBot inspects ad landing pages for policy compliance and ad relevance without feeding foundation model training.
  • GPTBot crawls content used to train foundation models. Disallowing GPTBot protects your text from model training without affecting search indexing.
  • ChatGPT-User executes user-initiated actions and browsing in ChatGPT and Custom GPTs. Because user requests initiate these visits, robots.txt rules do not govern them, and they do not influence search inclusion.

Site owners can configure permissions for OAI-SearchBot and GPTBot independently. Allowing OAI-SearchBot while disallowing GPTBot lets your site surface in ChatGPT search answers without your pages entering training datasets. OpenAI documentation states that changes made to a robots.txt file take roughly 24 hours to take effect across search systems.

Full disclosure: you are reading this on the blog of bunny.directory, a startup launch directory where founders submit products to earn permanent product pages and dofollow backlinks.

Step 1: Allow OAI-SearchBot through all four access layers

A permissive robots.txt file fails if your edge network or security tools reject crawler requests. OpenAI's advertiser guidance, updated three days prior to 2026-09-24, outlines the four distinct layers where web crawlers fail.

First, your robots.txt file must explicitly allow the search bot while disallowing training bots if you choose:

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

Second, web protection and bot mitigation services must allow OpenAI user agents. Platforms like Cloudflare and Akamai defend against denial-of-service attacks, but their automated rules can mistakenly return 403 Forbidden errors to search bots. Cloudflare has officially verified and allowlisted OAI-AdsBot, but network administrators must confirm that OAI-SearchBot also passes without triggering security rules.

Third, application-level human verification logic must exempt search crawlers. CAPTCHAs, JavaScript challenges, and behavioral anti-bot tests block automated visitors regardless of your robots.txt configuration.

Fourth, perimeter firewalls must accept requests from OpenAI IP address blocks. If your hosting environment restricts inbound traffic to published ranges, add the IP addresses listed at searchbot.json and adsbot.json to your firewall permit list. Large request spikes can also trigger HTTP 429 Too Many Requests rate limits, requiring your infrastructure team to inspect traffic throttles in server logs.

Placing concise answer capsules directly beneath question headings produces the highest citation frequency in ChatGPT. An audit of nearly two million sessions across 15 domains published by Adam Gnuse on Search Engine Land on 19 November 2025 showed that 72.4% of cited blog posts contained an identifiable answer capsule.

Search Engine Land defines an answer capsule as a self-contained explanation of roughly 120 to 150 characters, or about 20 to 25 words, positioned immediately after a question-shaped heading. This length gives sufficient context to answer the user query while remaining short enough for language models to lift into a summary.

The audit demonstrated that adding hyperlinks inside these capsules reduces quotation rates. Language models interpret anchor text as a sign that the definitive answer resides on a different page.

Capsule link configurationShare of cited capsules
No links~91%
Internal links only~5.2%
External links only~3.5%
Both internal and external links<1%

Over nine in ten cited capsules contained zero links. Keep your capsule paragraphs completely free of hyperlinks, reserving your citations and internal navigation for supporting paragraphs placed lower in the section.

Step 3: Back claims with original data and external corroboration

Publishing proprietary metrics and earning third-party mentions gives language models verified facts to synthesize into answers. Search Engine Land reported that 52.2% of cited posts included either original data or branded insights, with 34.3% combining both an answer capsule and proprietary findings.

Original data includes benchmarks, survey results, and proprietary metrics that originate on your page. Owned insights take proven practices and frame them under a branded recommendation, such as an explicit company tip. These structures provide models with an attributable fact rather than general advice.

Independent corroboration establishes the consensus language models look for when answering competitive questions. An analysis published by Quora Business on 11 June 2026 explains that ChatGPT rarely quotes a page in isolation, seeking claims corroborated across multiple authoritative sources. That study found ChatGPT cited brands in only about 0.59% of evaluated responses, compared with roughly 13% for Perplexity and 27% for Grok.

Founders can also structure public documentation for retrieval systems using tools like the llms.txt generator on bunny.directory. The site offers this utility to create clean markdown index files that help automated scrapers parse documentation.

Step 4: Track ChatGPT referral traffic in analytics

OpenAI appends a dedicated referral parameter to outgoing search links so site owners can monitor ChatGPT traffic in web analytics. According to OpenAI's publishers FAQ, inbound clicks from search results carry the tracking string utm_source=chatgpt.com.

This UTM tag allows site owners to create custom segments in Google Analytics to measure sessions, bounce rates, and conversions from ChatGPT citations. Google handles its own search AI features under a different reporting model. Google Search Central documentation, updated 10 December 2025, states that clicks from AI Overviews and AI Mode appear inside the standard Web search type in Google Search Console's Performance report, without a separate referral parameter.

Visibility shifts take time to materialize after updating pages. Quora Business found that traffic changes typically begin four to eight weeks after restructuring priority content, while broad coverage across competitive prompts requires three to six months.

What common mistakes block ChatGPT citations?

Relying on robots.txt disallow rules does not prevent your page title and URL from surfacing in ChatGPT search answers. OpenAI's publishers FAQ clarifies that if OpenAI discovers a URL through a third-party search provider or external links, ChatGPT Atlas can still surface the link and title as a navigational result.

To remove a page from search answers entirely, you must implement a noindex robots meta tag. The crawler must be allowed in robots.txt to fetch the HTML and read that noindex instruction. Disallowing the crawler in robots.txt blocks the bot from discovering the meta tag.

Another common error is inserting internal and external links directly into answer capsules. The Search Engine Land audit found that over 90% of quoted capsules contained no links. Hyperlinks signal that the definitive answer sits on another destination, causing the language model to skip the passage.

Finally, site teams often overlook application firewalls and anti-bot verification. Cloudflare or Akamai WAF rules that return 403 Forbidden status codes, or JavaScript challenges that block non-browser requests, stop OAI-SearchBot from reading your text even when robots.txt permissions are correct.

What's new as of 2026-09-24

Recent documentation updates have established clear separation between search crawling, advertising checks, and model training. As of 24 September 2026, OpenAI's crawler documentation specifies four distinct bots and confirms that opting out of OAI-SearchBot prevents inclusion in search answers within roughly 24 hours.

On 23 April 2026, Search Engine Journal reported that OpenAI added OAI-AdsBot to its crawler documentation following ad experiments begun on 9 February 2026. OAI-AdsBot verifies ad landing page safety and relevance without contributing data to model training.

OpenAI advertiser guidance, updated three days before 24 September 2026, confirms published IP address lists for OAI-SearchBot and OAI-AdsBot to help system administrators configure firewall allowlists. In addition, Quora Business reported on 11 June 2026 that ChatGPT cites brands in only 0.59% of evaluated answers, reinforcing the need for structured capsules and third-party corroboration.

Key takeaways

  • OAI-SearchBot manages content inclusion for ChatGPT search answers, and site owners can allow it while disallowing GPTBot to prevent model training.
  • Search Engine Land found that 72.4% of blog posts cited by ChatGPT included an answer capsule, and roughly 91% of those capsules contained no links.
  • The only way to stop a URL and page title from surfacing in ChatGPT search results is a noindex meta tag that the crawler is allowed to fetch and read.
  • Inbound visitors from ChatGPT search carry the tracking parameter utm_source=chatgpt.com for monitoring inside Google Analytics.
  • Firewalls and bot-mitigation tools must allowlist published OpenAI IP ranges to prevent automated 403 Forbidden errors.

FAQ

Which crawler does ChatGPT use for search results?

ChatGPT uses OAI-SearchBot to discover and surface web content in its search results. Opting out of OAI-SearchBot removes your pages from ChatGPT search answers, but site owners can allow OAI-SearchBot while independently disallowing GPTBot to keep content out of foundation model training.

How long does a robots.txt change take to update in ChatGPT?

A robots.txt update takes roughly 24 hours to take effect across OpenAI search systems. If your server previously blocked OAI-SearchBot, crawler access and subsequent citation eligibility will adjust after that window.

How do you track traffic from ChatGPT in Google Analytics?

You track traffic from ChatGPT by filtering referral sessions that contain the utm_source=chatgpt.com query parameter. OpenAI appends this parameter automatically to all outbound links clicked within ChatGPT search answers.

Does blocking OAI-SearchBot remove your website entirely?

Blocking OAI-SearchBot stops your content from appearing in summaries, but your URL and page title can still appear as a navigational link if discovered through third-party search providers. To suppress a page completely, add a noindex meta tag and allow the crawler to fetch the page so it can process that directive.

To plan your broader search and discovery strategy, read Answer Engine Optimization (AEO) Explained for Founders.

Keep reading