Most startups that ChatGPT never mentions are not losing a quality contest. They are failing a fetch. The crawler that would have read the page got a 403 from a firewall rule, or it got an empty <div id="root"> because the content only appears after JavaScript runs, and a page the engine cannot read is a page it cannot cite.
Fix the fetch, and you are competing on the part that actually decides citations: whether several independent sources say the same thing about your product. Here is the order to work through it, with the checks that take five minutes each.
How AI answers pick their sources
An answer engine can mention you in two ways, and only one of them responds to anything you do this quarter.
The first is training data: what the model absorbed before it shipped. Getting into that takes a training cycle, many months, and you cannot request it.
The second is retrieval. When someone asks ChatGPT, Perplexity or Google a question that needs current information, the engine runs searches, pulls a handful of pages, and writes an answer from the passages that agree with each other. ChatGPT search draws on OpenAI's own OAI-SearchBot crawl plus third-party search providers, historically led by Bing. Perplexity runs its own crawler and index. Google AI Overviews are built from Google's own index, which is why the same page that ranks tends to be the one that gets quoted.
The consequence is simple. A product named on five independent pages, each describing it the same way, beats a product described brilliantly on one page. Retrieval engines are looking for agreement, and your own site is only one vote.
Step 1: let the right crawlers in
Each engine uses different user agents for different jobs, and the difference matters because blocking the wrong one removes you from live answers.
| User agent | Operator | What it does | Block it? |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Builds the index ChatGPT search reads from | Never |
| ChatGPT-User | OpenAI | Fetches a page when a user's question needs it | Never |
| GPTBot | OpenAI | Collects model training data | Your call |
| PerplexityBot | Perplexity | Builds Perplexity's index | Never |
| Claude-SearchBot, Claude-User | Anthropic | Search indexing and user-requested fetches | Never |
| ClaudeBot | Anthropic | Collects model training data | Your call |
| Bingbot | Microsoft | Bing's index, which feeds Copilot and parts of ChatGPT search | Never |
| Google-Extended | Controls Gemini training use only, not AI Overviews | Your call |
Robots.txt is only half of it. The other half is the bot protection on your CDN or host, which can reject these agents before robots.txt is ever read. Cloudflare added a one-click switch for blocking AI crawlers, and on some zones it is on by default, so check it rather than assuming.
Test the real response, not the config file:
curl -s -o /dev/null -w "%{http_code}\n" -A "OAI-SearchBot/1.0" https://yourdomain.com/pricing
curl -s -A "PerplexityBot/1.0" https://yourdomain.com/pricing | grep -c "your product name"You want a 200 and a non-zero count. A 403, a challenge page, or a count of zero means that engine cannot see you.
Step 2: put the answer in the HTML
Googlebot renders JavaScript. Most AI crawlers do not: they fetch the raw HTML and move on. If your marketing site is a client-rendered single-page app, the pricing table, the feature list and the FAQ may not exist in the version those crawlers read.
The check is "view source", not the browser's element inspector. Open view-source on your homepage and your most important feature page, search for a sentence from the body copy, and if it is not there, the page is invisible to retrieval. Static generation or server rendering fixes it. Every page on this site is statically generated for exactly this reason, so the HTML a crawler receives is the full article.
Step 3: write answers an engine can lift
Retrieval engines quote passages, not pages. The passage they pick is usually the one that answers the question in a sentence or two without needing the paragraph before it.
So write the first two sentences under every H2 as if they will be pasted somewhere without context. Name the entity, state the specific fact, then explain.
Before: "There are a lot of ways to think about pricing, and it really depends on your team."
After: "Acme costs $12 per seat per month, with a free plan for up to three users. Teams over 50 seats get volume pricing on request."
The second version contains a product name, two numbers and a condition, and it survives being quoted on its own. Question-shaped headings help for the same reason: an H2 that reads "How much does Acme cost?" matches the question someone typed almost word for word. This is the same specificity that makes low-competition, question-shaped keywords winnable in Google at a low Domain Rating.
Paste your most important page into ChatGPT and ask it to summarize what the product does, who it is for and what it costs. If the summary is vague or wrong, the page is vague or wrong, and a retrieval engine will make the same mistake with less patience.
Step 4: add structure, and keep expectations honest
FAQ blocks with FAQPage schema are the cheapest structural win. Google restricted FAQ rich results to a small set of authoritative sites back in 2023, so do not add them for the search snippet. Add them because an explicit question paired with a two to four sentence answer is the easiest possible thing to extract. Every post on this blog, including this one, ends with one for that reason.
Organization schema with sameAs links to your X, LinkedIn, GitHub, Crunchbase and Product Hunt profiles tells engines that those profiles are the same entity as your site. It costs one JSON-LD block.
Then there is llms.txt. "llms.txt" gets around 5,400 US searches a month, more than any other query in this space, and as of this month none of the major answer engines has confirmed it influences which sources they cite. Add it if you like. Do not do it first.
Step 5: build consensus off your own site
This is where most of the result comes from, and it is the step founders skip because it does not happen in their codebase.
The sources that retrieval engines cite for "best X for Y" questions are overwhelmingly third-party: directories and review platforms, comparison posts, Reddit threads (Perplexity cites Reddit heavily), and YouTube videos with transcripts. You want to be named on as many of those as you can, with the same one-line description each time.
That last part is the lever. If G2 calls you "an invoicing tool for freelancers", Product Hunt calls you "AI billing for creators" and your homepage says "the financial OS for independents", no engine can confidently place you in a category. Pick one sentence with your category words in it, and use it on every listing.
The fastest way to get fifteen or twenty consistent third-party pages is a curated directory batch. The case for doing it early, and why platforms like G2 and Crunchbase double as consensus signals, is in the directory submissions playbook. The specific list and the order to submit in is the launch week checklist. The same links lift your Domain Rating, which is covered phase by phase in the DR roadmap.
Step 6: get indexed where the engines search
An engine that retrieves from Google's index cannot cite a page Google has not indexed, and the same goes for Bing. Two habits cover both:
- Google: submit your sitemap in Search Console and inspect every new URL on the day it ships. The weekly routine is in the five Search Console reports post.
- Bing: verify the site in Bing Webmaster Tools, then submit new and changed URLs through IndexNow. It is a single POST request and a key file in your public folder, and it gets pages into the index that ChatGPT search leans on without waiting for a crawl.
How to measure AI citations
Two numbers, checked monthly.
Referral traffic. In GA4, filter sessions by source for chatgpt.com, perplexity.ai, gemini.google.com and copilot.microsoft.com. ChatGPT also appends utm_source=chatgpt.com to the links it cites, so those visits are labelled even when the referrer is stripped.
A fixed prompt set. Write ten prompts your buyer would actually type ("best invoicing tool for freelance designers in the EU", not "Acme review"). Run each one in ChatGPT, Perplexity and Google once a month and record three things: whether you are named, which URL is cited, and which competitor is named instead. Keep the prompts fixed so month-over-month changes mean something.
A page that gets cited for a neighboring prompt but not the one you want is the AI version of a page-two ranking, and it responds to the same treatment: a targeted content refresh that adds the missing answer rather than a new page that starts from zero.
The order to do it in
| Step | Time | Why this position |
|---|---|---|
| Crawler access and CDN check | 30 minutes | Nothing else counts until the fetch works |
| View-source check, fix rendering | 1 hour to 2 days | Crawlers that cannot run JavaScript see nothing |
| Rewrite the first sentence of every key section | 2 hours | Makes your pages quotable |
| FAQ and Organization schema | 1 hour | Cheap structure for extraction |
| Consistent listings on 15 to 30 third-party sites | 2 to 4 weeks | This is what creates consensus |
| Google and Bing indexing, IndexNow | 1 hour, then weekly | Retrieval cannot cite unindexed pages |
| Monthly prompt set | 30 minutes a month | The only way to know if any of it worked |
Steps one to four are an afternoon of work on your own site. Step five is the slow part, and it is the one our directory submission service exists to take off your plate. For the rest of the system, from keyword research to measurement, grab the free guide and follow new breakdowns on X.
Frequently asked questions
- How do I get my startup mentioned by ChatGPT?
- Make sure OAI-SearchBot and ChatGPT-User can fetch your pages, serve the answer in plain HTML, and open every section with a self-contained sentence that names your product and a specific fact. Then get the same one-line description of your product onto independent pages: directories, comparison posts, Reddit threads. ChatGPT names products that several sources agree on, not the one with the best homepage.
- Does ChatGPT use Google or Bing to find sources?
- ChatGPT search combines OpenAI's own crawler, OAI-SearchBot, with third-party search providers, and Bing has historically been the main one. Google AI Overviews draw on Google's own index, and Perplexity runs its own crawler. In practice that means you want to be indexed in both Google and Bing, and submitting new URLs to Bing through IndexNow is the cheapest way to cover the second.
- Should I block GPTBot in robots.txt?
- GPTBot collects training data, so blocking it is a legitimate choice if you do not want your content in future models, and it does not remove you from ChatGPT search results. Blocking OAI-SearchBot or ChatGPT-User is the mistake that costs citations, because those are the agents that fetch pages for live answers. Most startups want visibility more than they want protection, so allowing all of them is the usual call.
- Does llms.txt help you get cited by AI?
- Not measurably, as of September 2026. It is a proposed convention for giving language models a map of your site, and none of the major answer engines has confirmed they read it when choosing sources. It takes ten minutes to add, so there is no harm in it, but it should come after crawler access, server-rendered HTML and third-party mentions, all of which demonstrably matter.
- How long does it take to get cited by Perplexity or ChatGPT?
- For live answers, days to a few weeks after your page is indexed by the search layer the engine uses, because those answers are retrieved at the moment someone asks. Showing up in what a model knows without searching takes a full training cycle, which is many months and outside your control. Work on the retrieval side first; it is the only one that responds to this quarter's effort.