← Blog
• AEO · AI visibility · ChatGPT

Where Does ChatGPT Get Its Information?

ChatGPT's answers come from two places: training data and live web search. What's in each, the crawlers that decide if you're a source, and how to become one.

By ClappX Team · August 19, 2026 · 8 min read

ChatGPT gets its information from two places: the data it was trained on, and — when it decides a question needs current information — the live web, which it searches and cites in real time. That split matters more than it sounds, because it decides whether your brand can ever appear in an answer. If ChatGPT can't read you in either place, it can't name you. Here's exactly where its answers come from, verified against OpenAI's own documentation, and what it takes for a brand to become one of the sources.

TL;DR

  • Every ChatGPT answer draws on training data (the model's built-in knowledge, frozen at a cutoff date), live web search (pages retrieved and cited in the moment), or both.
  • OpenAI says training data comes from three buckets: publicly available internet content, licensed third-party content, and content from users and human trainers.
  • When ChatGPT searches, it sometimes routes queries through third-party search providers — Bing is the one OpenAI has named — then reads the pages and cites them.
  • Three OpenAI crawlers decide whether your site can be a source: GPTBot (training), OAI-SearchBot (ChatGPT search), and ChatGPT-User (live fetches).
  • The most common own-goal we see: a robots.txt that blocks these crawlers by accident, making the brand invisible no matter how good its content is.

What are the two sources behind every ChatGPT answer?

Every answer ChatGPT gives is built from one or both of these:

  • Training data. A large body of text the model learned from before release. This is its built-in knowledge — broad, fluent, and frozen at a point in time.
  • Live web search. When a question needs current information, ChatGPT can search the web, retrieve real pages, and build the answer from them with citations. OpenAI's ChatGPT search announcement (October 2024) describes it plainly: ChatGPT chooses to search based on what you ask, or you can trigger a search yourself via the globe icon, and responses that used search show their sources.

Most answers about established topics come from training data alone. Answers about anything recent, changing, or specific — prices, "best X" lists, product comparisons — increasingly come from search. For a brand, the difference is the whole game, because each source has its own door you have to be let in through.

You can tell which one you're looking at. An answer with linked sources — inline citations or a sources sidebar — was built at least partly from live retrieval. An answer with no sources at all is the model speaking from training data, which means every fact in it is only as fresh as the model's cutoff.

What is actually in ChatGPT's training data?

According to OpenAI's own policy documentation, its foundation models are trained on three sources of information:

  • Publicly available information on the internet — public webpages, forums, blogs, and posts.
  • Information licensed from third parties — OpenAI has struck content deals with media outlets and other providers of specialized content.
  • Information provided by users, human trainers, and researchers — including ChatGPT conversations (which users can opt out of) and conversations written by AI trainers.

The critical limit is the knowledge cutoff: each model version stops learning at a certain date and knows nothing after it. The exact date varies by model — OpenAI publishes a cutoff for each release, and it moves forward with new versions — so "what does ChatGPT know about us" is really a per-model question. Without searching, ChatGPT can be confidently, fluently out of date — it will describe your product line as it existed at the cutoff, cite your old pricing, or repeat a stale third-party review as if it were current. That's not a bug you can file a ticket about; it's how training data works, and it's precisely the gap live search exists to fill.

For a brand, training data is the slow surface. You can't push an update into a model that's already shipped — you can only make sure the public web says accurate, consistent things about you, so the next training run learns the right facts.

When does ChatGPT search the live web instead?

Roughly: whenever the answer depends on being current. OpenAI's design is that ChatGPT decides for itself when a query would benefit from fresh information, searches automatically, and marks the response with clickable sources — inline citations plus a sources sidebar. You can also force a search manually.

For a brand, this is the winnable surface, and it's worth being precise about why. You don't have to wait for a training run to be included in a searched answer. If your page is readable, credible, and actually answers the question being asked, it can be retrieved and cited today. It's also why the same question can produce different answers on different days — searched answers depend on what was retrieved in that moment, which is one of several reasons ChatGPT gives different answers to the same question.

Does ChatGPT use Google or Bing?

Neither is the full answer. When ChatGPT searches, OpenAI's help documentation says it sometimes partners with third-party search providers — rewriting your prompt into targeted queries it sends them — and Bing is the provider OpenAI has publicly named. Independent testing (including a widely cited Semrush analysis) has found evidence of Google results appearing in ChatGPT's retrieval too, but OpenAI hasn't confirmed a Google partnership, so treat that as observed behavior rather than documented architecture.

The practical takeaway is less about which index and more about what it implies: your visibility in classic search still feeds your visibility in ChatGPT. A page that no search engine surfaces is unlikely to be retrieved when ChatGPT goes looking. AEO doesn't replace the crawlable, indexable foundation — it builds on it.

Which crawlers decide whether your site can be a source?

Whether ChatGPT can use your site at all comes down to the crawlers OpenAI operates, and — as of OpenAI's current crawler documentation (checked August 2026) — they are not interchangeable:

  • GPTBot crawls content that may be used to train future models. Block it in robots.txt and you're telling OpenAI to leave you out of the model's built-in knowledge.
  • OAI-SearchBot surfaces websites in ChatGPT's search features. Block it and you will not appear in ChatGPT search answers — full stop, regardless of content quality.
  • ChatGPT-User fetches a page live when a user's action in ChatGPT calls for it. Because these fetches are user-initiated rather than automatic crawling, OpenAI notes that robots.txt rules may not apply to it the way they do to the crawlers.

This is the quiet own-goal we see most often: a brand blocks AI crawlers in robots.txt — usually an over-cautious default applied years ago, or a blanket rule copied from a template — and then wonders why it never shows up in ChatGPT. Blocking GPTBot is a defensible content-licensing stance for a publisher. Blocking OAI-SearchBot is almost never what a brand actually wants, because it removes you from the one surface where inclusion is winnable this week.

While you're in robots.txt, it's worth knowing the adjacent convention: llms.txt is a proposed plain-text map of your site for AI systems. It's a low-cost bet, not a proven lever — but checking your crawler access and your llms.txt stance is the same ten-minute job.

How do you become a source ChatGPT actually uses?

Three moves, in order:

  1. Open the door. Confirm robots.txt isn't blocking GPTBot or OAI-SearchBot, and that your key pages are indexable by the classic search engines whose results feed retrieval.
  2. Publish answer-shaped content. Clear, factual answers to the real questions buyers ask, stated directly enough that a model can lift them without hedging. Specific, checkable claims survive synthesis; adjectives get compressed out.
  3. Build evidence beyond your own site. Models weigh what independent sources say about you far more than what you say about yourself. Reviews, comparisons, and community mentions are what make an engine comfortable naming you.

That's the skeleton of answer engine optimization. The full playbook — entity clarity, third-party validation, and the order to fix things in — is in our guide to getting recommended by ChatGPT.

And the starting point is always the same: find out what ChatGPT currently does with your brand. Ask it your category's buying question and read the answer honestly — are you named, misdescribed, or absent? If you'd rather not run that check by hand across every engine, that's exactly what the scan below does.

See how AI describes your brand today.

Free scan of your paid waste and your AI visibility. 60 seconds, no card, no call.

Run free scan →

Common questions

Where does ChatGPT get its information?

From two sources: its training data — which OpenAI says comes from publicly available internet content, licensed third-party content, and content from users and human trainers — and, when it searches, the live web, which it retrieves and cites for questions that need current information.

Does ChatGPT use real-time information from the internet?

Only when it searches. By default it answers from training data, which is frozen at a knowledge cutoff. When a question needs current information, ChatGPT can search the web automatically and show the sources it used as inline citations and a sources sidebar.

Does ChatGPT use Google or Bing for search?

OpenAI says ChatGPT search sometimes uses third-party search providers, and Bing is the provider it has publicly named. Independent testing has also observed Google results in its retrieval, but OpenAI hasn't documented a Google partnership — either way, classic search visibility feeds ChatGPT visibility.

Can I stop ChatGPT from using my website?

Largely, yes. Blocking GPTBot in robots.txt keeps your content out of future model training, and blocking OAI-SearchBot removes you from ChatGPT search answers — which is usually the opposite of what a brand wants. ChatGPT-User fetches are user-initiated, and OpenAI notes robots.txt rules may not apply to them.

See how AI describes your brand today.

Free scan of your paid waste and your AI visibility. 60 seconds, no card, no call.

Run free scan →
ClappX · Be found everywhere. Waste nothing.