“ChatGPT told me about you.” If you have not heard that sentence from a customer yet, you will, and the businesses hearing it today did specific, learnable things to be the name in the answer. This guide covers exactly what those things are: how ChatGPT actually decides what to say when someone asks for a recommendation, what the research shows makes content citable, and the step-by-step playbook we run for clients and on our own agency site, where ChatGPT is now a named referral source in our analytics.
First, the scale, so the effort makes sense. ChatGPT reported around 900 million weekly active users by early 2026, double its figure from a year before, and Semrush ranks it among the most visited websites in the world at over 5 billion monthly visits. More importantly for businesses: it increasingly sends people out. Semrush’s research measured outbound referral traffic from ChatGPT growing 206% during 2025, and by mid-2026 chatgpt.com’s US visits were still climbing 48% year over year (Semrush). And the visitors who arrive are unusually valuable: Semrush’s 2026 cross-industry data finds AI-driven visitors convert at roughly 4.4 times the rate of standard organic traffic, with SE Ranking measuring 68% longer time on site. Small stream, strong current.
How ChatGPT Decides Who to Mention
To be cited, it helps to know the machine you are persuading. When someone asks ChatGPT a question about businesses, products, or services, the answer is assembled from two layers, and each is a different optimization problem.
Layer 1: Trained knowledge. The model carries a compressed impression of the public web from its training data. If your business has years of consistent presence, a clear website, directories, press, reviews, all describing the same entity the same way, an impression of you exists in that memory. Thin or contradictory footprints produce hesitant or absent answers. You influence this layer slowly, through sustained consistency, which is why starting now beats starting when it is obvious.
Layer 2: Live search. For anything current, local, or specific, ChatGPT searches the web (drawing notably on Bing’s index) and reads a shortlist of pages before answering, citing the ones it leaned on. This layer moves on SEO timescales: if you are indexed, crawlable, and ranking respectably for the relevant queries, you are in the pool it reads. This is why “ChatGPT SEO” is not a metaphor; classic search optimization is literally the qualifying round.
Then: extraction. From the shortlist, the model repeats what it can safely repeat. The foundational research here is the Princeton-led GEO study (Aggarwal et al., KDD 2024), which tested nine content strategies across a 10,000-query benchmark and found that adding citations, statistics, and quotations boosted a source’s visibility in generated answers by up to 40%, while keyword stuffing performed below doing nothing at all. Machines cite evidence, not adjectives.
And: corroboration. Muck Rack’s 2026 analysis found 84% of AI citations pointed to earned editorial coverage, third-party publications rather than brands’ own pages. ChatGPT triangulates: what you say about yourself is a claim; what the web consistently says about you is a fact.
The 10-Step Playbook to Get Cited
This is the working sequence, ordered so each step feeds the next.
1. Run the baseline test. Ask ChatGPT (and Gemini and Perplexity, since answers differ) the five to ten questions a customer would ask before hiring you: “best [your service] in [your city],” “who should I hire for [problem],” “[competitor] vs alternatives.” Record who gets named, with which sources. This is your scoreboard, and re-running it monthly is your measurement.
2. Fix your entity everywhere. One canonical business description, name, services, locations, and differentiators, deployed identically across your site, Google Business Profile, LinkedIn, Bing Places (remember whose index ChatGPT reads), and the directories that matter in your category. Contradictions read as uncertainty, and uncertain entities do not get recommended.
3. Implement structured data. Organization, LocalBusiness, Service, FAQPage, and Article schema turn your pages from prose into machine-readable facts: who you are, what you do, where, at what rating. Schema does not guarantee citations; it removes the excuse for getting your facts wrong.
4. Publish pages that answer the baseline questions. Every question from step one deserves a page that answers it directly in the first paragraph, then earns the answer with depth. Conversational queries are long and specific, and the page that matches “how much does a Shopify store cost in Canada” beats the page that matches “our web services.”
5. Load your pages with extractable evidence. Apply the research mechanically: real statistics with named sources, one-line quotable answers near the top of each page, expert attribution, comparison tables, and clear headings that mirror questions. This article is doing it in front of you, sourced stats and all, because the tactic and the demonstration are the same thing.
6. Publish case studies with numbers. The most repeatable sentence a model can find about you is a quantified result. “452 leads at $19 each” (Landmark Real Estate) and “219 new patients at $25.87 per lead” (The Tooth Place) are exactly the kind of specific, safe-to-repeat claims assistants reach for when asked whether an agency actually delivers. Vague success stories are invisible; numbers travel.
7. Get into the third-party pages assistants read. Since earned coverage dominates citations, the highest-leverage work is often off your site: category listicles and roundups, local business journalism, industry publications, podcast appearances with show notes, and credible directories. When ChatGPT answers “best X in Y,” it is very frequently synthesizing precisely those pages, so being in them is being in the answer.
8. Accumulate reviews with substance. Review platforms are corroboration engines. Volume, recency, and detailed text (reviews that name the service and outcome) all feed the impression of an established entity. A 5.0 rating across dozens of named reviews is a fact pattern models treat as signal.
9. Keep the crawlers in. Check robots.txt: GPTBot (OpenAI), Google-Extended, PerplexityBot, and ClaudeBot must be allowed if you want the corresponding engines reading your pages. Keep the site fast and the HTML clean; extraction favours pages that parse easily.
10. Measure it like a channel. Segment AI referrals in GA4 (chatgpt.com and friends), add “AI assistant” to your how-did-you-hear options, and keep the monthly prompt test from step one. What follows is what that measurement looks like on our own site.
The Sources AI Models Trust (and How to Be in Them)
Citation analyses keep converging on the same source hierarchy, and it should shape where you spend effort. At the top: reference and community platforms. Pew’s research on AI summaries found Wikipedia, YouTube, and Reddit among the most-linked domains, with government sites over-represented in AI answers relative to normal results (6% of AI summary sources versus 2% in standard listings). In the middle: editorial coverage, the trade publications, local journalism, and expert roundups that Muck Rack’s analysis found supplying 84% of AI citations. At the base: high-authority directories and review platforms that corroborate the basics.
What that means in practice, source by source:
- Wikipedia and Wikidata. Most small businesses will not merit an article, and forcing one backfires. What is achievable: ensuring your company’s factual footprint (founding, location, services) is consistent everywhere Wikipedia-adjacent tools read, and earning the press coverage that notability eventually requires.
- YouTube. Video descriptions and transcripts are crawlable evidence. A channel answering your category’s questions gives assistants a second, corroborating voice that you control.
- Reddit and community forums. Assistants weight candid community discussion heavily. Never astroturf; do participate honestly where your category is discussed, and give customers reasons to mention you unprompted. One genuine “I used these guys, here were my numbers” thread outweighs pages of marketing.
- Trade and local press. The single highest-leverage earned source. Pitch data, not announcements: original statistics from your own work are the currency journalists and, downstream, models both cite.
- Directories and review platforms. Google Business Profile, Bing Places, industry-specific directories, and review sites are where models verify you exist, operate, and are rated. Completeness and consistency beat volume.
The pattern across all five: assistants trust what you cannot easily fake. Which is inconvenient for shortcuts and excellent for businesses that actually deliver, because the evidence trail of real work is precisely what the machines are built to find.
A Worked Example: From Invisible to Named in One Local Category
Here is the shape of the playbook applied, drawn from the pattern we run for local service clients. A clinic starts with the baseline test: assistants asked “best physiotherapy clinic near [city] for sports injuries” name two competitors, citing a local roundup, a directory, and review counts. The gap analysis writes itself. Month one is entity work: canonical description everywhere, schema deployed, Bing Places claimed, review velocity restarted with substance encouraged. Month two is content: a page answering the exact baseline question with treatment specifics, real pricing, clinician credentials, and a case-study page with numbers, the same structure that took our SEO client Get Back Physiotherapy to 42,000 clicks and 612,000 impressions from search, because ranking assets and citation assets are the same assets. Month three is corroboration: outreach lands the clinic in the local roundup that was carrying the competitors, and a data pitch (“what 500 running-injury assessments taught us”) earns a trade mention. Re-run the baseline test at day 90 and the clinic is appearing in some answers, cited alongside the incumbents; by month six, in thin local categories, it is often the consistently named option. No step required tricks; every step required actually being verifiable.
Our Own Receipts: ChatGPT in Our GA4
We tell clients not to trust agencies that only cite industry averages, so here is our first-hand data. ChatGPT appears as a distinct referral source in Infinity Digital’s GA4, sending [FILL: sessions per month] sessions in [FILL: most recent month], up from [FILL: baseline] in [FILL: baseline month], alongside smaller referral streams from [FILL: Perplexity/Copilot/Gemini as applicable]. Those sessions behave the way the industry data predicts: [FILL: engagement comparison vs organic, e.g. longer average engagement time], and they include real enquiries, prospects who arrive already briefed on our services and case studies because an assistant summarized them. The absolute numbers are modest next to organic search, exactly as the roughly-1%-of-web-traffic estimates suggest, and that is the honest pitch: this channel is small, compounding, disproportionately high-intent, and still nearly uncontested in most Canadian categories. The receipts are why we treat it as a channel rather than a curiosity, and they are the same receipts we build for clients through our GEO program.
The 90-Day Plan, Week by Week
For owners who want the playbook as a calendar:
Weeks 1 to 2: Baseline and audit. Run the prompt test across ChatGPT, Gemini, and Perplexity; record every citation. Audit your entity: every profile, directory, and description, listed against the canonical version. Check robots.txt for AI crawler access and verify Bing Webmaster Tools.
Weeks 3 to 4: Entity repair. Fix every inconsistency found. Deploy or correct Organization, LocalBusiness, Service, and FAQPage schema. Claim what is unclaimed, kill duplicate listings, and standardize the description everywhere.
Weeks 5 to 8: Evidence content. Publish or rebuild the pages answering your baseline questions: direct first-paragraph answers, sourced statistics, comparison tables, FAQ blocks. Ship at least one quantified case study. Add quotable one-liners to your money pages.
Weeks 9 to 10: Corroboration push. Identify the specific third-party pages your competitors’ citations came from; pitch or earn your way into the ones that accept businesses. Send the data pitch to one trade publication and one local outlet. Ask your five happiest customers for detailed reviews.
Weeks 11 to 12: Measure and iterate. Re-run the full prompt test, compare against the baseline, segment AI referrals in GA4, and note which new pages earned citations. The deltas tell you which lever to pull hardest in the next quarter.
Twelve weeks does not finish the job, the trained-knowledge layer keeps compounding for quarters after, but it reliably moves the search layer, and it produces the measurement habit that separates a channel from a hope.
The Measurement Toolkit
Because no Search Console exists for assistants yet, measurement is assembled from four parts, all doable in-house.
1. GA4 referral segmentation. Build a channel segment matching the AI referrers: chatgpt.com, gemini.google.com, perplexity.ai, copilot.microsoft.com, claude.ai. Track sessions, engagement, and conversions for the segment monthly. Two warnings from experience: a share of assistant-driven visits arrives untagged as direct traffic when users copy links rather than click, so your measured number is a floor, not the total; and dark traffic aside, the trendline is still the honest signal.
2. The prompt panel. A fixed spreadsheet of 10 to 25 prompts (your baseline questions plus variants), run monthly on each engine, recording mentions, citations, and the sources cited. Consistency matters more than sophistication: same prompts, fresh sessions, logged verbatim. Over quarters this becomes your rankings report for the generative era.
3. Citation monitoring. Brand-monitoring and AI-visibility tools now track when assistants cite your domain across query sets, useful once your footprint grows past what manual testing covers. Start manual; graduate when the panel gets unwieldy.
4. Attribution at the source. Add “AI assistant (ChatGPT, Gemini, etc.)” to every how-did-you-hear form and train whoever answers the phone to log it. This is the least glamorous instrument and the one that converts the channel from theory to revenue in your CRM, because “ChatGPT told me about you” is only data if somebody writes it down.
Review the four together quarterly: referrals up, panel mentions up, citations spreading, and source-attributed leads appearing is what winning looks like, and any one metric alone can mislead.
What Not to Do
The failure modes are as instructive as the tactics. Do not keyword-stuff for chatbots; the research shows it performs worse than nothing. Do not fake statistics or reviews; models corroborate across sources, and fabrications that get repeated eventually get traced, with your brand attached. Do not build “AI landing pages” no human is meant to read; extraction favours genuinely useful pages, and engines increasingly discount content that exists only to be cited. Do not block AI crawlers and then wonder about visibility. And do not chase one engine’s quirks: Similarweb’s 2026 data shows ChatGPT’s share of generative-AI traffic falling from about 76% to 53% in a year as Gemini and Claude grow, so the durable strategy is being citation-worthy in general, not gaming a single bot.
Where Ads Fit: The Other Way Into the Conversation
Citations are earned, but paid placement inside AI experiences is emerging as its own discipline, and it belongs in the same planning conversation. Sponsored visibility in AI surfaces, and campaigns built around how people phrase questions to assistants, are early but real; we run this as ChatGPT Ads management for brands that want presence while the organic footprint builds. The earned and paid tracks reinforce each other the way SEO and Google Ads always have: one buys the moment, the other compounds.
Frequently Asked Questions
Yes, in two senses: classic SEO puts your pages in the indexes ChatGPT searches, and citation optimization (evidence-dense content, entity consistency, earned coverage) determines whether it repeats you. The combined discipline is generative engine optimization, covered in full in What Is GEO.
Search-layer citations can appear within weeks if you publish a strongly evidenced page that ranks for the question. Trained-knowledge presence, being recommended without a live search, builds over months of consistent signals. Run the monthly prompt test and expect the search layer to move first.
Almost always corroboration: they are present in the listicles, directories, reviews, and press that the model reads, and you are not, or your entity information is inconsistent enough to lower confidence. The baseline test usually shows exactly which sources are carrying them.
Not directly; its live search leans on Bing's index among other sources. Practically, pages that rank well tend to do so everywhere, but it is worth claiming Bing Places and verifying your site in Bing Webmaster Tools, a step most Canadian businesses skip.
On the data, it is some of the best traffic there is: 4.4x conversion versus standard organic in Semrush's 2026 cross-industry measurement, with far longer engagement. These visitors arrive pre-sold by a recommendation, which is why we track them separately.
If you sell content itself, maybe. If you sell products or services, blocking the crawler is refusing the referral: the assistant cannot recommend what it cannot read.
Yes, disproportionately, because assistant queries are questions. FAQ content with direct first-sentence answers, marked up with FAQPage schema, maps one-to-one onto how people prompt.
Optimize once, benefit everywhere: the citation-worthiness fundamentals transfer, and the engine mix is shifting fast enough that single-engine strategies age badly. The baseline test should always run across at least three engines.
Slower, but yes, through the search layer: publish the best-evidenced answers in your niche, get into a few credible third-party pages, and keep your entity spotless. New businesses have won recommendation queries in thin local niches within a quarter.
The baseline test, because it converts anxiety into a to-do list: the sources citing your competitors are, item by item, the places you need to appear.