🎬 Want to become a clipper/ugc creator? Apply now!

Guide · 13 min read

Answer engine optimization, measuredHow to Get Cited by ChatGPT, and How to Prove It Worked

Analytics counts clicks, not the AI answers that named you without one. This is a hand-runnable protocol for finding out whether ChatGPT, Perplexity and Google AI Overviews actually cite your company, what the controlled evidence says about changing that, and how to avoid reading noise as movement.

Three rule-outs first. This is not citing ChatGPT in a bibliography, and not local NAP directory citations. Where this article says clipping it means short-form video clipping, cutting long video into short clips, not press clipping or hair clippers. The protocol here runs on ChatGPT, Perplexity and Google AI Overviews; Google AI Mode and Grok sit outside it.

One scope note on the best-known statistic in this space. Pew Research Center, 22 July 2025, drew on browsing data from 900 US adults fielded across March 2025, with the search results themselves collected 7 to 17 April 2025, and it covers Google AI Overviews only, not ChatGPT. In that April collection, 88 percent of the AI summaries cited three or more sources. Google AI Overviews behaved that way; no such rule holds for every engine.

01

How often ChatGPT actually names a B2B SaaS brand

Rarely, and the number most people quote belongs to a different engine.

Superlines sells AI-visibility software; the figures come from 34,234 responses in the window it labels 14 January to 13 February 2026, and it does not disclose its brand set, its prompt set, how many responses came from each platform, or a definition for either metric it reports. In the ChatGPT row, citation rate is 0.59 percent and brand visibility is 0.14 percent. Superlines defines neither column, so this article claims only the figures it publishes.

Superlines puts Perplexity's citation rate at 13.05 percent against ChatGPT's 0.59 percent, so the higher rate people quote is Perplexity's.

EngineFigure Superlines publishesColumn it sits in
ChatGPT0.59%Citation rate
ChatGPT0.14%Brand visibility
Perplexity13.05%Citation rate

Conductor sells SEO software; its benchmarks report dated 6 July 2026 covers 13,770 domains over an analysis period of May to September 2025 and reports ten industries, and its Information Technology category covers software and services, semiconductors and technology hardware. Conductor puts AI referrals at 1.08 percent of all website traffic blended across those ten industries, Information Technology highest at 2.80 percent, and ChatGPT at 87.4 percent of those referrals. Answer engine optimization for B2B SaaS, meaning work aimed at being named inside AI answers rather than ranked in blue links, runs against that 2.80 percent.

Gartner's press release of 19 February 2024 is a prediction with no sample behind it, made 29 months before this article, and Gartner has published no public re-test of it. The prediction was a 25 percent drop in traditional search volume by 2026.

A disclosure that applies to this whole page: six of the nine sources in this article are published by, or written by someone employed by, a company that sells AI-visibility or SEO software. Read them for direction, not for precision. We write separately about how B2B SaaS companies plan their marketing.

02

Why running the prompt once tells you nothing

Because the same prompt, run again with nothing changed, returns a different answer often enough that a single run is not evidence.

Zatuchin's preprint arXiv:2607.13304, 14 July 2026, is single-authored and not peer reviewed, and its author works for a company that sells AI-visibility software; it covers 12,933 responses from 20 Central and Eastern European brands across 8 languages and 3 models, and it measures sentiment polarity, not raw brand mentions. In the model fitted to the 7,173-response stability subset, within-prompt resampling accounted for 34.8 percent of variance, query language 32.0 percent, brand identity 1.6 percent. In the model fitted to the full 12,933 responses, query language was 26.5 percent and brand identity 1.5 percent, an intraclass correlation of 0.0146, meaning repeated answers to one prompt cluster barely more than answers to different prompts do. The two models are different and their percentages are not interchangeable.

1.5-1.6%of the variance in AI answers is explained by which brand is being asked about, in both Zatuchin models (arXiv:2607.13304, July 2026). Almost everything that moves is the run and the wording, not the brand.

MaxAEO sells answer-engine-optimization software and its founder Chris Han published this analysis on 21 July 2026; it reports 23,040 answers from 18 B2B SaaS brands across six engines, roughly 320 prompts run 12 times each over a 14-day window, and it publishes no dataset and no methodology appendix, so the sample is the vendor's own unaudited claim. Across those 23,040 answers, "a single run disagreed with the same prompt's 12-run majority verdict 19% of the time."

Zatuchin measures sentiment polarity on Central and Eastern European brands; MaxAEO measures brand mentions on B2B SaaS brands. Two different methods, one conclusion. My read, not theirs: agreement between independently flawed measurements is worth more than either alone. Neither is clean: one measures polarity rather than mentions, the other is unaudited.

03

The 300-run protocol: how to measure whether you are cited

Fifty prompts, run twice each, across three engines is 300 runs, the smallest test you can run by hand. It shows direction, not small moves. The arithmetic is 50 x 2 x 3 = 300 runs, and 300 / 3 = 100 runs per engine.

How the runs are split matters more than the total. MaxAEO models three designs at a 30 percent mention rate: 200 prompts x 3 runs = 600 runs at plus or minus 5.4 percentage points, 100 prompts x 6 runs = 600 runs at plus or minus 7.2, and 20 prompts x 30 runs = 600 runs at plus or minus 15.4. The same 600 runs, three designs, and the margin nearly triples on the split alone. My reading of both stability sources, not a finding either publishes: prompts stabilize your overall share, repeats stabilize the verdict on a single prompt. The protocol answers the first, hence 50 prompts and 2 repeats.

  1. Write 50 prompts across three types

    Category: "best customer support AI for mid-market SaaS". Comparison: "Intercom vs Zendesk vs Ada for customer support automation". Brand: "is Ada any good for customer support". Use the words your buyers would actually type.

  2. Run each prompt twice on each of three engines

    ChatGPT, Perplexity and Google AI Overviews. That is 300 runs total, 100 per engine, the smallest hand-runnable design.

  3. Count a hit when the answer names or links your company

    Naming without a link counts, and so does linking without naming. Record hits per engine, per prompt type.

  4. Track at least one competitor on the same 50 prompts

    Your share only means something against the share of a company your buyers actually compare you with.

  5. Budget four to six hours of operator time per window

    An estimate rather than a measured figure. Re-run when something material has changed on your site or in your market, not on a fixed calendar.

The worked example below is illustrative, with invented numbers, not measured ones. In the illustration, 1 + 15 + 8 = 24 hits out of 300 runs, which is 8.0 percent, against a tracked competitor's 12 + 22 + 19 = 53 out of 300, or 17.7 percent.

EngineYour company, hits per 100 runsOne tracked competitor, hits per 100 runs
ChatGPT112
Perplexity1522
Google AI Overviews819

MaxAEO's own recommendation is 150 to 200 prompts run 3 times each, per engine, per period, at roughly plus or minus 5 percentage points. On my read that is the better test; 50 prompts is the hand-runnable floor, not the optimum.

No margin of error is printed for a 50-prompt design, because MaxAEO publishes margins for 200 x 3, 100 x 6 and 20 x 30 and no other, and inventing a fourth would be the failure this article exists to name. Per-engine rows sit on a 100-run base and are the least stable numbers in the test.

04

What the controlled evidence actually says about getting cited

There is very little controlled evidence, and what exists is less encouraging than the category's marketing.

Ahrefs is an SEO tool company using its own crawl index; on 11 May 2026 it published an observational difference-in-differences study, not a randomized experiment, of 1,885 pages that added schema between August 2025 and March 2026 against the 4,000 controls it reports, all restricted to pages with 100 or more Google AI Overview citations in February 2025. Difference-in-differences compares the change in a treated group against an untreated one over the same period. Google AI Overviews citations moved -4.6 percent, which Ahrefs calls small but statistically significant at roughly 1 in 2,500; Google AI Mode citations moved +2.4 percent and ChatGPT citations +2.2 percent, both statistically indistinguishable from zero. One unreconciled detail: Ahrefs reports 4,000 control pages and also describes three control URLs per treated page, which would give 5,655, and the source does not reconcile them.

Aggarwal and co-authors published arXiv:2311.09735 on 16 November 2023, 350 days before ChatGPT Search launched; their main results were measured against an engine they simulated from the top five Google results plus GPT-3.5-turbo, and the Table 5 figures used here were run against live Perplexity.ai, both using two visibility metrics they defined themselves over a 10,000-query benchmark. The paper's subject is generative engine optimization, meaning page changes intended to make an AI engine's answer feature you. Inside Table 5, adding statistics to a page moves Subjective Impression from 24.7 to 33.9, a 37.2 percent relative lift, and Position-Adjusted Word Count from 24.1 to 26.2, an 8.7 percent relative lift.

TacticWhat the controlled evidence saysWhat it was measured on
JSON-LD schema markupMovement statistically indistinguishable from zeroChatGPT, 1,885 pages, observational (Ahrefs)
JSON-LD schema markupA 4.6% decline, small but statistically significantGoogle AI Overviews, same 1,885 pages (Ahrefs)
Adding statistics to a page37.2% on one author-defined metric, 8.7% on the otherLive Perplexity.ai, the paper's Table 5 (Aggarwal et al.)

What a SaaS marketing budget actually covers sets out the line items to plan against, including a first AI-visibility line.

05

Where AI citations actually come from

On two of the three engines, most of what gets cited is not a page the brand owns.

OtterlyAI sells AI-visibility monitoring; its AI Citation Economy report, published 1 February 2026 and updated 19 February 2026, covers more than one million AI citations across ChatGPT, Perplexity and Google AI Overviews from January to February 2026, and it does not publish the queries it ran.

59.8%of Google AI Overviews citations point at brand-owned domains (OtterlyAI, February 2026)
44.7%of ChatGPT citations point at brand-owned domains (OtterlyAI, February 2026)
28.9%of Perplexity citations point at brand-owned domains (OtterlyAI, February 2026)

The report also contradicts itself. It states in one place that community platforms hold 52.5 percent of citations against 47.5 percent for brand domains, and in another that brands hold 52.5 percent with the remaining 47.5 percent split across news sites and community forums. This article uses none of that aggregate split, only the three per-engine figures. A citation-source split also moves hard depending on which queries were run.

06

Reading your result, and what to do when the answer is zero

Read the overall 300-run share, ignore small month-over-month moves, and treat a zero as a coverage problem, not an optimization problem.

The noise floor below is my arithmetic, not any source's finding. On a 300-run base at an 8 percent share, a figure taken from the illustrative example rather than a measurement, the standard error of a single measurement is the square root of (0.08 x 0.92 / 300) = 1.57 percentage points. The difference between two independent measurements has a standard error of the square root of (1.57 squared plus 1.57 squared) = 2.22 points. A move callable at conventional confidence is 1.96 x 2.22 = 4.35, rounded up on purpose: do not act on a month-over-month move in the overall 300-run share of less than 4.4 percentage points. Quadrupling the base to 1,200 runs brings the smallest callable move only to 2.2 points.

One limit on that chain is mine, not any source's: it treats 300 runs as 300 independent observations, and repeats of the same prompt are not independent, so the true floor is wider than 4.4 points rather than narrower. No corrected figure is printed, because that needs a clustering figure no source here publishes. Treat 4.4 points as a minimum bar to clear, not a precise threshold.

Zatuchin puts brand identity at 1.5 to 1.6 percent of the variance in its models. A company appearing in 1 of 100 ChatGPT runs is effectively absent, not down from last month; per-engine numbers sit on a 100-run base and move too much to read month to month.

At zero, run three coverage checks.

How big a move is real on your run count?1.96 x sqrt(2) x SE
±1.57ppstandard error of a single measurement
4.4ppsmallest month-over-month move worth acting on

Mirrors the arithmetic above: standard error = the square root of (share x (1 - share) / runs); the callable move is 1.96 x the square root of 2 x that, rounded up.

This treats every run as independent, and repeats of the same prompt are not, so the true floor is wider than shown. Treat the result as a minimum bar to clear, not a precise threshold.

Are the 50 prompts the ones your buyers would type?

If the prompts are your words rather than theirs, a zero measures your prompt writing, not your visibility.

Does any page you own answer those questions in their words?

An engine cannot cite an answer that does not exist. Coverage comes before optimization.

Do the sources those engines cite name your company anywhere?

OtterlyAI puts brand-owned domains at 28.9 percent of Perplexity citations, so most of what gets cited there is somebody else's page. Being named in those pages is a distribution problem.

That third check is where distribution work connects to AI visibility: getting your product and your point of view into the pages and feeds that engines cite. We have also written about whether clipping works for brands.

How many prompt runs do I need to measure AI citations reliably?
Three hundred runs, split as 50 prompts x 2 repeats x 3 engines, is the hand-runnable floor. MaxAEO, 21 July 2026, an answer-engine-optimization vendor reporting 23,040 answers from 18 B2B SaaS brands, publishing no dataset and no methodology appendix, shows that at the same run count more prompts and fewer repeats buy more precision than the reverse. No margin of error is quoted for a 50-prompt design, because the source publishes none for it.
Does schema markup help you get cited by ChatGPT?
The one controlled read available says no. Ahrefs, 11 May 2026, an observational study by an SEO vendor of 1,885 pages, restricted to pages already carrying 100 or more Google AI Overview citations, found ChatGPT movement statistically indistinguishable from zero and a 4.6 percent decline on Google AI Overviews. The finding is scoped to that population and makes no general claim.
How often should I re-run the baseline?
Re-run when something material has changed on your site or in your market. On a 300-run base, treat any month-over-month move in the overall share smaller than 4.4 percentage points as inside measurement noise. The 4.4 is a rule derived by this article's author from the run count, not a figure published by any source.
Why is my company named on Perplexity but not on ChatGPT?
Because the engines cite different kinds of page. OtterlyAI, 1 February 2026, a visibility-monitoring vendor reporting on more than one million citations, which does not publish the queries it ran, puts brand-owned domains at 44.7 percent of ChatGPT citations and 28.9 percent of Perplexity citations, so Perplexity leans harder on pages a brand does not own.

Running the protocol tells you where you stand

Closing the gap takes content in front of the people asking those questions. Lumina Clippers runs managed short-form video clipping for B2B and consumer brands, from creator sourcing to multi-account posting across a 62,900+ creator network.

See managed short-form distribution
Rhys McKay

Rhys McKay · Founder & CEO, Lumina Clippers

Has led clipping campaigns delivering 18B+ views across a 62,900-clipper network

Rhys founded Lumina Clippers in 2024 and has run short-form distribution campaigns for crypto, SaaS, gaming, music and founder brands. He writes on clipping strategy, creator-led growth and brand visibility. Connect on LinkedIn · About the team →

More from the blog

← Back to all articles
All information on this page is fact-checked and kept up to date.