Aug 31, 2026

How to Verify Your Agency Is Actually Getting You Mentioned in ChatGPT Answers

Clarity Search AI Team

You verify it the same way you would verify any other claim: by asking for the raw evidence instead of the summary. That means a fixed list of buyer questions written down before the work starts, the full text of what each AI engine answered on a dated run, the source links that answer cited, and a number in your own analytics showing sessions and revenue arriving from AI engines. If your agency can produce those four things, they are working. If they can only produce a proprietary "AEO score" that went from 42 to 61, you have a slide, not a result.

Here is the audit, in the order you should run it.

Step 1: Lock the question list before anything else

Write down 10 to 25 questions your buyers actually type into an assistant, phrased the way a human phrases them. Not keywords. Real sentences: "best invoicing tool for freelance agencies," "alternatives to [competitor] for a 20-person team," "is [category] software worth it for under 50 employees."

Give that list to the agency and keep a copy. This is the only way to stop the most common form of AEO theater, which is reporting on questions selected after the fact because those are the ones you happened to win. A fixed question list turns a vibe into a measurement you can repeat next month.

Step 2: Get the baseline yourself, on day one

Before the agency reports anything, run those questions through ChatGPT, Claude, Gemini, and Perplexity yourself and screenshot the answers. Note which companies get named and in what order. Ten questions across four engines is about forty minutes of work, and it is the most valuable forty minutes in this whole process, because everything afterward is measured against it.

Most founders find the same thing here. The assistant confidently recommends three competitors by name and never mentions them once. That sting is the baseline. It also explains why this matters more than it did two years ago: 71 percent of B2B buyers now use AI chatbots in their buying process, so the answer you just read is the one your prospect is reading. Our free AI visibility check will run a pass for you if you would rather not do it by hand.

Step 3: Demand stored answers, not a score

This is where most agency reporting falls apart. Ask for, per question, per engine, per date: the full answer text, and the list of URLs that answer cited.

Why the full text matters: an answer that says "companies like Acme and Northwind are options in this space" is a mention. An answer that never says your name but lists your homepage as a footnote source is a citation, not a mention. Those are different outcomes with different value, and most tools blur them together into one inflated number. Ask your agency which one they are counting. If they cannot tell you, they are counting whichever is bigger.

Why the source links matter: they tell you why a competitor keeps getting named. If the same three review pages and roundups show up behind every answer in your category, that is a target list for off-site work, not a mystery. It is also how you find out which competitor is winning AI citations and on what criteria buyers are being told to compare you.

There is no industry-standard visibility score. Any composite number is your vendor's own math, and a number nobody can audit is not proof. We will not invent one either.

Step 4: Check the traffic in your own analytics, not theirs

Segment sessions by referrer for the four engines and look at the trend against the date the agency started. Two warnings before you read that chart.

First, the engines increasingly strip or obscure referral data, so a real ChatGPT visit can land in your reports looking like direct traffic. Your AI numbers are almost certainly undercounted, which is why so much AI traffic looks like direct. Ask whether the agency is tagging the pages it publishes so those visits can be resolved.

Second, sessions are not the finish line, but do not shrug at a small number either. Semrush measured 4.4 times higher conversion from LLM-referred traffic than from organic search, which means 200 visits from an assistant can outrun a few thousand from a search results page. So ask the harder question: which of those sessions fired a lead or purchase event, and which landing page they arrived on. If the agency's published pages are not in that list, their content is not what is earning you anything.

Step 5: Follow one answer all the way to a dollar

The complete receipt looks like this: a specific buyer question, an engine answer naming you on a dated run, a landing page that answer pointed to, a session in your analytics, and a conversion with a dollar amount attached.

That chain is the whole test. Bosten Shoes is the version we can name. Its founder, Rodrigo, started out unmentioned: buyers asking about leather shoes in El Salvador were getting answers that named other companies. Answer-first content went after those exact questions, the questions stayed monitored, and Bosten went from unmentioned to the number-one recommended leather shoe brand in El Salvador. The first sale traced back to ChatGPT landed attributed to revenue, not buried in direct traffic, in under 30 days.

What should make you nervous

  • Guaranteed placement. Nobody controls model outputs, including us. A vendor promising a spot in an AI answer is selling something they cannot deliver.
  • Reporting that changes shape month to month. Same questions, same engines, same metrics, or it is not a trend.
  • Mentions with no revenue line. Visibility that never produces a session or a sale is a vanity metric with better branding.
  • Refusal to hand over raw answers. The stored answer and its sources cost nothing to share. Reluctance is the tell.

Give it three to six months before you judge the content program, but judge the measurement in week one. The reporting either exists on day one or it does not.

Everyone shows you rank. The question worth asking your agency is whether they can show you revenue. If you want the baseline and the receipts running continuously instead of arriving as a monthly PDF, that is what Clark, our AI search specialist, does: it asks your questions on ChatGPT, Claude, Gemini, and Perplexity every day, records which companies each answer names and where you rank, and ties the sessions those answers produce back to the exact engine and post that earned them.

Written by the Clarity Search AI team.

AEOAI visibilityChatGPTAgencyAnswer Engine Optimization

Is AI recommending you?

See where your brand shows up across ChatGPT, Claude, Gemini, and Perplexity, and win back the customers AI is sending to your competitors.

Check my visibility

Win the customers
sent by
ChatGPT

Stop measuring AI visibility. Start turning it into leads and revenue you can prove.