Agencies: how to report AI visibility to clients without overpromising
Clients are asking what you are doing about ChatGPT. Here is a reporting framework for agencies: the KPIs that hold up, a sensible cadence, white-label delivery, and the language that avoids promising what nobody can guarantee.
By Seoptist Team · · 6 min read
Every agency has had the email: "Our MD asked ChatGPT who the best [whatever we do] in [our town] is and we weren't on the list. What are we doing about this?" The temptation is to answer with a promise. The better answer is a measurement, a plan and a report format that lets you show progress without claiming control you do not have.
This post sets out a framework we see working for agencies reporting to several clients: what to measure, how often, how to present it, and what to say when the numbers go the wrong way.
Set expectations in the first conversation
Three things to say before you sell anything:
- Nobody can guarantee a mention. AI answers are sampled, vary by engine and by run, and depend on sources you do not control. Anyone promising "we will get you into ChatGPT" is overpromising. What you can promise is to measure, to fix the things that are demonstrably in the way, and to show the trend.
- Most of the work is SEO you already do, applied more carefully. Crawler access, entity consistency, structured data, answer-shaped content. The client is not buying a new discipline so much as a new measurement and a sharper checklist.
- It takes weeks to months. Retrieval-based answers move within weeks of sources being corrected. Training-based beliefs move on the model provider's schedule. Report monthly, judge quarterly.
Put all three in the proposal. They protect you later.
The KPIs that hold up
Keep the headline set small and define each one in the report footer so a client who forwards it to their board is not asked questions you then have to answer.
Visibility (per engine): proportion of sampled answers that mention the client. Define the prompt set and sample count.
Share of voice (per engine): client mentions divided by client plus tracked competitor mentions. This is the chart clients understand instantly because it is market share.
Citation rate: proportion of answers that cite the client's own domain. The one metric fully under your control, and the clearest evidence that content work is landing.
Accuracy issues: count of answers that got a fact wrong (price, address, hours, product), open and resolved. Clients care about this more than about position.
Checklist progress: items completed, by category, and the health score. This is what you did, as distinct from what happened.
Secondary, for the appendix: average position in list-style answers, sentiment distribution, top cited third-party domains, AI referral sessions from GA4.
Leave out anything you cannot define in one sentence.
Cadence
- Weekly: data collection (sampled runs per engine). Not a report. Internal alerts only, for accuracy issues and sudden drops.
- Monthly: the client report. Four weekly points per engine on each chart, commentary on movement, checklist progress, next month's three priorities.
- Quarterly: trend review against the baseline, prompt set refresh, competitor set review, an honest paragraph on what did and did not work.
Weekly client reports in this channel are a mistake. The noise between individual weeks is larger than the signal, and you will spend the meeting explaining sampling.
Structure of the monthly report
1. Summary (five lines): headline SoV change, biggest win, biggest issue, what we did, what is next
2. Share of voice by engine, last 13 weeks, with two nearest competitors
3. Visibility and citation rate by engine
4. Accuracy issues: new, resolved, open (with the wrong fact and the source it came from)
5. Checklist: completed this month, in progress, top three for next month, health score trend
6. AI referral traffic from GA4 (with the caveat that it is a lower bound)
7. Notes on method: prompt count, engines, samples per prompt, data sources
Section 7 is not optional. It is what makes the rest credible, and it is where you state that answers are collected through the engines' official APIs and licensed data providers rather than by scraping consumer apps, which matters to clients with compliance teams.
Language that avoids overpromising
Some substitutions worth adopting across the team:
| Avoid | Use instead |
|---|---|
| "We got you into ChatGPT" | "Your ChatGPT visibility rose from 22% to 41% over the quarter" |
| "You now rank #1 in AI" | "You were the first brand named in 60% of list-style answers on Perplexity" |
| "AI traffic is up 300%" | "AI referral sessions rose from 12 to 48; this is a lower bound as most AI answers do not generate a click" |
| "Guaranteed results" | "We fix what is demonstrably blocking you and measure what changes" |
| "ChatGPT prefers our content" | "Three of your pages were cited this month, up from one" |
When numbers fall, say so, say why if you know, and say what you are checking if you do not. A share of voice drop is often a competitor appearing, not the client disappearing; show both lines.
White-label delivery
Clients want the report to look like it came from you. Practical requirements:
- Your logo and colours on the PDF and on any client-facing dashboard.
- A client-viewer login that shows only that client's data, with no access to your other accounts.
- Separate workspaces per client so prompt sets, competitors and fact cards do not bleed into each other.
- An export (PDF, CSV) you can drop into your own reporting pack if you already have one.
Seoptist's Agency plan (£399 a month, ex VAT) is built for this: up to 20 client workspaces, 300 tracked prompts across them, all engines, white-label reports with your logo and colours, client-viewer logins, and Slack or Teams alerts for accuracy issues. Larger agencies and those needing daily tracking or custom data retention use Enterprise. Details are on the pricing page.
Pricing the service
Agencies typically bundle AI visibility into an existing SEO retainer rather than selling it separately, for two reasons: the work overlaps, and a standalone line item invites the "guarantee" conversation. A reasonable structure is a baseline audit and prompt-set workshop as a one-off, then monthly tracking and checklist execution inside the retainer, with the tool cost passed through or absorbed depending on your margin model.
A first-month plan for a new client
Week 1: Fact card, prompt set (30-100 prompts), competitor set, baseline run, audit
Week 2: Fix crawler access, structured data errors, NAP mismatches (quick wins)
Week 3: Rewrite the top three service pages with answer-first paragraphs
Week 4: Correct third-party listings; first monthly report with baseline and actions
Clients who see a baseline, a list of concrete blockers and a fixed list of actions in the first report rarely ask for guarantees. They ask for the next report.
Frequently asked questions
How many prompts should I track per client?
Thirty is a workable minimum for a single-location service business; a hundred for a multi-service or e-commerce brand. Cover discovery, comparison, local and "is X any good" intents, and refresh the set quarterly.
What if a client wants a weekly report?
Offer a weekly alert for accuracy issues and major drops, and keep the full report monthly. Explain that weekly visibility numbers move for sampling reasons and that a monthly view is more honest.
Should I report AI visibility and Google rankings together?
Yes, in the same document, because the fixes overlap and the client experiences them as one channel. Keep the charts separate so nobody reads a share of voice figure as a ranking.