We Tracked Our Own Brand Across 4 AI Engines for 30 Days: 91 on Branded Prompts, 7 on Everything Else
We run an AI visibility platform. So before writing another guide about being visible in AI answers, we pointed the product at ourselves and left it running for 30 days.
The headline number is 34% visibility. That number is useless. Split into its two halves it becomes the most useful thing we have measured all year:
| Prompt type | Measurements | Average visibility |
|---|---|---|
| Branded (prompts naming us) | 384 | 91 / 100 |
| Unbranded (category questions) | 832 | 7 / 100 |
Ask any of the four major engines about us directly and they answer accurately, in detail, with near-perfect consistency. Ask them the question a buyer actually asks — best AI visibility software for B2B SaaS — and we are, to a very close approximation, not there.

The method, stated plainly
- Window: the 30 days to 14 September 2026
- Engines: ChatGPT, Gemini, Perplexity, Claude
- Prompts: 38 distinct, split across branded and unbranded intent
- Measurements: 24–28 repeat runs per prompt, 1,216 prompt–engine data points total
- Recorded per answer: mention, prominence, sentiment, competitors named, sources cited
The repetition matters more than the count. A single AI answer is a coin flip; 28 consecutive runs across four engines is a finding. Everything below survived that repetition.
Twelve prompts scored exactly zero
Not “low.” Zero — no mention, in every run, on every engine, for 30 days.
Here is the part that should bother anyone running a content programme, and the reason we are publishing this at all. For eight of those twelve prompts, we had already published a dedicated article targeting that exact question.
| Prompt (0.0 visibility, 100% zero-mention) | Our published article on it |
|---|---|
| How to track brand mentions in ChatGPT and Perplexity | Brand mention tracking guide |
| How to fix negative brand sentiment in ChatGPT responses 2026 | Fix negative brand sentiment in AI |
| Best platforms for auditing brand visibility in LLM results | Best AI visibility tools 2026 |
| Why is my brand not showing up in AI search | Brand visibility in ChatGPT |
| Top rated AEO platforms for marketing agencies | Best AI search optimization agencies |
| Best GEO software for B2B SaaS 2026 | AI SEO software tools guide |
| Measuring AI search visibility for multi-location businesses | Franchise & multi-unit AI search |
| Automated tools to track brand mentions in Claude and Gemini | Claude, Gemini & DeepSeek tracking |
Eight articles, written specifically for those questions, indexed and live. Eight zeros.

So what actually failed
We can rule out most of the usual suspects with the data itself.
It is not a crawl or index problem. An engine that scores us 91 on branded prompts — describing our features, pricing and positioning accurately — has plainly read the site. You cannot be simultaneously unreadable and accurately summarized.
It is not a content-quality problem in the way people mean it. Those eight articles are thorough, structured and answer-first. They rank. They just do not get retrieved into an answer where the model has to choose between vendors.
It is a corroboration problem. When the model answers “best AI visibility software for B2B SaaS”, it is not summarizing vendor websites. It is synthesizing from the sources it already trusts to adjudicate that category — review platforms, comparison posts, roundups, community threads. On those sources, our competitors are named and we largely are not. Our own article arguing we belong in the category is, from the model’s perspective, the least credible possible witness.
This is the uncomfortable structural fact of AEO: your own content makes you eligible; other people’s content gets you chosen.

The sentiment number nobody wants
Across every mention we recorded:
| Sentiment | Count |
|---|---|
| Positive | 273 |
| Neutral | 803 |
| Negative | 0 |
Zero negative mentions reads like a clean bill of health. It mostly is not. Neutral at that ratio means the engines state facts about us when asked and almost never advocate. Negative sentiment is a reputation problem with a known fix — find the source, displace it. Overwhelming neutrality is a positioning problem: nothing the engines have read gives them a reason to prefer us in a comparison.
If your sentiment report is 70% neutral, you do not have a sentiment problem. You have an advocacy vacuum.
What we are changing
Publishing this costs us something, so here is what it bought us — a plan that is not “write more articles.”
1. Source-first, not content-first. For every zero-scoring prompt we now log the URLs the engines cited instead of us, ranked by frequency. That list — review sites, comparison pages, a handful of community threads — is the work queue. It is a PR and placement roadmap, pre-ranked by how much each engine already trusts the source.
2. Stop counting aggregate visibility. We report branded and unbranded separately from now on. A composite score is a number that averages a solved problem with an unsolved one and tells you nothing about either.
3. Track exclusions, not presence. The most actionable number in the dataset is not 34%. It is the twelve prompts where competitors appear and we do not. That converts directly into assignments; a percentage does not.
4. Prune before publishing. Eight articles that produced zero citations are evidence that volume is not the lever. We would rather fix the corroboration under five pages than publish five more.
How to run this on your own brand
For the other side of this — what it looks like when the work pays off — see our SEO + LLM visibility case study: 0 to 1,000+ organic clicks a day in six months.
Reproduce it — the method is not proprietary.
- Write two prompt lists. Branded: “is [you] worth it”, “[you] vs [competitor]”, “[you] reviews”. Unbranded: “best [category] for [use case]”, “how do I [problem]”, “top [category] tools 2026”. Fifteen minimum in each.
- Control the session. Logged out or memory off, location fixed. Checking from an account that has discussed your company forty times measures your history, not the model.
- Repeat, do not sample. Same prompts, same conditions, weekly. One run is trivia.
- Record four columns per answer: named or not, prominence, competitors named, sources cited. The fourth is the one everybody skips and the only one that tells you what to do next.
- Report the split. Branded and unbranded, separately, every time.
If your branded score is high and your unbranded score is near zero, you are where we are: known to the engines, absent from the decisions. That is a fixable problem, but not by writing about yourself.
Run a free AI visibility audit to see your own split in about two minutes.
The bottom line
We measured 1,216 AI answers about our own category and found a 13x gap between how well the engines know us and how often they recommend us. Eight dedicated articles on our eight worst prompts moved nothing.
The lesson we are taking, and publishing against our own interest: in AI search, being the best-informed source about yourself is worth almost nothing. The engines were never going to take your word for it.
Frequently Asked Questions
A branded prompt names your company — 'is Sanbi.ai worth it', 'Sanbi.ai vs Profound', 'reviews of Sanbi.ai'. An unbranded prompt describes the problem or category without naming anyone — 'best AI visibility software for B2B SaaS', 'how do I track brand mentions in ChatGPT'. Branded prompts are asked by people who already know you exist; unbranded prompts are where new buyers are acquired. In our own 30-day measurement the two scored 91 and 7 out of 100 respectively, which is the gap most brands never see because they only ever check the branded ones.
In our data the most common reason is not a technical block — it is that the engines have no third-party corroboration for your claim to the category. Every one of our twelve zero-scoring prompts was an unbranded buyer question, and for eight of them we had already published a dedicated article on that exact question. The article ranked; the answer still named someone else. Publishing on your own domain establishes that you talk about a topic. Being named in an answer requires the sources the engine already trusts for that category to talk about you.
There is no published industry benchmark yet, which is part of why we released our own numbers. What matters more than the headline score is the split: our composite was 34%, which sounds middling until you separate it into 91% on branded prompts and 7% on unbranded ones. A brand reporting a single aggregate number is almost always reporting its branded performance with the unbranded failure averaged into invisibility. Measure and report the two separately or the number will mislead you.
More than most people run. AI answers vary between runs, so a single check is a coin flip. We measure 38 distinct prompts at 24-28 repeat measurements each across four engines, which produced 1,216 data points in 30 days. That repetition is what lets you distinguish a real zero — a prompt that returned no mention in 28 consecutive runs across every engine — from an unlucky sample. Below roughly 15 prompts with weekly repeats, you are reading noise.
Not on its own. We published dedicated articles targeting eight of our twelve worst-performing prompts and still scored zero on all of them across ChatGPT, Gemini, Perplexity and Claude. On-domain content makes you eligible and gives the engines something to cite once they decide you belong in the answer. What decides that is third-party corroboration: review sites, comparison posts, community threads and industry coverage that name you in the category. Content is necessary and not sufficient.
In our 30 days we recorded 273 positive mentions, 803 neutral and zero negative. Zero negative sounds like a win and mostly is not — neutral at that volume means the engines describe us factually when asked directly but rarely characterize us as a recommendation. Negative sentiment is a reputation problem you can fix with sources. Overwhelming neutrality is a positioning problem: nothing the engines have read gives them a reason to advocate for you over an alternative.
Build two prompt lists. The branded list names your company in the ways real buyers would — 'is X worth it', 'X vs competitor', 'X reviews'. The unbranded list is the category questions a buyer asks before they know you: 'best [category] for [use case]', 'how do I [problem]', 'top [category] tools 2026'. Run both across every engine on a fixed schedule, from a clean session with location set, and record mention, prominence, competitors named and sources cited. Report the two scores separately. Sanbi automates this, or you can start with a free AI visibility audit to see the split for your brand.
Check the technical layer first, because it is cheap and binary: crawlable, indexed, snippet-eligible, rendering without JavaScript, and AI crawlers getting 200s in your server logs. But a zero on unbranded prompts while scoring 91 on branded ones rules most of that out — an engine that can describe your product accurately when asked directly has clearly read your site. That pattern is a corroboration problem, not a crawl problem.
Treat the cited sources as the to-do list rather than the prompt. For every prompt you lose, record which URLs the engine cited instead. Those pages are the ones the engine already trusts for that category, ranked by how often they appear. Getting named on them — through reviews, comparison inclusion, expert roundups or genuine community presence — moves the answer faster than another post on your own blog. That is the change we made to our own programme after this measurement.
Weekly for a stable prompt set, with a monthly review of the trend. AI answers shift for reasons you do not control — model updates, index refreshes and unannounced retrieval changes, as the August 2026 collapse in Reddit's ChatGPT citations demonstrated. A single measurement tells you almost nothing; four weeks of the same prompts tells you whether anything you did worked.