A brand can win every AI answer in its category and still have no control over what the answer says. That is what a GEO audit I ran this month on a UK women’s activewear brand showed. Across twelve buyer prompts on Claude, the brand was mentioned in all twelve, sat in the top three names every time, and led share of voice at 32 percent against rivals with far bigger budgets. It was cited as a source in two answers. The other ten were assembled from magazines, comparison blogs and a forum thread, and the story those pages told was consistent and not one the brand chose.
I am not naming the brand because this was research, not an engagement, and it would be unfair to publish a diagnosis of a company that did not ask for one. The method, the numbers and the fixes are all here, because the pattern applies to almost any consumer brand. I built the audit as a Claude skill after eight years of doing the equivalent work for Google rankings, and this was its first full run.
In numbers
- 12 buyer prompts, one engine, one afternoon
- 100 percent mention rate, 100 percent recommendation rate, 32 percent share of voice
- Only 17 percent of citations from the brand's own domain; 41 of 52 were editorial
- The sizing answer came from four third-party sites because the help centre renders no text
- Eight fixes, ordered by evidence and effort; re-run in 90 days and diff the scores
How the audit was built
The audit is a fixed prompt set run through an engine with search on, with every answer and every cited URL logged. Twelve prompts across five intent buckets: discovery (“best gym leggings uk”), comparison (“brand A vs brand B which is better”), brand (“is brand A good quality”), use case (“what to wear for reformer pilates uk”), and logistics (“does brand A run true to size”). Four competitors were fixed before the runs so share of voice had a denominator.
Each prompt ran in a fresh session with web search on. For every answer I recorded whether the brand was mentioned, its rank among the brands named, whether its own domain was cited, the full brand order, and every cited domain. That produced a run log and a scoreboard. The same set is written so it can be repeated on ChatGPT, Perplexity and Google AI Mode from a UK browser in about fifteen minutes per engine, and the scores diff against this edition.
The scoreboard
On Claude the brand scored 100 percent mention rate, 100 percent recommendation rate, 17 percent own-domain citation rate, and 32 percent share of voice. Share of voice across the set: the brand 12 mentions, the largest premium rival 9, the largest performance rival 7, a UK heritage brand 7, a smaller UK label 3. Mention by bucket was perfect: discovery 3 of 3, comparison 3 of 3, brand 2 of 2, use case 2 of 2, logistics 2 of 2.
The number that mattered was the 17 percent. Of 52 citations across the twelve answers, 41 were editorial (two fashion magazines carried ten between them), 5 were the brand’s own domain, 3 were community (a parenting forum and TikTok), 2 were a review platform, and 1 was a competitor’s collection page. The engine trusted editors, not the brand.
What the engine actually said
The engine repeated the same split on every comparison: the brand owns “everyday”, the premium rival owns “training”. One answer opened “They are built for different jobs” and closed with buy the rival for hard gym sessions and this brand for everyday and low-impact wear. The brand prompt carried the same caveat: pieces built for lower-impact workouts, some leggings ride up when running. Every one of those lines was sourced from a magazine or a comparison blog. Nothing on the brand’s own site argued the other side, because nothing on the brand’s site was being read for those prompts.
The sizing question was worse. Asked whether the brand runs true to size, the engine answered “runs small in the waistband, size up if between sizes”, sourced from four third-party sites and a forum thread. The brand’s own sizing article was fetched during the audit and returned only a title and meta tags. The body is rendered by JavaScript inside a help-centre widget, so a text fetcher never sees it. The returns answer was correct only because a plain HTML returns page exists. Sizing has no such page, so the brand has no voice on the question that decides its returns volume.
Want this done for your business?
An audit, a build, or a morning report on your own accounts. Tell me what you are trying to do and I reply myself within one working day.
The eight fixes
The fix list is ordered by evidence and effort, and the first three are the ones that move the number.
- Publish a plain-HTML sizing page with per-product fit notes, the “size up if between sizes” guidance in the brand’s own words, and FAQPage schema. Link it from every product page. Own-domain citation on logistics prompts goes from one of two to two of two.
- Move every help-centre article that carries a buying fact (sizing, returns, delivery, fabric care) out of the JavaScript-only widget into static pages, or server-render the help centre. Verify with curl that the body text returns without JavaScript.
- Add a “training or everyday?” section to the leggings collection page and the hero product page: fabric weight, compression, what each legging is designed for, with the brand’s own evidence for high-impact use where it exists. Crawlable HTML, not an accordion that renders empty.
- Build one owned comparison page per top rival, honest on both sides, with a fit and price table. The comparison prompts are currently answered entirely by two blogs, and these pages are the exact format the engine already cites.
- Product schema audit: Product, Offer, AggregateRating, plus MerchantReturnPolicy and OfferShippingDetails so returns and delivery facts are machine-readable per product.
- Digital PR against the source map: keep the two magazine positions live by sending new-season samples before their re-test cycles, and pitch the comparison-blog tier with fit data and a “for the gym” story.
- A community programme that seeds real sizing and “worth it?” threads through customers and ambassadors, brand-identified where the rules allow. No fake accounts.
- A robots.txt decision on GPTBot, ClaudeBot, PerplexityBot and Google-Extended, made deliberately and written down, because it is the prerequisite for the other three engines reading any of the above.
The 90-day plan runs fixes 1, 2, 5 and 8 in weeks one and two, fixes 3 and 4 in weeks three to eight, and fix 7 in weeks nine to twelve, then re-runs the same twelve prompts on all four engines. The next review opens with two numbers: own-domain citation rate, target 17 percent to 50 percent, and comparison-bucket own citations, target zero of three to two of three.
What this means for any brand
Visibility and control are different scores, and most GEO conversations only measure the first. A brand that is named everywhere and cited nowhere is renting its reputation from whichever magazine last ran a round-up. If that magazine re-tests and drops the brand three places, the discovery bucket drops with it.
The cheap fixes are almost always the same: a buying fact locked in a help widget, a comparison the brand never wrote, a sizing or pricing page that does not exist as HTML. None of them needs a new tool. They need someone to read the answers, trace the citations, and write the pages the engine is looking for.
Limitations, stated
This edition ran one engine of four, from a cloud session rather than a UK IP, one run per prompt. The comparison framing is decisive enough to act on, but I would re-run it three times before a client relied on it. The prompt set and scoring are fixed so the other three engines produce comparable numbers, and the run sheet for them is a fifteen-minute job each.
Key takeaways
- Mention rate and citation rate are different numbers; this brand scored 100 percent on one and 17 percent on the other.
- Forty-one of 52 citations were editorial. The engine trusts editors when the brand gives it nothing to read.
- Buying facts inside JavaScript help widgets are invisible to answer engines; test with curl.
- Owned comparison pages and a plain sizing page are the highest-return fixes, and neither needs a new tool.
- Fix, wait 90 days, re-run the same prompts, diff the scores. Anything else is guessing.
If you want to know what the engines say about your brand, ask me to run the same twelve prompts. The prompt set and the scoring sheet are in the free toolkit.