track brand mentions in ai, AI SEO, share of voice, LLM monitoring, answer engine optimization

How to Track Brand Mentions in AI Step by Step

Written by LLMrefs TeamLast updated September 24, 2026

A 2025 Semrush analysis of one million non-branded queries across five AI engines found brand mentions in 26% to 39% of AI responses, depending on the engine. That range changes how SEO teams should think about visibility. Your brand may be present in a meaningful share of answers, yet absent from another engine, region, or buyer scenario.

To track brand mentions in AI reliably, you need more than a few screenshots. Build a recurring measurement system that records mentions, citations, position, competitors, source domains, language, and geography under consistent conditions. The objective isn't to create an impressive dashboard. It's to identify where AI recommends your brand, where it overlooks you, and which content or third-party sources could change that outcome.

Why Tracking Brand Mentions in AI Matters Now

A brand mention in an AI answer is any appearance of your company, product, service, or domain, whether the response links to you or not. That distinction matters because AI systems can name a brand while citing another website, giving you visibility without a directly attributable visit.

The difference isn't theoretical. A 2025 Ahrefs study based on more than 31,000 brand mentions from a database of 150 million prompts found that AI systems frequently mention brands without linking to them. The reported unlinked mention rates varied by engine, from 10.7% in Google AI Overviews to 51.6% in Perplexity, and the study found that brand mentions appeared without a link 50% to 90% of the time on average.

A one-off prompt check can't tell you whether a result is durable. AI responses vary as models change, retrieval sources rotate, competitors publish new material, and prompt wording shifts. A screenshot can confirm that your brand appeared once, but it can't establish a trend or explain whether a competitor is consistently displacing you.

The three signals worth tracking

A useful program should expose three business signals:

  • AI share of voice: How often your brand appears compared with competitors across the same prompt set.
  • Citation source concentration: Which domains AI engines rely on when answering questions in your category.
  • Competitive displacement: Which rivals appear in prompts where your brand is absent, especially on commercial or comparison questions.

These signals connect AI visibility to practical SEO work. If competitors repeatedly appear beside a particular review site, that site becomes a potential outreach or partnership target. If your brand is named but your domain isn't cited, your reputation and content distribution may be working on different tracks.

For broader context on how structured monitoring supports reputation work, the guide to the business benefits of social listening is useful because it treats mentions as signals to interpret, not merely counts to collect.

Practical rule: Treat AI mention tracking as a weekly measurement loop. Don't treat it as a quarterly audit.

Picking Keywords, Geos, and Languages for AI Tracking

Start with buyer questions, not the largest head terms in your SEO platform. Export your existing keyword list, remove branded queries, and sort the remaining terms into three intent buckets:

  1. Problem-aware: The buyer describes a difficulty, such as “how can a SaaS team reduce onboarding friction?”
  2. Solution-aware: The buyer knows the category but is exploring approaches, such as “customer onboarding software for a remote sales team.”
  3. Comparison-aware: The buyer is choosing among vendors, such as “alternatives to spreadsheet-based onboarding.”

These buckets produce different answer patterns. Problem-aware prompts reveal which brands AI associates with expertise. Solution-aware prompts show category inclusion. Comparison-aware prompts expose recommendation order and competitor displacement.

For geo selection, begin with the countries where revenue or pipeline matters most. Then add expansion markets where competitors are already visible. Don't assume one English-language prompt represents every market. A buyer in the United Kingdom may use different terminology from a buyer in the United States, while a French-language response may draw on a different source mix entirely.

Pair every location with its working buyer language. Machine-translated prompts often sound unnatural and can produce results that don't reflect local search behavior. Keep the same intent and decision criteria across languages, but let native speakers adapt the phrasing.

A keyword research workflow such as Keyword Kick can help turn a broad list into a more usable set of topic and intent inputs before you build the tracking configuration.

A clean configuration

Intent Bucket Seed Keyword Example Target Geo Language
Problem-aware reduce SaaS customer churn United States English
Solution-aware customer retention platform United Kingdom English
Comparison-aware customer retention tools for mid-market teams Germany German
Solution-aware logiciel de fidélisation client France French

Your configuration should also specify the engine list, such as ChatGPT, Perplexity, Gemini, Claude, and Copilot, along with the run cadence. Keep geo and language paired in every record. If you mix locations or languages, you won't know whether a visibility change came from your content, the engine, or a sampling inconsistency.

Generating Conversation-Based Prompts That Reflect Real Buyers

Keyword strings are convenient for SEO systems, but they underrepresent how people use AI assistants. A buyer usually supplies context, asks for a decision, and adds a constraint. Build prompts around those three components:

  • Context: Who is asking and what situation are they in?
  • Intent: What decision do they need to make?
  • Constraint: What limits the answer, such as geography, budget, team size, or software stack?

Suppose your seed keyword is “customer retention software.” You could generate a prompt set like this:

  1. “I run a growing subscription business and our customers are leaving after the first renewal. What types of tools can help diagnose and reduce churn?”
  2. “Which customer retention platforms are suitable for a mid-market SaaS company with a small lifecycle marketing team?”
  3. “What should a European SaaS company compare when evaluating customer retention software that integrates with Salesforce?”
  4. “Compare the strongest customer retention tools for a B2B company that needs behavioral segmentation and clear reporting.”

Each prompt asks one decision question. Avoid compound clauses such as “Which tool is cheapest, easiest to integrate, and best for enterprise reporting?” Split that into separate prompts or you'll struggle to interpret the answer.

Prompt hygiene controls

Keep the brand name out of the prompt body when measuring unaided visibility. Use consistent phrasing across geographies, while adapting idioms naturally for each language. Assign every prompt a stable ID, intent bucket, locale, and version number.

Start with 30 to 50 prompts per geo. That gives you enough coverage to identify patterns without creating an unmanageable review queue. Expand only after you can explain which prompts produce useful commercial signals.

Prompt quality is the largest variable in tracking accuracy. If your prompt set contains generic, artificial, or overly branded questions, the dashboard may look precise while measuring behavior no real buyer exhibits.

For a practical way to turn seed terms into natural questions, use the AI prompt generation guide. Version prompts whenever you alter wording, constraints, or intent. Otherwise, a change in the prompt library can look like a visibility improvement caused by a content update.

Keep the metrics separate

For every response, record at least three distinct outcomes:

  • Mention: Was the brand named?
  • Citation: Did the response link to the brand's domain or supporting content?
  • Rank: Where did the brand appear relative to other named brands?

A brand can score well on one and poorly on the others. The Ahrefs findings above show why mention rate and citation rate shouldn't be combined into a single unexplained number.

Aggregating Responses, Mentions, and Citations

Capture the raw response before you extract metrics. The raw text is your audit trail, especially when a parser misses a product name, misreads a citation, or classifies a neutral statement as a recommendation.

Every run should preserve three layers:

  1. Raw response text, including the prompt and answer.
  2. Detected brand and product mentions, including the exact string used by the engine.
  3. Cited URLs, meaning the sources surfaced in the answer.

Create a brand alias map before counting. “HubSpot,” “Hub Spot,” and “hubspot.com” should resolve to one entity, while a product line may need its own classification. Run a separate pass for parent-brand and product-name detection so a response listing several products doesn't inflate the company's total mention count.

Tag each record with the engine, prompt ID, locale, date, and mention position. Store both linked and unlinked mentions. Also record two useful exceptions:

  • Mention without citation: The AI names your brand but links elsewhere.
  • Citation without mention: The answer cites your content or domain but doesn't name your brand.

That second category matters because source visibility can precede explicit brand visibility. It also helps content teams find pages that influence answers even when the company name isn't visible in the generated text.

Metric What It Captures Answers
Mention Brand or product appears in the answer Is the brand part of the conversation?
Citation A URL is surfaced as supporting evidence Does the answer provide a clickable source?
Rank Position among named brands How prominently is the brand presented?

A simple worked calculation

Imagine a prompt batch produces 18 mentions of your brand across 60 tracked prompts, alongside 142 competitor mentions. Your share of voice is:

18 ÷ (18 + 142) × 100 = 11.25%, which rounds to 11.2%.

If your brand positions across the responses where it appears average 2.4, your mean position is 2.4, with 1.0 as the ideal. Keep absent results as null rather than assigning them a low rank. Otherwise, you blur the difference between “not present” and “present but poorly positioned.”

The brand monitoring framework for AI results provides a useful reference for organizing these fields. The important operational choice is to retain the raw response alongside the normalized record. Aggregated metrics tell you that something changed. Raw answers help you determine what changed and why.

Computing Share of Voice and an Aggregated AI Rank

Share of voice is straightforward only when the denominator is clear. For each engine, calculate your brand's mentions divided by all tracked brand mentions in the same prompt set. Then calculate an aggregate across engines, geographies, and intent buckets.

Don't let the engine with the largest response volume dominate the headline metric. Use equal weighting when each engine represents an intentional monitoring surface, or apply a documented weight when business evidence shows that some engines matter more to your audience.

Rank needs similar discipline. If your brand appears first among five brands, assign position 1. If it appears last, assign 5. If it doesn't appear, record null and exclude it from the mean position calculation. Report presence separately, so a lower average rank can't hide a declining mention rate.

A four-step infographic explaining how to calculate AI share of voice and brand ranking metrics.

A weekly workflow inside LLMrefs

A practical run starts with the same prompt IDs used in prior weeks. Trigger the engines, collect the responses, normalize aliases, and review the resulting mention and citation records. Then export the data for analysis rather than relying only on the dashboard headline.

For example, a team might review the overall share of voice, then filter by “comparison-aware” prompts in Germany. If the brand's aggregate number is stable but German comparison visibility falls, the next action is regional, not a site-wide content rewrite.

The share of voice measurement guide is helpful when you need to formalize denominator rules, competitor sets, and reporting views. Build rollups at four levels: prompt, topic cluster, locale, and total footprint. That structure keeps a small geo spike from distorting the company-wide view.

Measurement discipline: A missing mention should remain missing. Never convert absence into an artificial rank.

Use the rank output to prioritize work. A brand appearing second on a high-value comparison prompt may need better differentiation, while a brand appearing fifth may need stronger category evidence and clearer entity signals. The metric doesn't prescribe the fix, but it points you toward the right diagnostic question.

Running the Workflow Weekly With LLMrefs and Exporting Results

A recurring operating rhythm keeps the data usable:

  • Monday: Refresh the prompt list, review version changes, and update brand aliases.
  • Tuesday: Trigger a full run across ChatGPT, Perplexity, Gemini, and Copilot.
  • Wednesday: Validate a 10% sample by hand to catch parser regressions, unexpected answer formats, and missed product aliases.
  • Thursday: Calculate share of voice, mention rate, citation rate, and average AI rank.
  • Friday: Send a concise Slack digest and email a CSV to stakeholders.

Place the schedule where the team can inspect it before the run begins.

Screenshot from https://llmrefs.com/dashboard/weekly-run.png

The exports should match different users. SEO leads need the flattened mentions CSV, content teams need prompt-level answers and cited URLs, and outreach teams need a citations-only CSV that groups third-party sources by topic and competitor inclusion. Per-prompt JSON is useful for engineering and audit logs because it preserves the full response structure.

An API connection can trigger scheduled runs and pull the latest share-of-voice data into a Looker Studio or Notion dashboard. Keep the source data immutable, then calculate reporting views from a separate layer. That makes historical comparisons possible when your scoring logic changes.

Alerts that deserve attention

Set alerts around business meaning, not every fluctuation:

  • Share of voice falls more than 2 points week over week.
  • A top-three prompt loses all brand mentions.
  • A new competitor enters the cited-source set.

These conditions should open a review task, not trigger an automatic content rewrite. First inspect whether the prompt version, locale, engine behavior, or parser changed.

Before locking the weekly run, confirm that the engine list is unchanged, prompt versions are documented, aliases include recent product names, locales are paired correctly, and the competitor set hasn't changed without approval. Also check a sample of raw responses. Clean data is more valuable than a larger data volume.

Benchmarking Against Competitors and Troubleshooting Common Traps

A competitor benchmark should compare the same prompt set, engines, locales, and time window. Select three to five direct competitors, then record their mention rate, share of voice, average AI rank, and strongest cited source. A dashboard that shows your brand improving without showing who gained or lost around you is incomplete.

Brand Share of Voice Mention Rate Avg. AI Rank Top Cited Source
Your brand 11.2% 24% 2.4 Industry publication
Competitor A 18.6% 31% 1.8 Review platform
Competitor B 14.1% 27% 2.1 Specialist forum
Competitor C 9.7% 19% 3.0 Competitor domain

The values in this example are illustrative calculations, not market benchmarks. Your reporting system should populate them from your own fixed prompt set.

Use brand stature as a baseline lens

AI visibility isn't distributed evenly across brand sizes. An arXiv paper on measuring brand visibility across AI search engines reported that global household brands appeared in 73% of relevant answers, established mid-market and regional brands in 44%, and niche or small brands in 11% on first visibility runs.

That ladder changes how you interpret a gap. A smaller brand shouldn't compare only with a global incumbent and conclude that every missing mention reflects a content failure. Benchmark within the relevant category, market, and maturity tier, then target the structural reasons for exclusion, such as weak entity clarity, limited third-party coverage, or absent comparison content.

For category and customer-feedback context, research into leading Clarabridge competitors can help expand the competitor set and reveal the sources buyers use when evaluating adjacent tools.

Diagnose before optimizing

When rank drops suddenly, check:

  • Sampling: Did the prompt wording, locale, or engine change?
  • Entity resolution: Did the model confuse your brand with another company or product?
  • Source mix: Did a previously cited review, forum, or publication disappear?
  • Competitive content: Did a rival publish a comparison page or earn new third-party coverage?
  • Framing: Is the brand still mentioned, but described less clearly or less favorably?

The most common traps are predictable. One-off snapshots create false positives, unlinked mentions get counted as citations, geo variance disappears inside global averages, and teams overreact to temporary spikes. Recurring sampling solves the first and fourth problems, while separate mention and citation fields solve the second.

The optimization moves should follow the evidence. Refresh sources that AI already cites, fill prompt gaps with direct answers, improve entity disambiguation, and pursue credible third-party coverage where competitors dominate. Over the next quarter, the advantage comes from repeating the same measurement and action cycle until the team can distinguish durable movement from noise.


LLMrefs helps teams generate conversation-based prompts from keywords, monitor brand mentions and citations across AI answer engines, and compare share of voice and aggregated position by competitor, geography, and language. Set up a recurring tracking workflow, inspect the cited sources behind each result, and visit LLMrefs to start measuring where your brand appears and where it needs stronger visibility.

How to Track Brand Mentions in AI Step by Step - LLMrefs