/interfacer.
FeaturesLong read

What Hits-Out-of-Asks Actually Means for Measuring AI Visibility

A brand mention by AI means nothing without knowing which questions were asked and how many times.

Senior Writer · · 10 min read
Cover illustration for “What Hits-Out-of-Asks Actually Means for Measuring AI Visibility”
Features · August 20, 2026 · 10 min read · 2,212 words

Hits-out-of-asks measures how often an AI answer mentions your brand out of every prompt you tested. That's the whole idea. But the number is only worth reading if you can see what's behind it, because a mention rate without its sample size is just a percentage floating in space, and percentages floating in space can be made to say whatever the person reporting them wants.

Here's why this metric exists at all. Google used to hand users ten blue links and let them pick. Now the AI Overview picks for them, names a brand or two, and the user often never scrolls further. Pew Research Center analyzed 68,879 real Google searches and found that when an AI Overview shows up, the external click rate drops from 15% to 8%, nearly cut in half. Only 1% of users click the links buried inside the Overview itself. More strikingly, 26% of AI Overview searches end the whole session right there, versus 16% when there's no Overview. The person got their answer and left. No click, no visit, nothing to track in your analytics dashboard.

Ahrefs found the same erosion accelerating, not leveling off. As of February 2026, AI Overviews cut click-through for the top-ranking organic result by 58%, up from 34.5% in April 2025. And this isn't some fringe behavior confined to recipe searches or "what time is it in Tokyo." B2B technology queries trigger an AI Overview 82% of the time. Healthcare queries hit 88%. Rank #1 all you want; if the summary doesn't say your name, you're invisible to most of the people who searched.

So the question every brand actually cares about has quietly changed. It used to be "did they click our link?" Now it's "did the AI say our name?" That shift is what makes a mention-based metric necessary rather than a nice-to-have dashboard add-on. Hits-out-of-asks is the plain-language name for the number that answers it.

What hits-out-of-asks actually measures

Strip away the jargon and it's simple arithmetic. Hits are the AI answers that mention your brand. Asks are the total prompts you ran. Divide one by the other, multiply by 100, and you've got your mention rate. Some vendors call it Mention Rate, others call it Brand Visibility Score, but it's the same math wearing different clothes.

Think of it as the click-through rate of the generative-search era. Old CTR asked whether someone clicked your link out of everyone who saw it in search results. This asks whether the AI said your name out of every question you threw at it.

There are two flavors worth telling apart. Unweighted mention rate just counts frequency: does your brand show up, yes or no, across all the runs. Position-weighted AI Visibility Score cares about where you show up too; a mention buried third in the answer counts less than one that leads. You divide each mention's value by its position, add them up, divide by total responses. Tracking both side by side makes sense, because a brand that gets mentioned constantly in passing looks nothing like a brand that gets named first when it counts.

And the raw number hides three very different events that all get lumped under "mention." A mention just means your name showed up somewhere in the text. A citation means the AI linked to your actual domain, which can happen with no mention at all, or a mention with no citation. A recommendation means the AI told the user to go with you, which is a much stronger signal than just being name-dropped. Per Visiblie's framework, if you run 100 prompts and get 35 mentions but only 12 of those are genuine recommendations, your recommendation rate is much lower than your mention rate. For a B2B SaaS company trying to connect this metric to pipeline, that 12% is the number that matters.

Table: What a Trustworthy Mention Rate Discloses. Compares What It Measures, Best For, Key Limitation and When to Prioritize by Unweighted Mention Rate, Position-Weighted AI Visibility Score and Recommendation Rate.

Why the denominator matters more than the number itself

Here's the part almost nobody talks about: the "asks" half of the ratio is the single biggest decision in the whole measurement exercise, and it's usually the part that gets hidden.

Most measurement tools ship with a set of auto-recommended prompts. Convenient, sure. But per Search Engine Journal's reporting, those defaults are rarely the questions your actual buyers are typing into ChatGPT while trying to solve a real problem. A tool's convenience set is not a buyer's question set, and treating them as interchangeable is how brands end up optimizing for the wrong thing entirely.

Graph Digital's 2026 research found that 82% of B2B manufacturing and industrial brands are invisible at the exact moment that matters most: early-stage, when a buyer is describing a problem without naming any vendor yet. That gap exists because most measurement programs skip unbranded discovery prompts and just track head terms with the company name baked in.

A credible prompt set has to span more ground than that. It needs awareness-stage prompts, where the buyer describes a problem with zero vendor names attached. It needs comparison-stage prompts, the "X vs Y" and "alternatives to Z" phrasing. It needs pricing and feature questions, and it needs persona-specific and use-case-specific versions of all of the above.

Why does the mix matter so much? Because per MaxAEO's benchmarks, branded prompts produce a median mention rate around 64%, while comparison prompts land around 18%. Same brand, same underlying visibility, but a rate built from branded prompts alone comes out roughly three times higher than one built from a realistic mix. Nothing about the brand changed. Only the questions changed.

So before you trust any mention rate, ask which prompts and how many. Graph Digital's sizing guidance gives a rough map: simple brand or category checks need around 50 prompts, serious B2B measurement across buyer stages and personas needs 200 to 500 or more, and complex industrial B2B with geography or certification variants needs 500 or more per cycle. Layer3 Labs offers a leaner working range for most teams, 40 to 100 prompts split across three or four intent clusters. Fewer than 40 and week-to-week noise drowns the signal. More than 100 and the tracking cost starts outpacing what the metric actually tells you.

If the denominator isn't visible, the numerator can't be trusted.

Diagram: Why Prompt Mix Triples Your Apparent Mention Rate. Visualizes: Show the dramatic difference in mention rate produced by different prompt types for the same brand.

Why running a prompt once tells you almost nothing

AI models don't give the same answer twice. Ask ChatGPT the exact same question today and tomorrow, and you'll get different brands named, sometimes a different order entirely. This isn't a glitch. It's baked into how the model works.

Under the hood, an LLM predicts the next likely token given everything before it, and sampling techniques like top-k and nucleus sampling deliberately introduce randomness at every single step of generation. Research on LLM output variance backs this up directly. Which means one run of one prompt is a single draw from a probability distribution, not a fact you can hang your strategy on.

MaxAEO's tracking shows just how much this randomness swings the number. A brand sitting near the category median commonly sees its weekly mention rate move up or down 8 to 15 percentage points from model variability alone, before anyone has touched their content strategy. That's not signal. That's the model rolling dice.

A July 2026 arxiv preprint on variance decomposition put actual numbers on this. The reliability of a brand's ranking from a single answer comes out to around 0.01, which is functionally noise. Push the design to its max, testing across 8 languages, 3 models, and 15 paraphrase variants, and reliability climbs to about 0.36. Still not great, but a real improvement. Interestingly, once you've run a prompt five times, running it a sixth barely helps; the paper found each additional repeat past the fifth reduces relative-error variance by only 0.0003. Diminishing returns kick in fast on repetition alone. The better move is spreading your measurement across languages and models rather than hammering the same prompt over and over.

And the engines don't even agree with each other. MaxAEO found that ChatGPT and Perplexity share fewer than 1 in 5 cited domains, meaning a visibility win on one platform frequently doesn't show up at all on the other. The practitioner standard that's emerged from all this: run each prompt three to five times per cycle per engine before you aggregate anything. A single-run result reported as a "rate" is a coin flip wearing a lab coat.

What a trustworthy figure actually discloses

A hits-out-of-asks number is only evidence when it comes with its full paper trail. Strip the context away and it's just a number, and numbers without context can be tuned to say almost anything by adjusting what went into them.

Any figure worth trusting should tell you the total prompt count and how it breaks down by intent (branded, unbranded, comparison, and so on), the number of runs per prompt per engine, which engines got queried, the date range, and whether you're looking at unweighted mention rate or the position-weighted version. Leave any one of those out and you're reading a headline, not a measurement.

There's a statistical point buried in here too: a raw count without its sample size isn't evidence of anything. "We appear in the majority of answers" means nothing until you know how many answers, generated from how many prompts, run how many times. And because mention rates are proportions estimated from a sample, they come with sampling uncertainty attached. A rate reported as a bare point estimate, no confidence interval in sight, is claiming more certainty than the underlying data can back up.

A zero deserves special scrutiny. A 0% mention rate might mean your brand genuinely has an authority gap in that space. Or it might mean the measurement itself broke: a probe that never reached the model, a prompt set too narrow to catch anything, a single run that happened to come back null. Those two situations look exactly the same on a dashboard. They need to be reported very differently.

Sentiment has to ride along with the rate, or the rate lies by omission. A 20% visibility score paired with predominantly negative sentiment describes a completely different business problem than a 20% score paired with predominantly positive sentiment. The rate alone can't tell you whether the AI mentioning you is helping or actively hurting. AirOps research adds a related wrinkle: brands that earn both citations and mentions are 40% more likely to resurface across repeated runs than brands that only get cited without being named. Track them together and you get a sturdier picture than either one gives you alone.

The practical filter, honestly, is simple. If someone hands you a headline visibility percentage with no denominator, no run count, and no confidence range attached, that's marketing copy. Not measurement.

Where hits-out-of-asks sits in the bigger picture

Mention rate is the foundation. It's where you start, not where you stop. On its own, it tells you whether you're in the conversation at all, and that's useful, but it doesn't tell you much about whether being in the conversation is doing anything for you.

Four companion metrics turn the foundation into something actionable. Share of Voice divides your brand mentions by total brand mentions across every tracked prompt, times 100; if the engines produce 200 total brand mentions and 50 belong to you, your SOV is 25%, which tells you how you stack up against the field rather than just how you look in isolation. Prompt Coverage measures breadth: what share of your prompt library returns your brand at least once, as opposed to mention rate, which measures depth across repeated runs. Recommendation Rate, per Visiblie, divides answers containing an explicit recommendation for you by total prompts tested, and for B2B SaaS this often predicts pipeline better than raw mention counts ever could. Citation Share tracks how often the AI links to your domain versus a competitor's; a lopsided citation count in a competitor's favor is worth paying attention to well before you even calculate a rate off it.

None of this is a niche concern anymore. Gartner projects that 25% of total search volume will shift to AI interfaces by the end of 2026. That's not an early-adopter curiosity; that's infrastructure, and B2B content teams need to treat it that way.

Put together, the pieces tell a coherent story. Mention rate tells you whether you're in the answer. Coverage tells you which questions you're missing entirely. Recommendation rate tells you whether the mention actually carries weight. Sentiment tells you what the AI is actually saying about you when it does mention you. Share of Voice tells you how all of that stacks up against everyone else fighting for the same attention.

Letterbrace treats every one of these as a first-class tracked outcome alongside traditional search signals, not as a separate side project bolted onto the reporting deck. Every figure gets reported as hits-out-of-asks with its sample size and confidence context attached, and a human reviews every strategic move the numbers suggest before anyone acts on it. Measuring this stuff honestly takes more work than pulling a flattering percentage and slapping it on a slide. But a system built to do it honestly hands you a real signal on what to fix. A system built to flatter you just hands you confidence you haven't actually earned yet.

Sources

  1. visiblie.com
  2. pixis.ai

More in Features