The Revenue Signal · Issue 08


In April 2026, the agency DerivateX published the 2026 AI Visibility Benchmark. The Demand Gen Report covered it.

Fifty named B2B SaaS companies. Fourteen hundred buyer-intent prompts. Each prompt runs across multiple LLMs.

Clio scored 89 out of 100. LeadSquared scored 2 out of 100. The category average sat at 56.9. The 87-point spread inside a single benchmark is the actual shape of B2B SaaS citation in April 2026.

What's worth a CEO's attention isn't the spread. It's the measurement gap behind it.

Most B2B teams check their AEO position with a single run of a single prompt on a single platform. That is a snapshot of a stochastic system. It is not a score. And it is why companies look at their AEO dashboard and see numbers that bear no resemblance to what their buyers actually encounter.

This issue covers one thing: why running the same AI prompt once is statistically meaningless, what running it multiple times reveals, and how the top of the citation benchmark got there.

Reading time: ~12 minutes. Exercise time: ~45 minutes.

The Signal

Single-run AEO measurement produces noise, not signal

LLMs are stochastic. The same prompt run twice in ChatGPT, Claude, Perplexity, or Gemini can return different citation sets—and the variance is largest when live search, fan-out queries, or platform-specific retrieval are involved.

The BeVisibleIQ 2026 B2B SaaS AI Citation Study ran 75 buyer-intent prompts across four AI platforms in March 2026 (Perplexity Sonar-Pro, Gemini 2.5-Flash, Claude Haiku-4.5, and ChatGPT GPT-5.4) and classified 2,391 citations collected on a single day. The headline finding split by platform: 79% of citations on Perplexity, Gemini, and Claude point to third-party sites (review pages, comparison content, and listicles), while on ChatGPT, 74.6% of citations point directly to the vendor's own website. Four engines. Four different definitions of "who gets cited."

That platform divergence is the entry point. Most AEO measurement tools still pull from search-style snapshots. They run the prompt once, screenshot, count, and move on. The output looks like a Google rank report, which is what tool buyers expect, but it is not what the underlying system produces.

Two things go wrong with single-run measurements.

First, a citation that appears in run 1 may not appear in run 2. The same BeVisibleIQ study found that Gemini generates 3.7 fan-out sub-queries per prompt on average, rising to 5.0 for comparison queries and as high as 7 for a single "X vs Y" prompt. Every sub-query consults different sources, most of which never appear in the final visible citation list. The retrieval shifts between runs, and so does what gets surfaced.

Second, citation position matters more than presence. DerivateX co-founder Apoorv Sharma told Demand Gen Report that mention rate and position carry 60 of the 100 available points in the benchmark's scoring framework: "That is where the optimization work happens." The implication: a positive mention that appears once a quarter does nothing for the pipeline. Position is what moves demos.

For a CEO, this means the AEO score on your dashboard is probably wrong. Not because the tool is broken. Because one measurement of a stochastic system is noise. You need at least three runs per prompt per platform to separate signal from noise, and many AEO tools still treat prompt results like static rank snapshots—not all disclose whether they use repeated runs per prompt.

That single decision (how many runs per measurement) is the difference between knowing whether your category surface is locked in or guessing.

The Build

How Clio reached the top of the 2026 AI visibility benchmark

Clio, the legal practice management SaaS, was the top scorer in the DerivateX 2026 AI Visibility Benchmark. A composite score of 89 out of 100, against a category average of 56.9.

The visible pattern on Clio's site suggests that its strength comes from category-level publishing depth, not only from technical AEO checks.

Clio's GEO for Law Firms guide breaks down how AI search works for legal buyers, what GEO is, how AI decides which firms to mention, and how to measure visibility. It's written for their buyers (small law firms) but functions simultaneously as Clio's own demonstration of category authority. The piece is structured for chunk-level retrieval: short answer blocks, named entities, specific stats (AI-generated answers now appear in 16.48% of US Google searches, more than double their earlier 2025 level), inline source attribution.

Across their content library, the pattern repeats. The ChatGPT Search for Lawyers post, the Lawyer's Guide to AI Citation, and the Deep Research for Lawyers explainer each target a different buyer-stage question with a structured, citable answer.

The result DerivateX measured: Clio's 89 sits more than 32 points above the 56.9 category average across all 50 companies, and a notable detail from the Demand Gen Report coverage explains why composite scores spread so widely: 44 of the 50 companies scored 19 or 20 out of 20 on sentiment. The visibility gap is not a perception problem. It is a frequency and position problem, and that is what Clio's publishing depth solves for.

What a CEO can take from Clio: AEO position is a function of category-level publishing depth, not technical SEO checkpoints. Clio publishes the actual questions their buyers ask AI. The technical readiness comes second.

The companies sitting near the bottom of the same benchmark appear to have much weaker coverage for AI buyer questions, often relying on older SEO content with FAQ schema layered on top.

The Move

Run a three-run reproducibility check on your top 10 prompts this week

Plan 45 minutes. No new tools. Three steps.

Step 1. List your 10 most commercially valuable buyer prompts. These are the prompts where a prospect researching your category would be likely to enter you into the consideration set. Include 3 research-intent, 4 comparison-intent, and 3 decision-intent. Write them in the exact phrasing a real buyer would use, not your marketing terminology.

Step 2. Run each prompt three times in ChatGPT and three times in Perplexity. Don't refresh, don't reword. Same prompt, three runs. Note for each run: (a) whether you were cited, (b) whether you were cited inline in the body of the answer or only in a "sources" list at the bottom, and (c) which competitors were cited in the same answer. Use a simple spreadsheet. At roughly 30–40 seconds per run if you move quickly, the full 60-run sweep takes 30–45 minutes.

Step 3. Calculate your reproducibility rate. For each prompt, count how many of the six runs (three on each platform) cited you. The above four mean you've locked the slot. "Two or three" means partial position. One or zero means you're not actually winning that prompt. You got lucky if a screenshot exists.

This is a real diagnostic, not a vanity score. The prompts where you scored zero across six runs are the prompts your category-leading competitor is winning by default. Those are your priority targets for the next quarter of content work.

If you want this run systematically (the same sweep, four LLMs, three runs each, reported with per-prompt diagnoses on an ongoing cadence), Revenue Experts AI's AI Experts are built for it. A purpose-built AEO Citation Expert handles the runs, the scoring, and the weekly delta so a revenue team doesn't have to.

Elizabeta Kuzevska Co-Founder, Revenue Experts AI https://revenueexperts.ai

How Revenue Experts AI works on this

Three ways to take action on what's in this issue. They map to three different commitments.

1. Free 60-second AI Visibility Audit. Drop in a URL or up to 10 pages across your domain. The tool runs an 8-strategy crawl that simulates how each AI engine sees the site, scores citation readiness across five categories, and returns a prioritized fix list with industry benchmarks. No call, no signup, no payment. If your readiness score is under 60, fix that before anything else. Run the free audit →

2. $497 AI Visibility Audit (full Citation Audit Method). A custom 50-prompt sweep designed around your actual category, your actual competitors, and your buyer's actual question patterns is run across ChatGPT, Claude, Perplexity, and Gemini, with three runs per prompt per LLM. The output is a research-grade report covering citation rate per prompt, position (inline vs. see-also), competitive map by LLM, and a per-prompt diagnosis explaining where you're losing and why. 5-7 day turnaround. This is the methodology behind the framework in this issue, executed for your business. Book the $497 audit →

3. AI Experts — ongoing citation measurement. A purpose-built AEO Citation Expert running your prompt sweep every two weeks, reporting the delta, flagging where competitors are gaining ground, and recommending the next content move. For teams that want this as a permanent function rather than a one-off audit. See AI Experts →

A practical sequence for most B2B SaaS teams: free audit this week to baseline technical readiness, a $497 audit next month to map actual citation position across the category, and AI experts as the ongoing function once you have something worth tracking weekly.

Elizabeta Kuzevska Co-Founder, Revenue Experts AI https://revenueexperts.ai

Sources

Keep reading