Half of today's executives start their research on AI platforms. We wanted to know what they were seeing. What we found instead was that it depends on when they ask.
How often are you running with the first thing AI shares? Is AI omitting the right answer to your question? In our attempts to streamline research and buying journeys, what could we be losing?
See for yourself and select a C-Suite persona below to see the same prompt run twice — and what changes between the two responses.
What metrics should I prioritize to prove the ROI of AI-generated content?
Our research found that 50.4% of identical prompts, from C-Suite personas, produced a substantively different response. This same-session variance measures what happens in real time.
In other words, if you’re a CTO evaluating SaaS platforms and you ask a question, there’s a 50% chance the answer changes if you pose it again, mere seconds later. But you're probably not repeating prompts. So, how much of your online journey is dictated by chance? And are you comfortable with that?
The ways that an AI platform formats its answers can tell you a lot about what it looks for online. Structure is a leading indicator. Of the “core four” AI platforms — ChatGPT, Perplexity, Claude, and Gemini — Gemini and Claude are the most heavily structured.
Most consistently structured LLM. Avg 2,871 chars per response.
By far the longest — avg 7,100 chars, nearly 3× Claude. Most consistent opening structure of all four platforms with the lowest content variance.
The most heavily formatted platform — format is highly stable. But content isn't (68% content variance rate).
Structurally distinct: 100% bold, 98% bullet lists, inline numbered citations. Shortest avg response at 2,362 chars.
Because you can’t control what an AI says, but you can control how easy you make it to quote you. While AI answers are unpredictable, the way platforms organize their answers is rigid. Gemini and Claude default to strict section headers, ChatGPT writes long-form essays, and Perplexity delivers rapid-fire bullet points with footnotes.
AI engines don’t read the web like humans; they break articles into small, bite-sized fragments to build structured summaries. If your key claims and data are buried inside long, winding paragraphs, the engine’s extraction system chops them up and may miss the point. But if you organize your content into clean, self-contained sections with clear subheads and direct facts, you make it effortless for any AI model to lift, synthesize, and cite your brand.
Peer-reviewed research in Generative Engine Optimization (GEO) confirms why structural engineering directly impacts brand presence. In the foundational study published at KDD 2024 by researchers from Princeton University, Georgia Tech, the Allen Institute for AI, and IIT Delhi (Aggarwal et al., GEO: Generative Engine Optimization), researchers demonstrated that deliberate content-optimization strategies can boost visibility inside generative search engines by up to 40%. The highest-performing levers were not traditional keyword density, but architectural extractability: fluency optimization (+28% visibility lift) and the inclusion of standalone statistics and authoritative quotation additions (+30% to +40% visibility lift).
In modern Retrieval-Augmented Generation (RAG) pipelines, automated chunking algorithms fragment narrative documents into isolated semantic windows (typically 300–800 tokens). When a brand’s claims are dispersed across narrative prose, that semantic context is severed during retrieval. Conversely, when content is engineered into modular, answer-first units with descriptive H2/H3 headers, self-contained metrics, and verified attribution, it survives chunking intact — seamlessly slotting into the exact comparative matrices and cited footnotes these platforms construct for executive buyers.
If you're in the C-Suite, there's a 50% chance you're seeing content that won't be there in a second. If you're a CHRO, that drops to 35%. But if you're a Chief Medical Officer, that jumps to 67%.
Your C-Suite audience isn't a monolith; their experiences with AI platforms vary widely.
average same-session variance across all nine personas and four platforms — meaning one in two identical prompts returned a meaningfully different answer.
LinkedIn appeared 3,836 times across 66,030 total citations for C-Suite persona queries — the top source for every persona except our Chief Medical/Clinical officer, where PubMed leads. For PR and content teams, this is immediately actionable: LinkedIn posts are direct inputs to what AI tells B2B buyers when they research a category.
LinkedIn appearances across 66,030 total citations for C‑Suite persona queries
| Persona | Primary | Secondary |
|---|---|---|
| CISOSteve | Microsoft Learn | |
| CTO / CIOIda | CIO.com | |
| CMOMicah | Gartner | |
| CFOPetra | Deloitte | |
| CDODenise | Medium | |
| Chief MedicalHenry | PubMed | |
| CHROCandace | SHRM | |
| Global CommsCaroline | Forbes | |
| CROPeter | Gartner |
Same-session variance tells you what happens in real time. Week-over-week variance tells you something else: whether AI answers to the same questions actually drift across time as models update, new content is indexed, and the information landscape shifts. That’s what we’re currently analyzing. Once we’re done, findings will be shared here (so keep an eye out)!
Subscribe to PAN’s newsletter — AI citation data, findings, and analysis delivered monthly.
SubscribePAN’s Brand to Demand approach connects citation presence to pipeline. See what that looks like in practice.
See the approachPAN’s new brand-to-demand intelligence layer, dropping soon.
Comprehensive findings. Full data visualization. Get it in your inbox the day it publishes.
Join the waitlistBulleted domain lists reflect ChatGPT citation behavior from the original C-Suite Signals research — a ChatGPT‑specific dataset; they do not represent citation behavior on Claude, Gemini, or Perplexity. Top sources are pulled from Perplexity responses. All same-session variance rates shown are confirmed values, calculated from Layer 1 collection across 4,500 unique prompt/platform pairs (9 personas × 4 platforms × 125 prompts × 3 runs). Variance is measured using Dice word-set similarity at a 0.50 threshold. Week-over-week variance data, which may diverge from same-session rates, is currently being compiled. Perplexity source hierarchy derived from 66,030 citations extracted from Perplexity responses to persona-specific queries in Layer 1.
The original C-Suite Signals report asked ChatGPT where executives get their information — across six personas and 10,000+ cited links. The AI Citation Index extends that question to four platforms and adds two new measurements: same-session variance (does the same prompt return a different answer when run again in the same session?) and week-over-week variance (do AI answers drift across time as models update and content indexes change?). The two measurements are complementary but distinct. Same-session variance reveals real-time instability. Week-over-week variance reveals whether AI’s view of your category, competitors, and brand shifts meaningfully from one week to the next.
Same-session variance analysis ran 13,500 total runs: 125 prompt pairs per persona per platform, across nine personas and four platforms (ChatGPT, Claude, Gemini, Perplexity). For each pairing, the identical prompt was submitted three times in the same session and the responses compared — producing the 50.4% same-session variance finding. Citation source data was extracted separately from Perplexity responses and aggregated across 66,030 individual citations, then source-coded by domain and persona to produce the source hierarchy findings. Week-over-week variance analysis will submit the same prompts again across a second wave, separated from the originals by four one-week intervals, to measure how much AI answers drift across time rather than within a session.
Combined, the two analyses will cover 40,000+ discrete AI responses — on top of the original C‑Suite Signals research, which analyzed 15,000+ ChatGPT-cited links from 1,500 searches across nine personas using SparkToro-informed custom GPTs. When week-over-week findings are complete, they will be read directly against the same-session rates: the relationship between the two numbers is the full picture of AI answer stability.
15,000+ ChatGPT-cited links across nine personas — the source hierarchy that the four‑platform Citation Index builds on.