Methodology

Built for a probabilistic search surface

AI Search does not behave like a ranked list of blue links. Answers vary from run to run. Citations shift. Brands appear, disappear, and get replaced by competitors without a clear “#1 position” to manage. A single screenshot of ChatGPT is not a measurement—and a mystery visibility score is not a strategy.

Other sites earn authority by exposing a detailed methodology across their product and documentation rather than asking visitors to trust an unexplained score. They describe criteria extraction, retrieval analysis, fact checking, site structure, and citation diagnostics. Kelsey takes the same posture—and goes further by connecting measurement to the change you should ship next, then proving whether that change moved the answer.

Overview

What Kelsey does

Kelsey is an AI search visibility platform for teams that need buyers to recommend them in ChatGPT, Perplexity, Gemini, and other answer engines—not only rank on Google.

You add your website. Kelsey helps you track the buyer questions that matter commercially: who should I use, what is best for X, how do you compare to Y. For each prompt, we run real searches, record who gets recommended, which sources get cited, and where competitors win. From that evidence we surface the gaps on your site, prepare a concrete content change, and monitor whether the intended AI answers move after you publish.

In short: detect a commercially important recommendation gap, explain the claims and sources causing a competitor to win, prepare the smallest useful change, let your team review and publish, then measure for weeks whether the answer actually changed.

Positioning

What makes us different

Most AI visibility tools stop at tracking. They show mention charts, competitor lists, and a composite score—and leave your team to interpret what to do next. That is useful for reporting. It is incomplete for growth.

Kelsey is built around action and proof, not a fake AI visibility score. Measurement is multi-probe and intent-organized so you are not managing noise. Research questions that uncover “what criteria does the model use?” stay separate from the score you track over time. Recommendations are grounded in answers, citations, competitor pages, and your own site—not invented from a model’s memory. Content drafts go through claim verification. Experiments start with a hypothesis tied to specific prompts. Results compare repeated before-and-after probes and separate target prompts from unrelated ones so normal variation is not mistaken for a win.

That is where Kelsey can appear more scientifically responsible than competitors without pretending AI Search is perfectly deterministic.

01

One answer is not a measurement

Large language models are stochastic. Ask the same buying question twice and you may get a different shortlist, different wording, or different citations. Betting a narrative—or a board update—on one reply is like judging an ad campaign from a single impression.

Kelsey runs multiple independent probes for the same buyer intent and reports how often a brand appears. Mention rate, spread across probes, and movement over time are the unit of analysis—not whether you got lucky once in a cherry-picked screenshot. Inside the product, the methodology for a scan is visible: same core question, probes spaced in time, and which model answered.

We treat visibility the way serious teams treat media and SEO: frequency and pattern, not a one-off sample. When probes disagree more than usual, we surface that volatility so you treat movement as signal, not noise from a single run.

02

Prompts are organized by buyer decision

Buyers do not ask one perfect keyword. They ask discovery questions, use-case questions, comparison questions, decision questions, and FAQs—often with small wording changes that should not rewrite your strategy each week.

Related prompts are grouped into topics so one wording variation does not distort the result. Hierarchy matters: topic → buyer stage → prompt family → wording variation → scan. You can see whether you win the decision, not only whether a particular sentence happened to include your brand.

Where it matters, Kelsey also supports buyer segments and personas. The same prompt can produce different recommendations for a solo founder versus an enterprise IT buyer. Measuring without that lens conflates audiences and hides where you actually lose.

03

Research is separate from tracking

Understanding why a model recommends someone is different from measuring how often you appear for the prompts you care about. Mixing those jobs pollutes the metric.

Kelsey uses additional research questions to uncover recommendation criteria—what attributes, proof points, and comparisons show up when assistants shortlist vendors. Those research answers inform content strategy and opportunity prioritization. They do not get mixed into your visibility score.

Tracking stays clean: commercially important prompts, repeated probes, mention patterns over time. Research stays exploratory: criteria extraction, competitive patterns, and gaps to close. Teams can trust the score as a tracking series and still dig into qualitative “why” without the two contaminating each other.

04

Every recommendation requires evidence

Pure LLM advice about your site hallucinates. Kelsey connects findings to the original answers, cited sources, competitor pages, and your own website so a recommendation is a trail you can audit—not a black-box tip.

When a competitor wins a prompt, we surface the claims and sources that show up alongside them: comparison pages, review sites, Reddit threads, pricing transparency, entity consistency, and other citation signals. When we suggest refreshing or creating a page, the reason types are explicit—patterns from winning answers, patterns from competitor pages, or a gap on your site—tied back to URLs and prompts.

Site audits follow the same rule. We crawl real pages, check technical and AI-indexing signals, match coverage to the questions buyers ask, and ground synthesis in evidence URLs. The goal is not a prettier dashboard. It is a next action you can defend in a content review.

05

Claims are verified before publication

AI Search amplifies whatever is easy to retrieve and easy to quote. Publishing unsupported claims—pricing, integrations, guarantees, competitive comparisons—can get repeated by assistants even when it is wrong. That hurts trust faster than silence.

If Kelsey cannot support an important claim from your site or from evidence we can show, it asks your team to confirm it or leaves it out. Drafts are built to be reviewable: what changed, why it matters for the prompts you care about, and which statements still need human confirmation.

We would rather ship a smaller, accurate page than invent facts that look good in a demo and fail in production.

06

Experiments begin with a hypothesis

Publishing more content is not a methodology. Every content change in Kelsey is tied to the criteria it addresses and the prompts it is expected to influence. Before you publish, you should be able to say: “We are missing proof point X that shows up when assistants recommend Y; this change is meant to move prompts A, B, and C.”

That framing keeps experiments honest. It stops vanity refreshes disconnected from buyer intent. It also makes failure useful: if the target prompts do not move, you learned something about the criteria or the competitive evidence—not that “content didn’t work” in the abstract.

The product loop is intentionally small: the smallest useful change that closes a commercially important gap, then proof—not a sprawling content calendar justified by a single score.

07

Results account for normal variation

Because AI answers vary, a one-scan “before” and one-scan “after” can lie in either direction. Kelsey compares repeated pre- and post-publication probes and separates target prompts from unrelated prompts so you can see whether the change you shipped moved the intents you meant to move.

We do not treat a single probe as the result. When there are not enough probes yet, the product is explicit that certainty is still building. Stronger reads come from more independent runs, not from dressing a thin sample as a definitive win.

Unrelated prompts act as a control lens. If everything moves together, you may be seeing platform-wide noise. If target prompts move while controls stay flat, you have a stronger story that the change mattered. That is how we evaluate experiments without claiming AI Search is a deterministic A/B test.

How it fits together

From measurement to a change you can prove

Put together, the methodology is a closed loop:

  1. Measure how often you appear across independent probes for buyer-decision topics—not one answer.
  2. Research recommendation criteria separately so the score you track stays uncontaminated.
  3. Explain who wins with evidence from answers, citations, competitor pages, and your site.
  4. Prepare a hypothesis-linked change with claims your team can verify.
  5. Publish, then compare repeated pre/post probes on target prompts versus unrelated prompts.

This is how Kelsey stays more scientifically responsible than tools that hide randomness behind one score—and more useful than tools that only report. We expose the method. We still respect that AI Search is probabilistic.

Ready to see your AI visibility?

Run real customer-style searches across ChatGPT, Perplexity, and Gemini.

Run a free scan