AI Notes · What changed

Why your AI answer changes every time you ask, and what to do about it

Ask ChatGPT the same question twice and you may get two different lists. Owners see that and assume the whole thing is random. It isn’t — but it means a single reading is a reading, not a verdict, and it changes how you should measure.

The first thing a sceptical owner does after reading a Duly Noted report is open ChatGPT and ask the question themselves. Good. Everyone should. About a third of the time the list they get back is a little different from the one in the report — a competitor swapped, an order changed — and they come back asking whether the report is wrong.

It's not wrong. It's a reading of something that moves. Here's why it moves, how much, and what that means for measuring it honestly.

Three sources of variation

The model is probabilistic. Language models generate text by choosing among likely next words, and most assistants run with a little randomness on purpose so answers don't feel canned. Two runs of the same prompt can take slightly different paths and land on slightly different lists, especially at the margin — the third name on a list of three is much less stable than the first.

The retrieval changes. For the answers that name specific local businesses, the assistant is searching the live web first. What it finds depends on the search index that morning, which pages are fresh, what the search engine ranked. A competitor publishes a new page, a directory updates, a review lands — the retrieved set shifts, and so does the summary of it.

The phrasing matters. "Best plumber in Tucson" and "top-rated plumbing companies in Tucson" pull on slightly different sources. This isn't noise; it's real, and it's why the report asks three phrasings rather than one and averages them. A business that appears under all three is on much firmer ground than one that appears under one.

How much it actually moves

Less than the sceptic fears, more than the enthusiast admits. In the reports we run, the first-named business for a local query is usually stable across runs and phrasings. The second is fairly stable. The third and any fourth are where the churn is — a set of four or five businesses trading the last slot. Assistants that cite sources (Perplexity, Gemini with grounding) are more stable than ones that don't, because their answer is anchored to specific pages.

The practical consequence: being reliably first or second in an assistant's list is a position that persists. Being occasionally third is a coin flip that will look different next week.

What it means for measuring

A single report is a reading. It tells you, truthfully, what the assistants said on that day for those questions. It is not a verdict on your business, and treating it as one — in either direction — is the mistake.

Two things make a reading into evidence. Repetition over time: run it monthly, and the trend across six readings is far more meaningful than any one of them. And breadth: twelve cells (four assistants, three questions) smooth out the luck of a single wording or a single model's mood. A business that scores 15 one month and 20 the next hasn't necessarily improved; one that goes 15, 22, 31, 38 over four months has.

This is why the dashboard keeps every run and draws the trend, and why the monthly tier exists. It's also why I'm uneasy about tools that present a single AI visibility number with two decimal places. Precision isn't the same as accuracy, and a number that pretends to be exact about a probabilistic system is hiding the uncertainty rather than showing it.

What to do with variance

Use it. The churn at the bottom of the list is opportunity. If your business is one of the four or five trading the third slot, you are close — the assistants know you exist and consider you a candidate. The difference between you and the business that holds the slot is usually legibility: a clearer title, a schema block, a FAQ page with prices, consistent name-address-phone across the directories. Small, specific work moves a borderline candidate into a stable position, and the trend chart is where you'll see it happen.

If you're absent across all twelve cells and stay absent on a second run, variance isn't your problem. Presence is, and that's a different fix list — the one in the teardown notes.

Re-running before you follow up

A rule I keep for myself and recommend to anyone using these reports in a sales context: re-run before you follow up. Assistants change; a report from three weeks ago may already be stale. Every conversation should carry a fresh reading, and every fresh reading is another point on the trend. The variance that seems like a weakness of the method is, used this way, the thing that makes the method persuasive: you're not showing someone a snapshot, you're showing them a direction.

The honest summary

Answers vary. That's not a flaw in the meter; it's the nature of the thing being measured. Ask three ways, ask four assistants, ask every month, and read the trend rather than the number. That's also, not coincidentally, how anyone who's done search for thirty years learned to read rankings.

Filed under What changed · All AI Notes
The AI Citation Report
Is AI recommending you? Now you know.
41
Overall
15
AI visibility
80
Readiness

What ChatGPT, Claude, Perplexity and Gemini say when a customer asks for a business like yours — who they name, who they link, and the 19 things on your site keeping you out.

Check my business See a sample report
Free readiness scan · Full report $99, once · $39/mo to track
When you want it fixed
Duly Noted is the meter. The SEO Savant is the mechanic.

Thirty years in search, three clients at a time, diagnosis before prescription. Rankings are vanity. Revenue is sanity.

  • Citation Fix Sprint — the report's technical fixes, done, in three weeks$1,950
  • Revenue Audit — two weeks, eight areas, written by hand$3,500
  • The Engagement — full implementation, three seats$6,000/mo
Ask about a Fix Sprint theseosavant.com ↗
More notes
[duly_noted_posts count="4" columns="1" offset="0"]

All AI Notes →

Who writes this

Yishai — in search since 1996, Southern Arizona. Built Duly Noted so owners could see what the assistants say without paying for an agency retainer to find out.

Read enough? Measure yourself.

Free readiness scan. Full report $99. Every claim with the evidence beside it.

Leave a Reply

Your email address will not be published. Required fields are marked *