Live The AI Search Visibility & Ranking Factors reports are now live. Click here to view them →
AI Search Visibility

How One “Meaningless” Word Completely Changes Your AI Recommendation. Eg: The Car You Own Changes Your Hotel Recommendations

Change one word in a prompt, a word that has nothing to do with what you are asking, and the AI recommends a different set of brands.

I ran one prompt 1,374 times across six models. It asks for the top 10 hotel brands in Los Angeles. In passing, it mentions what car the traveler is flying in to pick up. The car has nothing to do with hotels. It is the smallest irrelevant detail I could put in a prompt.

Say the car is “a used Honda Civic” and marriott.com appears in 98.7% of the answers, bestwestern.com in 79.2%, fourseasons.com in 25.2%. Say it is “a brand new Ferrari” and fourseasons.com appears in 99.1%, marriott.com falls to 10.4%, and bestwestern.com never appears at all.

Same question, same models, same settings. The only thing that changed was a detail nobody would think to control for.

1,374answers collected, one prompt template
6models across OpenAI, Anthropic, Google
400distinct hostnames recommended
39different sites held the #1 spot
61.4%chance the same #1 comes back when you ask twice with identical settings
19.3%chance two random runs share the same #1
1,230web searches the models typed while answering
0domains that appear in 90% of answers

Every call used the same template. Only the vehicle text changed:

I am flying into Los Angeles for a night to pick up my new car. It is {vehicle}. Recommend to me, in rank order, the top 10 hotel brands you would suggest I stay at. I’m happy to stay anywhere in the greater LA area. Give me the websites only, no extra text.

Six wordings went into that slot: “a used Honda Civic”, “an off-lease BMW 3 Series”, “a brand new Ferrari”, and the one-word versions “a Civic”, “a BMW”, “a Ferrari”. That makes three pairs, and the two wordings in each pair differ by nothing but a few extra words about the same car.


The raw results, by model, setting, search, and wording

Everything below this section is my read of the data. This section is the data. Pick a model, a setting, whether web search was on, which wordings to compare, and whether to read the answers as hotel names or as domain names. The grid redraws from that exact slice of the 1,363 usable answers.

One row per hotel, in the order the columns introduce them: the first column’s top 10, then whatever the second column adds, then whatever the third adds. Each cell is that hotel’s rank for that column, and a blank cell means it did not make that top 10 at all. Teal is a top-3 finish, blue is 4th to 7th, purple is 8th to 10th, and the shading runs dark at rank 1 to light at rank 10 straight through, so only the hue changes at a band edge. Read a row across to follow one hotel, or step back and look at the shape: a solid corner that empties out in the next column is the group of hotels the car removed.

The default view pools each car across both of its wordings (“a Civic” with “a used Honda Civic”, and so on), so the comparison is car against car with the phrasing averaged out. Switch it to compare the one-word wordings on their own, or the described ones. Hotel names come from a researched lookup covering 99.7% of the 13,646 addresses the models returned. The rest are shown as the raw address, and almost all of those are domains the models invented, like parkerhyatt.com or thepenisula.com. Those are left exactly as the model wrote them: a made-up address is a real result, and repairing it would hide it.

The grid shows the top ten. Underneath each of these columns is a full ranked table, every domain and every hotel the slice ever returned, with how often each one made the top 10 and how often it was the #1 pick, plus every web search the models typed to get there.

Those tables are too wide for this column, so they live in the guide: answers from memory alone and what search changes. Same dropdowns, same data, room to read it.


The models write their own web searches before they answer

With web search turned on, the model does not hand your prompt to a search engine. It writes its own search queries first, reads what comes back, then answers. That step is called fan-out, and it is the first place the car shows up. These are the real strings, pulled from the raw API responses, not a summary of them.

Say “a used Honda Civic” and the model searches for the best hotels in Los Angeles. Say “a brand new Ferrari” and it searches for luxury. The word “luxury” appears in 4 of 332 Civic queries, 60 of 324 BMW queries and 123 of 574 Ferrari queries. That is 1.2%, then 18.5%, then 21.4%.

The plain phrasing moves the opposite way, and it moves further. Queries asking for the best or top hotels without naming a luxury tier go 34, then 11, then 1: 10.2% of the Civic’s searches, 3.4% of the BMW’s, 0.2% of the Ferrari’s. By the time the traveler is picking up a Ferrari, the model has stopped asking which hotels are good and started asking which hotels are expensive.

Nobody asked for a luxury hotel. The traveler mentioned a car.

Every web search the models issued for this selection, counted rather than ranked, most-used first. Search was enabled on every call counted here, so there is no web-search dropdown. A count above 1 means separate calls independently invented the same query, and it is counted every time: a query that ran five times adds five. Above each column: how many queries that slice produced, how many of those contain the word “luxury”, and how many ask for the best or top hotels without naming a luxury tier.

Luxury words climb to 23.4% of searches while plain “best hotels” falls to 0.4%

Out of every web search the models ran, how many contained a luxury word (luxury, 5-star, boutique, upscale), a plain “best hotels” or “top hotel brands” with no luxury tier named, or a bargain word (budget, cheap, affordable, economy). One bar per wording, so these six are the split-out version of the three pooled numbers above. The two leading bars trade places as the car gets more expensive: luxury climbs from nothing to 23.4% while the plain phrasing falls from 11.1% to 0.4%. The described Civic never once produced a luxury word, and it is the only wording that ever produced a bargain word, at 2.1%.

The words “best” and “top” themselves stay flat across the three cars. What falls is “best” with no luxury tier attached: the Ferrari prompt writes “best luxury hotel brands” and the Civic prompt writes “best hotel brands”. Both ask for the best. Only one of them has already narrowed the request to luxury.


Three of the six models rewrite almost the entire top 10 when the car changes

To compare two lists of ten brands I use Rank-Biased Overlap, which scores two ranked lists from 0.0 to 1.0. Identical lists in identical order score 1.0. Lists with nothing in common score 0.0. Agreement near the top of the list counts for much more than agreement near the bottom, which matches how people read a recommendation. Every number below is on that scale, and every one of them is judged against the same reference point: how much two runs agree when nothing at all was changed.

Gemini 3.6 Flash keeps 0.14 of its own top 10 when the car changes

How much of its own top 10 each model keeps when the only thing that changed is the car. The green bar is how much that model already disagrees with itself when nothing changes, so any bar below green is damage the car did. Claude Haiku 4.5 moves the least, 0.55 against its own 0.68. Gemini 3.6 Flash keeps 0.14, ChatGPT keeps 0.16, and Gemini 3.5 Flash-Lite keeps 0.17. Those three rewrite almost the entire list.

“A BMW” and “an off-lease BMW 3 Series” are not the same question

Same car, two ways of saying it, all six models pooled. The dashed line is how much two answers agree when the exact same question is asked twice, so anything under it is caused by the extra words alone. Adding “off-lease” costs the most, dropping the BMW pair to 0.50, because it pulls the answers down-market. “A Ferrari” already tells the model as much as “a brand new Ferrari”, so that pair only falls to 0.64.

Dropping two words from “a used Honda Civic” to “a Civic” takes fourseasons.com from 25.2% of answers to 56.9%. Nothing about the request changed. The traveler still wants a hotel in Los Angeles.


Web search and reasoning effort change the answer more than asking twice does

The prompt is the part you control. There are other inputs you usually do not set, cannot see in the answer, and may not know exist.

Five of the six models change more from the search toggle than from asking the same question twice

The same model asked the same question, once allowed to search the web and once forced to answer from memory. Purple is how much that model’s answers agree when the same question is asked twice. Five of the six models sit clearly below their own purple bar, and Gemini 3.5 Flash-Lite drops to 0.51 against its own 0.63. Search and memory return two different sets of brands.

Claude and Gemini usually answer from memory even with search switched on

How often each model actually ran a web search when it was allowed to. GPT-5.6 Sol searched every time. ChatGPT searched 94.4% of the time. The two Gemini models searched 22.2% and 24.1% of the time, Claude Haiku 4.5 13.6%, and Claude Sonnet 5 10.0%. So most Claude and Gemini answers with web search switched on came from memory anyway, and nothing in the answer tells you which one you got.

The reasoning dial takes GPT-5.6 Sol from 2.67 web searches per answer to 16.67

Reasoning effort is a dial on the API that tells the model how much to think before it answers. Most chat products pick it for you. For every one of the five models that expose it, changing it alone moves the answer further than asking the same question twice does. Gemini 3.5 Flash-Lite falls to 0.51 when asking twice gives it 0.64. GPT-5.6 Sol falls to 0.59 against 0.65.

It also changes how much the model searches. GPT-5.6 Sol averages 2.67 web searches per answer at its lowest reasoning setting and 16.67 at its highest. Your brand has to be findable across all of those searches.

Temperature is the dial everyone knows about, and it is the one that does the least. Across the three models that accept it, changing temperature moves the answer no further than asking the same question twice with identical settings.


Two identical calls name the same #1 pick only 61.4% of the time

An LLM generates an answer one token at a time, sampling from a probability distribution at every step. Two runs of the same prompt on the same model with the same settings are two samples, not two readings of a fact.

Same model, same settings, same question, asked three times. The steadiest result in the whole study is Claude Sonnet 5 at 0.78 with search off. ChatGPT, with nothing touched at all, sits at 0.63.

How often two identical calls named the same brand first. GPT-5.6 Sol keeps its own winner 42.2% of the time. Claude Haiku 4.5 is the steadiest at 79.1%. Across all 1,363 answers, 39 different sites held the #1 spot at some point, and two runs picked at random share a #1 only 19.3% of the time.


Every change in this study, compared with asking the same question twice

Each bar is how much two answers still had in common after changing one thing. Higher is more stable.

All six models pooled. In the top bar nothing was changed at all: the same question was asked twice, and the two answers had 0.67 in common. Every bar below it is a change that pushed the two answers further apart. Temperature and reasoning stay close to that 0.67. Toggling web search and adding a few words about the same car both land at 0.58. Swapping the car outright takes it to 0.28.

Read the top and bottom bars together. Asking the same question twice costs you 0.33 of the list. Changing the car, a detail with no bearing on hotels, costs you 0.72 of it.

No measurement of this prompt can be more reliable than 0.67, because 0.67 is what the measurement returns when nothing is being measured.


How to measure your own AI search visibility

Write the prompts you test the way your customers actually write them, including the details that seem irrelevant. Those details are doing more work than the question is.

Getting those prompts right is most of the job. You are writing from the customer’s side of the conversation: who they are, what they need, the situation they are in, phrased the way they would phrase it. The chat products also carry a profile of the user, built from memory and chat history, and a test prompt has to stand in for that too.

And the prompt cannot lean toward your own business. A persona written so that your product is the natural answer will produce a report where your visibility looks strong, because the prompt made it strong. If the prompts do not describe a real customer, you are paying to monitor metrics that may not matter at all.

Run every prompt many times, not once. One run is one sample. The trend across 5-10 runs of the same prompt is a measurement; a single run is not.

Run each prompt twice over, once with search off and once with search on, and keep the two results apart. Search off tells you what the model has learned about your category. Search on tells you what retrieval is currently adding. Those are separate problems with separate fixes, and averaging them together hides both.

Query through the API rather than a logged-out web session, so the model version, the reasoning effort and the search setting are recorded values rather than unknowns. Then vary the prompt on purpose: several personas, several phrasings, run repeatedly. What you get back is a distribution and a share of answers over time.

The OppAlerts AI Search Visibility report is this same method run across 100 industries instead of one prompt. It is free to read.


Methodology

Six models: gpt-5.6-sol and chat-latest from OpenAI, claude-sonnet-5 and claude-haiku-4-5 from Anthropic, gemini-3.6-flash and gemini-3.5-flash-lite from Google. All through their APIs, in July 2026.

For each model I swept every dial its API exposes. Temperature where it is accepted, which rules out chat-latest, the gpt-5.6 models, and claude-sonnet-5. Reasoning effort or thinking budget where supported. Web search on and off. Six vehicle wordings. Three repeats of every exact configuration. That is 1,374 result slots.

Some responses came back unusable: alphabetized instead of ranked, or a refusal instead of a list. Those slots were deleted and collected again, up to 20 attempts, for 1,694 API calls in total. Eleven slots refused on every attempt, all of them Claude Haiku 4.5 with temperature set and web search on. Those eleven are excluded from every similarity number here. That leaves 1,363 usable answers.

Each response was parsed into an ordered list of hostnames, with the scheme and any leading “www.” stripped. Similarity between two lists is Rank-Biased Overlap, extrapolated, p=0.9, over the top 10 after removing repeats.

Repeats are never averaged into one list before comparing. Every similarity number is the mean over all cross pairs of actual responses. Every dial’s effect is compared against the agreement of repeats in the same slice, where nothing was changed, so a dial only counts as doing something if it moves the answer further than doing nothing does.

Domain-level tables roll subdomains up to the registrable domain. The hotel-name view uses a researched lookup instead, so marriott.com/en-us/brands/westin counts as Westin rather than disappearing into Marriott. Web-search queries were pulled from each provider’s raw response: web_search_call actions from OpenAI, server_tool_use inputs from Anthropic, groundingMetadata from Gemini.

One caveat on the numbers above. This is one prompt, in one category, in one city, on one set of model versions. The effect is large, and it points the same direction across all six models, so I am confident the effect is real. The exact sizes are specific to this test. A different category with a different amount of established consensus would move them, and re-running this against the same models three months from now would move them again.

Written by Ben Wills, 26+ years across marketing and engineering.

Let's discuss how we can work together.

An industry report, a custom dataset, a partnership; a couple of sentences is plenty. Every message comes straight to me, and I read all of them.

If it's a fit, you'll hear back quickly with next steps or a time to talk.

Website Contact Form