LLM Ranking Factors Clarity & Direction for AI Search Visibility & SEO Campaigns
A new, large-scale correlation study of what actually moves the needle in ChatGPT, Claude, and Gemini answers, rebuilt on a dataset roughly 10× larger than our May 2026 release, across 100 industries and 1,100 buyer personas.
Get Expanded Industry Research
This post covers the all-industry numbers and one industry example. The live platform has this same breakdown, plus over/under-performer tables, fanout queries, and persona pages, for all 100 industries, refreshed roughly every one to three weeks as new link, crawl, and Reddit data comes in.
| Plan | Launch price | Regular price | What you get |
|---|---|---|---|
| 3 industries | $374.25 / qtr | $499 / qtr | Any 3 of the 100 industry reports, swap your picks each renewal |
| 10 industries Most popular | $749.25 / qtr | $999 / qtr | Any 10 of the 100, best value per report |
| All 100 industries | $2,624.25 / qtr | $3,499 / qtr | Everything unlocked, nothing to pick |
Coupon FOUNDERS25 at checkout.
Need this built for your industry specifically?
I’m taking on custom work directly: custom personas, competitor sets, or an industry that isn’t in the 100 covered here. I’m limiting this to 5 clients at a time. Expected turnaround is 4 to 8 weeks per project, the goal is one month, but give it room to run longer.
The All-Industry Report
How this research was built, what changed since May, and every signal measured, before we zoom into one industry example.
This Is Correlation, Not Causation
Every number in this report describes what moves together. None of them prove what causes an LLM to recommend one domain over another.
Every ρ in this report is a Spearman correlation. It measures how strongly a signal and LLM visibility move together across thousands of domains, not whether one causes the other. A domain that scores well on a signal tends to also get recommended more, for reasons this report cannot fully isolate from the outside.
That distinction has a real consequence: improving a signal that correlates strongly with visibility is not a guarantee that your own visibility improves. It means you’re moving in the direction that’s associated with getting recommended more often, not pulling a lever with a known, guaranteed effect.
The single strongest signal we measure, Search Engine Outbound Links, explains about 11% of the variance in recommendation behavior on its own (R² = 11%). Even accounting for all 18 signals together, most of what determines an LLM’s recommendation happens inside the model, in places external data can’t see: training data composition, fine-tuning, RLHF preferences, and brand familiarity built up over years. We can measure what correlates with getting recommended. We cannot see the causal machinery that actually produces the recommendation.
Treat this report as a map of what to test, not a guarantee of what to fix. If you improve a signal that correlates strongly with visibility and nothing changes, that’s not a contradiction. That’s correlation and causation working exactly as different things.
How To Use This Report
Turn the data into sales conversations, client priorities, and the next three to six months of work.
| Use case | How to use it |
|---|---|
| Sales | Show prospects exactly where their AI visibility is weak, which competitors get recommended instead, and what gaps your team can close. |
| Agency / service | Use the industry and persona pages to scope the next three to six months of client work: search, content, entity cleanup, Reddit and community presence, backlinks, and reputation monitoring. |
| In-house | Use the persona breakdowns to decide which buyer segments matter most, and where your current AI visibility is missing. |
The benefit is focus. Instead of generic AI visibility tactics, you get the exact industry, the exact persona, and the signals that actually correlate with getting recommended in that market.
The Largest LLM Ranking Factors Report, To Date
I believe this is the largest LLM ranking factors analysis published anywhere. If something bigger or more comprehensive exists, let me know.
- 100 industries, each scored against its own vocabulary and its own buyer language.
- 1,100 industry/persona combinations: 10 targeted personas plus one neutral persona per industry.
- 403K+ prompts run across ten models: Claude, GPT, Gemini, DeepSeek, and GLM.
- 145,289 distinct domains recommended by an LLM at least once, scored against 18 external signals.
- 150K+ organic searches, across Google and Bing.
Compared to our own May 2026 release
| May 2026 | July 2026 | |
|---|---|---|
| Industries | 145 | 100 see below |
| Buyer personas | 1,595 | 1,100 |
| LLM prompts | 105K+ | 403K+ |
| Domains recommended and scored | 29,562 | 145,289 |
| External signals measured | 13 | 18 |
| Models sampled | 1 (ChatGPT 5.4) | 10, across 5 providers |
| Total data analyzed | 500TB | 3PB |
The industry count went down. Everything else went up, by a lot: this edition samples 10 models instead of one, analyzes 6x the raw data, and the underlying dataset is roughly 10x larger overall.
The Scale
The research required fast code, not a giant cloud budget: 3 petabytes of data analyzed, up from 500TB in May.
| Input | Scale |
|---|---|
| Total data analyzed | 3PB |
| Web pages crawled | 15B+ |
| Links analyzed (web graph) | 282B+ |
| Reddit submissions | 3B+ |
| Reddit comments | 26B+ |
| Wikipedia articles | 25M+ |
| Wikidata entities | 120M+ |
| Organic search results collected | 150K+ queries |
| Tracked hostnames | 352,000 |
| Weighted vocabulary phrases | 110,221 |
Despite the size, this pipeline runs without a large cloud spend. That’s a result of years spent on fast parsing and a hand-written, high-throughput HTML/text extraction pipeline. Changing how a signal is scored, as this edition did for two of them, is a recompute measured in minutes once the raw scans exist.
Research Process
How the recommendation dataset was built.
- 100 industries, each broken into 10 targeted buyer personas plus one neutral, industry-wide persona.
- Every persona gets its own weighted vocabulary: 110,221 distinct phrases across 1,200 industry/persona segments, averaging 2.7 segments per phrase.
- 352,000 tracked hostnames, mapped to the industries they compete in.
- LLM recommendation runs across 10 models from 5 providers, plus the web searches those models issue on their own while answering (fanout).
- Every recommended domain is cross-referenced against search results, homepages, web crawl data, Reddit, Wikipedia, Wikidata, and backlink graph data.
| Data source | What it measures |
|---|---|
| LLM prompts | Which domains get recommended, by industry and persona |
| LLM search fanout | The literal web searches models run while answering |
| Google SERPs | Which domains appear for the same industry and persona language |
| Downloaded search-result pages | Outbound links and phrase usage from the pages that win search rankings |
| Common Crawl | Open-web co-occurrence of a domain and the industry’s language |
| Community mentions and phrase co-occurrence | |
| Wikipedia & Wikidata | Encyclopedia citations and structured entity data |
| Backlink graph | PageRank, harmonic centrality, and backlink counts, at domain and host level |
| Downloaded homepages | How closely a domain’s own homepage speaks the industry’s language |
How Signals Are Scored
Every correlation in this report is a Spearman ρ between one signal and a domain’s LLM recommendation score.
- Rank-based: only the ordering of domains matters, not the scale of the raw numbers.
- R² (ρ squared) is the share of rank variation that signal explains on its own.
- A domain missing a signal is excluded from that signal’s correlation. It is never scored as zero.
- Every signal is scored globally, per industry, and per persona, using that scope’s own vocabulary. The same domain can have a different score for the same signal in three different tables.
Improved Keyword Relevance Scoring
Search Engine Outbound Links and Homepage Keyword Relevance use an improved scoring approach this edition: instead of a flat keyword match, a phrase now counts 3x if it appears in a page’s title tag, 2x in its meta description, and 1x in the visible body. A title tag holds a handful of words, so what a page spends those words on is a stronger signal of what the page is actually about than a body-text mention is. This more accurate scoring is a large part of why both signals moved so much since May, see Why Search Signals Jumped.
How The Rankings Moved
All 14 signals we can directly compare across both editions. ▲ teal = strengthened since May. ▼ red = weakened.
| Signal | May ’26 ρ | Jul ’26 ρ | Δ | Trend |
|---|---|---|---|---|
| Search Engine Outbound Links | 0.230 | 0.331 | +0.101 | ▲ |
| Homepage Keyword Relevance | 0.072 | 0.204 | +0.132 | ▲ |
| Search Engine Appearances | 0.241 | 0.165 | -0.076 | ▼ |
| Best Search Engine Rank | 0.238 | 0.148 | -0.090 | ▼ |
| Backlink Count (Domain)split host/domain in Jul | 0.204 | 0.160 | -0.044 | ▼ |
| Backlink PageRank (Domain)split host/domain in Jul | 0.194 | 0.192 | -0.002 | ● |
| Backlink PageRank History (Domain)split host/domain in Jul | 0.200 | 0.183 | -0.017 | ▼ |
| Backlink Harmonic Centrality (Domain)split host/domain in Jul | 0.169 | 0.151 | -0.018 | ▼ |
| Common Crawl Presence | 0.123 | 0.165 | +0.041 | ▲ |
| Wikidata Entities | 0.120 | 0.151 | +0.031 | ▲ |
| Reddit Comment Mentions | 0.111 | 0.148 | +0.037 | ▲ |
| Reddit Submission Mentions | 0.096 | 0.128 | +0.033 | ▲ |
| Average Search Engine Rank | 0.096 | 0.077 | -0.019 | ▼ |
| Wikipedia Presence | 0.077 | 0.055 | -0.023 | ▼ |
Rows marked with a methodology note compare May’s single blended backlink metric against July’s domain-level equivalent; July also reports a separate host-level variant of each (see the full table below). Two of these rows also changed because of the improved scoring described above, not just dataset size.
Why The Rankings Moved
Two structural changes explain most of it, visualized below: every signal’s change in ρ since May, sorted worst to best.
The dataset grew roughly 10x since May, which tightens every correlation by reducing noise, and two signals were rescored with the title/meta/body weighting described above. Both are real, and together they explain the movement.
All 18 Signals, Ranked
All 18 signals, ranked by ρ against LLM visibility. Color marks the tier.
These are the all-industry pooled tier boundaries. Individual industry pages (see the Airlines example below) use the same signals but their own tier cutoffs relative to that industry’s own distribution.
Full July 2026 Data Table
The complete correlation table, all-industry pooled, for readers who want every number.
| # | Signal | ρ | R² | Coverage | Tier |
|---|---|---|---|---|---|
| 1 | Search Engine Outbound Links | 0.331 | 11% | 28% | Dominant |
| 2 | Homepage Keyword Relevance | 0.204 | 4.2% | 28.6% | Strong |
| 3 | Backlink PageRank (Domain) | 0.192 | 3.7% | 67.2% | Strong |
| 4 | Backlink PageRank History (Domain) | 0.183 | 3.3% | 67.8% | Strong |
| 5 | Search Engine Appearances | 0.165 | 2.7% | 6.4% | Strong |
| 6 | Common Crawl Presence | 0.165 | 2.7% | 42.8% | Strong |
| 7 | Backlink Harmonic Centrality (Host) | 0.164 | 2.7% | 67% | Strong |
| 8 | Backlink Count (Domain) | 0.160 | 2.6% | 51.5% | Strong |
| 9 | Backlink PageRank (Host) | 0.156 | 2.4% | 67% | Strong |
| 10 | Backlink Harmonic Centrality History (Domain) | 0.153 | 2.4% | 67.8% | Strong |
| 11 | Backlink Harmonic Centrality (Domain) | 0.151 | 2.3% | 67.2% | Strong |
| 12 | Wikidata Entities | 0.151 | 2.3% | 8.5% | Strong |
| 13 | Backlink Count (Host) | 0.150 | 2.2% | 45.2% | Confirmed |
| 14 | Best Search Engine Rank | 0.148 | 2.2% | 6.4% | Confirmed |
| 15 | Reddit Comment Mentions | 0.148 | 2.2% | 33.2% | Confirmed |
| 16 | Reddit Submission Mentions | 0.128 | 1.6% | 28.9% | Confirmed |
| 17 | Average Search Engine Rank | 0.077 | 0.6% | 6.4% | Emerging |
| 18 | Wikipedia Presence | 0.055 | 0.3% | 16.9% | Emerging |
ρ = Spearman correlation vs. LLM visibility score. R² = share of variance explained by this signal alone. Coverage = % of sampled domains with a non-null value for this signal.
What’s Actually Driving This
The 18 signals fall into four natural families. Here’s how each one is trending.
Search Engine Signals
Search Engine Outbound Links, Search Engine Appearances, Best Search Engine Rank, Average Search Engine Rank. Being cited (outbound links) now outperforms merely appearing (appearances) by a wide margin, the reverse of May. Part of that is the title/meta/body scoring change described above, not just market movement. Rank position within results (average rank, best rank) matters less than either.
Backlink & Authority Signals
PageRank, Harmonic Centrality, and Backlink Count, each split into domain-level and host-level variants. Nine of the eighteen signals in this table are backlink-graph metrics. Domain-level cuts beat host-level cuts on every single pair, the split itself is the finding.
Content & Relevance Signals
Homepage Keyword Relevance. The one on-page signal we measure, and the biggest riser in the whole study. A domain’s homepage now needs to actually be topically relevant, not just technically indexed.
Community & Knowledge Graph Signals
Reddit Comment Mentions, Reddit Submission Mentions, Wikidata Entities, Common Crawl Presence, Wikipedia Presence. Reddit and Wikidata both strengthened. Wikipedia article presence, on its own, weakened: structured data and live discussion appear to be more useful to LLMs right now than encyclopedic prose alone.
Why 100 Industries, Not 145
This edition covers 100 industries, down from 145 in May. 45 were cut, 0 were added, the new list is a strict subset of the old one.
The 45 industries removed since May 2026
- Agricultural equipment
- Apartment rentals & multifamily leasing
- Auto repair & maintenance
- Beer brands
- Beer, wine & liquor stores
- Bottled water & functional beverage brands
- Budget hotel chains
- CRO / clinical services
- Car-wash chains
- Collision-repair centers
- Colocation interconnection services
- Commercial HVAC equipment
- Commercial mortgage lending
- Commercial real-estate listing marketplaces
- Commercial solar EPC
- Coworking / flex office
- Cruise booking sites
- Dental services
- Garage-door services
- Gas stations & fuel retail
- General contractors (commercial)
- HVAC service contractors
- Hair-salon & barber chains
- Home services (plumbing, HVAC, remodeling)
- Hospital staffing agencies
- Hospitals
- IP / patent law firms
- Luxury hotels
- M&A advisory boutiques
- Managed network services
- Office furniture / workplace equipment
- Office supplies wholesalers
- Oil-change chains
- Pest-control services
- RV dealerships
- Roofing services
- Sales-outsourcing / SDR services
- Senior home-care services
- Server hardware for enterprises
- Soft drink brands
- Spa & massage chains
- Theme parks & amusement parks
- Trade media / B2B publishers
- Veterinary services
- Workers’-comp insurance
Every Industry In This Report
Click any industry to jump straight to its live report.
- Accounting software
- Affiliate-marketing networks
- Airlines
- Athletic apparel brands
- Auto OEM brands
- Auto insurance
- Auto-glass repair
- B2B ad agencies
- B2B marketing data providers
- Baby care & diaper brands
- Beauty & cosmetics retail
- Behavioral-health / therapy platforms
- Brokerage & wealth-management apps
- CRM software
- Car rental brands
- Clothing & apparel retail
- Cloud infrastructure services
- Colleges & universities
- Commercial banking
- Commercial solar / energy services
- Consumer banking
- Consumer legal services
- Consumer wealth advisors / RIAs
- Corporate tax advisory
- Cosmetic dentistry
- Cosmetic-surgery clinics
- Cosmetics & makeup brands
- Credit cards
- Credit monitoring services
- Cruises
- Customer support / contact center software
- Data analytics / BI software
- Data center colocation
- Data-warehouse / lakehouse platforms
- Debt settlement & credit repair services
- Debt-consolidation lenders
- Dermatology groups
- E-signature & document workflow
- ERP software
- Enterprise AI platforms
- Enterprise search & knowledge copilots
- Executive search
- Fertility clinics
- Fitness clubs
- Food delivery platforms
- Furniture stores
- GPU / AI infrastructure vendors
- Golf equipment brands
- HR consulting
- HR/payroll software
- Haircare brands
- Health insurance
- Healthcare IT for providers
- Healthcare practice-management software
- Home centers
- Home exercise equipment
- Homebuilders / new homes
- Homeowners insurance
- Hotels & resorts
- Household cleaning brands
- IVF & reproductive medicine
- Influencer-marketing platforms for brands
- Insurance brokerages / benefits brokers
- Jewelry, luggage & leather goods retail
- Legal services for businesses
- Life insurance
- Live entertainment & ticketing
- Luxury fashion & accessories
- Managed legal services
- Market research & insights firms
- Marketing automation software
- Martech / CDP / attribution software
- Mattress stores
- Med spas & aesthetic clinics
- Media-buying agencies
- Mortgage lenders
- Nutrition & supplement retailers
- Online travel agencies
- PR & communications agencies
- Payroll / PEO services
- Personal checking accounts & neobanks
- Personal loan & fintech lenders
- Procurement software
- Renters insurance
- Residential real estate brokerages
- Residential solar installers
- SEO / content-marketing agencies
- Scientific instruments
- Self-storage brands
- Ski resorts
- Skincare brands
- Sneaker brands / athletic footwear
- Staffing agencies
- Streaming video services
- Tax-prep services for consumers
- Used car dealers
- Vacation rental platforms
- Weight-loss clinics & GLP-1 telehealth
- Wine & spirits brands
- Wireless carriers
Example Industry: Airlines
Everything so far has been the all-industry view. What follows zooms into one industry, end to end, so you can see exactly what the full paid report looks like.
Airlines
A concrete walkthrough of what a single industry report looks like, end to end.
Airlines is a useful example because the brands are ones almost everyone recognizes, and because the industry shows nearly every finding in this report at once: a search-visibility mismatch, a persona-driven recommendation set, and a wide gap between LLM recommendations and Google’s AI Overviews.
| Domains tracked | 507 |
| Buyer personas | 11 (10 targeted + neutral) |
| Top signal | SE Outbound Links, ρ = +0.552 |
| Top LLM-recommended domain | delta.com |
The following sections use this one industry to show what the full paid report looks like: the signal correlations, the cross-channel comparison, the over/under-performer analysis, the literal search queries models run, and how much the answer changes by persona. Every number below is pulled from the live report site, the same thing a paying visitor sees.
What Predicts Airline Recommendations?
Every top-15 signal reads “Dominant” on this industry’s own scale, more than usual is explained by what we measure.
| Signal | ρ | R² | n | Rank Influence |
|---|---|---|---|---|
| SE Outbound Links | +0.552 | 30.5% | 296 | Dominant |
| Reddit Comments | +0.551 | 30.4% | 358 | Dominant |
| Common Crawl | +0.517 | 26.7% | 375 | Dominant |
| Reddit Posts | +0.508 | 25.8% | 327 | Dominant |
| Wikipedia Citations | +0.452 | 20.4% | 280 | Dominant |
| Domain PageRank | +0.396 | 15.6% | 448 | Dominant |
| PageRank History | +0.385 | 14.8% | 453 | Dominant |
| Domain Backlinks | +0.379 | 14.4% | 407 | Dominant |
| Host Backlinks | +0.369 | 13.6% | 375 | Dominant |
| Harmonic Centrality History | +0.354 | 12.6% | 453 | Dominant |
| Host PageRank | +0.341 | 11.6% | 451 | Dominant |
| Search Engine Appearances | +0.338 | 11.4% | 80 | Dominant |
| Host Harmonic Centrality | +0.338 | 11.4% | 451 | Dominant |
| Domain Harmonic Centrality | +0.315 | 9.9% | 448 | Dominant |
| Homepage Keywords | +0.301 | 9.1% | 203 | Dominant |
| Wikidata | +0.275 | 7.6% | 194 | Strong |
| Best Search Engine Rank | +0.255 | 6.5% | 80 | Strong |
| Avg Search Engine Rank | +0.061 | 0.4% | 80 | Emerging |
Reddit is essentially tied with search citation as the top predictor, both Reddit signals outrank every backlink metric. Source: oppalerts.com/AI-Search-Visibility/airlines/.
Four Channels, Four Different Winners
The same airline domains, ranked four different ways: what LLMs recommend directly, what LLMs search for while answering (fanout), what shows up in Google’s AI Overviews, and plain organic search rank. Each column is a 0–100 percentile within that channel, a dash means the domain never appeared there at all.
| Domain | LLM | Fanout | AI Overview | Organic |
|---|---|---|---|---|
| alaskaair.com | 96 | 92 | 79 | 26 |
| emirates.com | 91 | 74 | 87 | 15 |
| qatarairways.com | 90 | 97 | 66 | – |
| aircanada.com | 87 | 63 | – | 90 |
| britishairways.com | 80 | 45 | – | 45 |
| ana.co.jp | 79 | 61 | – | 60 |
By average organic rank, YouTube and Travelocity lead the airlines category, not an airline. By LLM recommendation, Delta leads outright. Google’s AI Overviews cite a mix of airline sites, aggregators, and ranking authorities that barely overlaps with either list. Each channel measures something different, none of them is a reliable proxy for the other three.
Who AI Actually Recommends
Recommendation score is a weighted reciprocal-rank sum across every model and run, scaled by each model’s influence weight.
| Domain | Score | Appeared in |
|---|---|---|
| delta.com | 768.8 | 6.4% of runs |
| southwest.com | 555.9 | 5.9% |
| united.com | 408.9 | 6.5% |
| google.com | 362.2 | 3.3% |
| alaskaair.com | 342.7 | 5.2% |
| aa.com | 219.8 | 4.2% |
| singaporeair.com | 216.4 | 3.3% |
| studentuniverse.com | 176.5 | 0.9% |
| qatarairways.com | 165.4 | 3.0% |
| skyscanner.com | 144.7 | 2.5% |
Delta, Southwest, United, and Alaska lead, in line with conventional US market share. The outlier: google.com ranks #4, ahead of American Airlines, models frequently recommend “search Google Flights” as the answer itself, not just an airline brand. Full ranked list (25 domains) available in the paid report.
Where LLM Recommendations Diverge
Domains ranked by percentile within each channel. Over-performers rank far better in pure LLM recommendations than across fanout, AI Overviews, and organic combined. Under-performers are the mirror case.
| Over-performers | LLM | Fanout | AI Overview | Organic | Gap |
|---|---|---|---|---|---|
| studentuniverse.com | 92 | 24 | – | – | +84 |
| sta-travel.com | 71 | – | – | – | +71 |
| spirit.com | 68 | – | – | – | +68 |
| lufthansa.com | 84 | 50 | – | – | +68 |
| Under-performers | LLM | Fanout | AI Overview | Organic | Gap |
|---|---|---|---|---|---|
| reddit.com | – | – | 100 | 100 | −67 |
| youtube.com | – | – | 99 | 93 | −64 |
| worldairlineawards.com | – | – | 98 | 94 | −64 |
| nerdwallet.com | – | – | 91 | 98 | −63 |
Generic platforms and aggregator/content sites dominate organic results and AI Overview citations but barely register as direct LLM recommendations. Niche, persona-specific travel brands overperform in pure LLM recommendations relative to their weaker footprint everywhere else.
What The Models Search For
When a model needs the web to answer, it writes and runs its own search queries. These are captured verbatim.
| Buyer question (persona) | Query the model actually ran |
|---|---|
| Business Road Warrior, best on-time airline | OAG Punctuality League 2025 airlines on-time performance Delta United Alaska Qatar Emirates ANA Singapore |
| Premium Leisure Flyer, best business class | 2025 Skytrax World Airline Awards best business class first class premium economy airlines |
| Student Abroad Flyer, cheap student fares | site:studentuniverse.com student flights official flexible ticket terms baggage airlines |
| Senior Comfort Traveler, accessibility | AARP senior travelers best airlines customer service accessibility |
| Family Vacation Planner, seating policy | US airline family seating dashboard DOT children sit next to parents |
| Miles Maximizing Loyalist, best rewards program | The Points Guy best airline loyalty programs 2026 |
Notice how many of these queries target a specific third-party authority by name, Skytrax, OAG, The Points Guy, NerdWallet, J.D. Power, AARP, not the airline’s own site. Ranking well with the review sites and industry-authority lists your buyers’ models actually search for matters as much as your own homepage.
Same Industry. Different Buyer, Different Answer.
The industry-wide #1 recommended airline is Delta. No single airline is #1 for every persona.
| Persona | #1 recommended domain |
|---|---|
| Business Road Warrior | delta.com |
| Senior Comfort Traveler | delta.com |
| Family Vacation Planner | southwest.com |
| Miles Maximizing Loyalist | united.com |
| Student Abroad Flyer | studentuniverse.com |
The Student Abroad Flyer result is worth calling out: it’s a student-fare booking site, not an airline, and it was the #1 recommended domain for this persona in our May 2026 report too. That’s a real, repeated finding, not noise: a persona built around a specific need pulls in a specialist brand that never shows up as a top recommendation for anyone else asking about airlines.
This is the central measurement problem this whole report keeps coming back to: a single aggregate “who’s #1 in airlines” number hides five different correct answers, depending on who’s asking.
Back To The Full Report
That’s Airlines end to end. The rest of this report returns to the all-industry findings and what they mean for you.
Key Findings From The Report
The data doesn’t point to one universal ranking factor. Industries lean on different evidence layers, and persona context changes the recommendation set entirely.
- Search engine outbound links are now the single strongest predictor across all industries, ρ = 0.331, up from rank #3 in May, partly a real shift and partly a scoring change.
- Homepage relevance went from the weakest signal we measured to a top-2 signal, a 3× increase, for the same reason.
- Backlink authority got more granular this edition (domain-level vs. host-level) and mostly held its ground, the extra detail is additive, not a correction.
- In Airlines specifically, Reddit discussion is essentially tied with search citation as the top predictor, ahead of every backlink metric.
- A brand’s LLM recommendation score can diverge sharply from its organic rank and its AI Overview citations, in either direction.
- The literal searches models run while answering name specific third-party authorities far more often than they name the brand’s own site.
- No single brand is the top recommendation across every buyer persona in an industry, even a category as consolidated as US airlines splits five different ways by persona.
- Persona coverage, how many distinct buyer segments recommend you at all, is a different and arguably more useful number than one aggregate visibility score.
Why Search Signals Jumped
Two of the biggest movers since May jumped because we improved how we score them.
In May, a keyword match anywhere on a page counted the same whether it was in the title tag or buried in the footer. In July, a phrase match counts 3x if it’s in the page’s title, 2x if it’s in the meta description, and 1x if it’s in the visible body. The reasoning: a title tag holds a handful of words, so what a page spends those words on is a much stronger signal of what the page is actually about than a body-text mention is.
That improved scoring explains a large part of why SE Outbound Links jumped from rank #3 (ρ = 0.230) to rank #1 (ρ = 0.331), and why Homepage Relevance jumped from the weakest signal in the whole May table (ρ = 0.072) to a top-2 signal (ρ = 0.204), a 3× increase.
The dataset also grew roughly 10x since May, which independently tightens every correlation by reducing noise. Both the better scoring and the bigger dataset are real, measurable improvements, and together they explain the jump.
Persona Coverage Is A New Kind Of Market Share
A brand’s AI visibility is a distribution across personas, not one score.
| Airline | Personas where it’s the #1 recommendation |
|---|---|
| delta.com | 2 of 11 shown (Business Road Warrior, Senior Comfort Traveler) |
| southwest.com | 1 of 11 shown (Family Vacation Planner) |
| united.com | 1 of 11 shown (Miles Maximizing Loyalist) |
| studentuniverse.com | 1 of 11 shown (Student Abroad Flyer) |
Persona coverage measures how many distinct buyer contexts can see you at all, not just your single best-case score. For a brand that competes across multiple segments, broad persona coverage may matter more than one high aggregate number that’s actually being carried by a single strong segment.
What Conclusions Can We Draw?
The research doesn’t produce one magic ranking factor. It produces a measurement framework.
- Search visibility matters, but it explains part of the picture, not the whole thing.
- Backlink authority, Reddit discussion, Wikipedia and Wikidata presence, and homepage relevance vary dramatically by industry, Airlines leans on Reddit and search citation, other industries lean elsewhere.
- The best signal for one industry can be minor in another.
- The most useful unit of measurement isn’t the keyword. It’s the persona.
That means the right question isn’t “what’s the LLM ranking factor?” It’s “for this industry, for this persona, what actually distinguishes the domains that get recommended?”
SEO Principles Still Apply
Search isn’t dead. It’s one layer of AI recommendations.
The data shows a consistent relationship between search visibility and LLM recommendations, and models are already searching the web live while answering. That makes SEO fundamentals more important, not less: authority, relevance, crawlability, citation breadth, and useful content still feed the evidence layer LLMs draw from.
The mistake is treating search rank as a stand-in for AI visibility. It’s one signal, not the answer.
Personas Change The Measurement
You can’t monitor AI visibility like a single search-ranking report.
In SEO, personalization was a modifier on top of a keyword rank you could still track. With LLMs, the persona and the context of who’s asking are the center of the recommendation, not a modifier on it. Airlines alone splits five different ways by persona in this report.
A single aggregate score isn’t a meaningful thing to monitor. If you’re tracking AI visibility without specifying who’s asking, you’re measuring noise.
Methodology & Dataset Scale
Every correlation in this report is drawn from a single pooled, all-industry dataset rebuilt from scratch for July 2026.
Models Sampled
LLM visibility scores are pooled across every model listed above, sampled against 403K+ prompts spanning 1,100 buyer personas across 100 industries. Backlink, search, and content signals are cross-referenced against nearly 15 billion crawled web pages, the full history of Reddit, all of Wikipedia, and 282B+ links.
How To Work With Me
Custom analysis for industries, portfolios, agency clients, link prospecting, and reputation monitoring. I’m taking on 5 clients at a time, expected turnaround is 4 to 8 weeks per project.
- Industry-specific AI visibility analysis: define industries and personas, get your sign-off, run the prompts, collect the data, deliver the reports.
- Multi-client agency research: repeat the process across several large clients or related industries.
- Large-scale link prospecting: Common Crawl, backlink graph data, and industry relevance scoring to prioritize outreach targets.
- Large-scale reputation monitoring: Reddit, search results, Wikipedia/Wikidata, web crawl data, and backlink signals combined to monitor category reputation at scale.
- Custom high-throughput data pipelines: process large datasets quickly without turning the project into a large cloud spend.
Fixed costs depend on scope and data sources, but the engineering approach is built to keep variable processing cost low, this project analyzed 3PB of data without a large server bill.
Get The Full Picture For Your Industry
This post is the free, all-industry summary plus one industry example. The live platform has this same depth of report, per industry, for all 100 industries, refreshed roughly every one to three weeks.
Launch pricing only, this rate won’t come back.
- $374.25/quarter for 3 industries
- $749.25/quarter for 10 industries
- $2,624.25/quarter for all 100 industries
This launch pricing ends 5 PM Pacific, July 31, 2026. LLM visibility, fanout queries, AI Overview citations, organic search, Reddit, Wikipedia, backlinks, and content ideas, uncapped and exportable.
OppAlerts.com • Ben Wills
Want this built custom for an industry that isn’t in the 100, or want it done as a consulting engagement? I’m taking 3 to 5 clients at a time, 4 to 8 weeks per project. OppAlerts.com/contact/
SEO & Engineering Background
Building the data pipeline, running the analysis, and turning it into strategy.
This project sits exactly where my background overlaps: SEO, search systems, high-throughput data engineering, and practical business reporting.
From about 2001 through 2013, I worked in the SEO industry: directed over 1,400 clients, a team of more than 70, and spoke at a number of the major conferences. I designed and executed SEO projects for Lowe’s Home Improvement, Motorola, Marriott, Salesforce, and others.
Since about 2012, I’ve been on the engineering side: everything from search engines indexing over 10M documents, to full-stack development, to a year writing embedded software for a high-end audio company.
If you’re looking for large-scale data collection and analysis like this, large-scale link prospecting, or large-scale reputation monitoring, let’s talk about your project.