This chapter right-sizes its own topic. The industry pitch says video is the future of AI search and brands that skip it will vanish from the answers. The citation data says video barely appears in them. Both halves of the work here are honest sizing: what images, video, and audio actually do for AI visibility, and what budget that justifies.
What the citation data says about video
In OppAlerts' study, video draws 0.5% of what the search-grounded models cite, against 9.2% for Reddit and 2.6% for Wikipedia, and most of that sliver comes from one model.From the citation analysis in the OppAlerts study, summarized in What Actually Correlates with AI Search Visibility; method and tables in The Research Behind This Guide, part of the AI Search Visibility research. Even the model that does read video at meaningful volume does not convert that reading into recommendations: video presence fails to add independent predictive power for who gets recommended.
Video presence still correlates with AI recommendations overall, and the reason is the trap this chapter exists to flag. Famous brands dominate the video shelf, so a brand with a big video footprint tends to be recommended, the way a brand with a big billboard budget tends to be recommended. The footprint mirrors the fame; nothing in the data shows it driving the recommendation. In the mirror-versus-lever language of What Actually Correlates with AI Search Visibility, video is the clearest mirror in the study.
One boundary on that assessment: it describes chatbot-style answers. Inside Google's own products, the picture differs; Pew found Wikipedia, YouTube, and Reddit are the most commonly linked sources in Google's AI summaries, together about 15% of listed sources.Pew Research Center, analysis of Google AI summaries from March 2025 browsing data. Google owns YouTube, which is the obvious reading of the gap. If Google's AI products matter for your category, YouTube presence earns a larger weight than the chatbot data alone implies.
Where multimodal wins
Ruling out citation miracles leaves the real cases, and they are worth naming plainly.
- Visual intent: some queries are inherently visual ("mid-century living room ideas", "what does a deer tick bite look like"), and answers to them include images. Product imagery also feeds shopping results, where a listing without good images loses regardless of any model's opinion; the commerce side is in Commercial Visibility: Products, Comparisons, and Agentic Buying.
- Instructional intent: how-to queries retrieve video because video genuinely answers them, and a good walkthrough gets linked in AI summaries for those queries.
- Entity recognition: logos, product shots, and consistent visual identity across your properties help systems bind images to your entity, part of the work in Entities: Becoming a Thing the Machine Knows.
- The audience itself: people watch the videos and listen to the podcasts. Demand generated there shows up later as branded prompts and community discussion, which are levers this book covers elsewhere. Media that serves an audience needs no citation to justify itself.
Transcripts and metadata: the indexable payload
Whatever media you produce, the machine-consumable layer is text, so ship the text. Models and search pipelines work mostly from transcripts, captions, titles, descriptions, and surrounding page copy, not from watching pixels. A video without a transcript is, to most of these systems, a title and a thumbnail.
The checklist: publish full transcripts on the same page as the embedded video, caption everything, write titles and descriptions that state the actual subject in category vocabulary, use VideoObject structured data per Structured Data and Machine-Readable Facts, and give images descriptive alt text and file names. Podcast episodes get show-notes pages with transcripts, which turn audio into crawlable pages. This is the cheapest work in the chapter and the only part that makes the rest legible to machines; a transcript is also co-occurrence text, your brand next to category vocabulary, feeding the same corpus mechanics as Links: The Signal That Refuses to Die.
A right-sized multimodal budget
Fund multimodal in this order. First, the text layer for everything you already produce: transcripts, captions, metadata, markup. It is cheap and it is the only part with a direct retrieval mechanism. Second, media that serves demonstrated audience intent, visual and how-to queries in your category, sized by whether those queries matter to your revenue. Third, brand media for its own audience, judged as audience-building, with any AI-visibility effect treated as a side benefit.
What the data does not support is producing video because AI visibility demands it. It does not, and the correlation that seems to say otherwise is fame reflected back. Spend the difference on the levers with independent signal behind them: the links, community, and press of the three chapters before this one.