Key takeaways
- YouTube is now the most-cited domain in Google's AI Overviews. Ahrefs' tracking of AI Overview citations put youtube.com top of the list through 2026, at roughly a fifth of the citations among leading sources, with Reddit second and every news publisher well behind both.
- Being cited and being watched are different events. Pew Research Center found people clicked a result in 8% of visits where an AI summary appeared, against 15% where it did not, and clicked a link inside the summary itself in about 1% of visits.
- The retrieval layer reads your transcript, title and chapters. It does not look at your thumbnail at all. Two different audiences now decide whether a video gets found, and only one of them has eyes.
- Thumbnails still decide the surfaces that actually deliver views — Home, Suggested, search results, Discover cards and the video carousels Google is testing inside AI answers. The thumbnail's job did not shrink; a second, text-only job was added beside it.
- Almost everything sold as "AI search optimisation" for video is ordinary packaging hygiene: say the answer out loud, caption it accurately, chapter it at the answer boundaries, and write one video per question.
- Google Discover started carrying YouTube content and a follow button in September 2025, and it is the external surface most creators never check in Analytics.
For twenty years the question "how do people find my video" had two answers that mattered: inside YouTube, and through Google. Both worked the same way. A query went in, a ranked list came back, and the list was made of titles and thumbnails, which meant packaging decided everything after ranking. Every piece of advice written for creators since 2012 assumes that shape.
That shape now has a third layer stacked on top of it, and the layer does not render images. An AI summary answers the question in prose, cites a handful of sources, and hands back one link if the reader wants more. Google said at I/O in 2026 that AI Overviews had passed 2.5 billion monthly users and that AI Mode had passed one billion. Whatever you believe about how good those answers are, that is not a test surface any more.
The surprising part, for creators, is who is winning it. The most-cited domain in Google's AI Overviews is not a newspaper, an encyclopaedia or a forum. It is YouTube. That fact is worth understanding precisely, because it comes with a catch that decides whether any of it is worth your time.
Why the answer machines reach for video
Ahrefs has been tracking which domains AI Overviews cite, ranking them by share of citations across a large set of United States queries — more than three million queries in its mid-2026 run. Through 2026 youtube.com sat at the top of that list, at around a fifth of citations among the leading sources, with Reddit close behind and everything else well back. A separate study covered by Search Engine Land found the same pattern across AI engines generally: Reddit, YouTube and LinkedIn cited more than anything else.
Treat the exact percentage as a moving number measured one way by one company — the study updates as it re-runs, and "mention share among top sources" is a methodology, not a law of nature. The ordering is the durable finding, and the reasons behind it are mechanical rather than mysterious.
A retrieval system needs text it can quote with attribution. Every YouTube video carries one: the caption track, generated automatically when you do not supply your own. That transcript is often the only text in existence that answers a narrow practical question in sequence — how to reseat a toilet flange, what a specific error code means on a specific washing machine, which settings changed in a game patch. Written publishers do not cover most of that; creators cover it in fifteen minutes of speech, with the steps in order. For a system assembling an answer, a transcript reads in milliseconds and comes pre-attributed to a source with a public URL.
There is a structural advantage too, and it would be strange to pretend otherwise: Google owns YouTube, indexes it completely, and does not need permission or a crawl budget to read it. Nothing about that is available to you as a lever, but it explains why the ceiling for video citation is high and stable rather than something a policy change is likely to take away.
A citation is not a view
Here is the catch, and it is the reason to read the rest of this with a cold eye.
In July 2025 the Pew Research Center published an analysis of 68,879 Google searches made by 900 United States adults who had agreed to share their browsing activity. Where an AI summary appeared, users clicked through to a website in 8% of visits. Where no summary appeared, they clicked in 15%. The share who clicked a link inside the AI summary itself was about 1%.
Google disputed the study publicly, calling the methodology flawed and the query set unrepresentative of real Search traffic. That objection is not unreasonable — a panel's queries are not Google's query mix — and Pew's numbers are one measurement rather than the settled truth. But no serious reading of the evidence lands on "AI summaries send more clicks than a list of ten links did". The honest summary is that citations are cheaper to earn than views, and convert into views at a much lower rate than a ranked result used to.
So the strategic conclusion is narrow. Getting quoted in AI answers is worth doing when the work is work you should do anyway. It is not worth reorganising your channel around, and it is certainly not worth paying an agency for.
The surfaces, and what each one actually pays
"Being found on Google" is now at least four different events with different mechanics, sitting alongside the two surfaces inside YouTube itself. Separating them is most of the clarity available here.
| Surface | What the viewer sees | Thumbnail shown? | What a hit is worth |
|---|---|---|---|
| AI Overview / AI Mode citation | Prose answer, your channel as one of several cited sources | No | Low click rate, occasional high-intent visitor, brand exposure you cannot measure |
| Video carousel inside an AI answer | Tappable clips with AI-written descriptions | Yes, small | A real view, often starting mid-video at the cited moment |
| Google video results and key moments | Thumbnail, title, channel, seekable chapter links | Yes | A full click-through, the closest thing to old-fashioned search traffic |
| Google Discover | Large card in a scrolling feed, no query behind it | Yes, large | Volume in bursts, low intent, thumbnail-decided |
| YouTube search and its AI carousel | Ranked results, sometimes an assembled clip carousel above them | Yes | Durable, compounding traffic — still the best-value search surface you have |
| Home and Suggested | Thumbnail and title in a grid | Yes | The majority of most channels' views, decided almost entirely by packaging |
Read down the thumbnail column. One row out of six does not use your image — and it is the row with the lowest click rate attached to it. That single observation should settle most of the anxiety currently being sold to creators about AI search making thumbnails irrelevant. What has actually happened is that a text-only reader was added to an audience that was previously all eyes, and the text-only reader controls a small, growing share of discovery.
Where the thumbnail is not even rendered, and where it still decides
It helps to be concrete about the mixed cases, because they are the ones changing fastest.
Adweek reported that YouTube has been testing video carousels inside Google's AI Overviews for product and location queries — the sort of question where someone wants to see the thing, not read about it: best noise-cancelling headphones, museums worth visiting in a city. Those carousels show thumbnails at small size with machine-written descriptions beside them. YouTube has run the equivalent inside its own search results too, an AI-assembled carousel of clips with short descriptions, tested first with Premium members in the United States.
Two things follow. First, on these surfaces your thumbnail competes at a smaller size than anywhere else, against a text description you did not write and cannot edit. A thumbnail that depends on a readable six-word headline loses there; a thumbnail with one legible subject and hard contrast survives. The discipline is the same one the television screen already demanded, applied at the opposite end of the size range. Second, the clip being surfaced may be from the middle of your video, which means the segment does the packaging job the thumbnail normally does.
The one-sentence version
The retrieval layer decides whether your video is eligible to be shown. Your thumbnail still decides whether a human picks it once it is. Optimising for the first while neglecting the second is a way to be cited constantly and watched rarely.
Discover is the surface creators under-read
In September 2025 Google announced that Discover — the queryless feed on the Google app's home screen — would start carrying more creator content, including YouTube Shorts and posts from other platforms, along with a follow button for publishers and creators. Reporting through 2026 has described YouTube absorbing a substantial share of Discover's positions as that change rolled out.
Discover matters for three reasons. It is enormous. There is no query behind it, so intent is low and the card is doing all the persuasion, which makes it the most thumbnail-dependent external surface in existence. And there is nothing to submit, optimise or claim: eligibility follows from publishing on YouTube, and the feed decides.
What you can do is notice it. In Studio, Analytics under the Reach tab lists external traffic by site, and Google properties show up as their own rows there. A video with an unusual spike from an external Google source has usually been picked up by Discover rather than by search, and the two behave nothing alike: search traffic accumulates slowly for years, Discover arrives in a day and leaves. Treating a Discover burst as evidence that a format is working is one of the easier ways to mislead yourself. Our traffic sources guide covers how to read each bucket without drawing that kind of false conclusion.
What makes a video quotable
Everything below is retrofittable to videos you have already published, and none of it requires believing any specific claim about how retrieval systems rank sources. The logic is only this: systems that quote video quote its text, so the video's text should contain a clean answer.
Say the answer out loud, early, in one self-contained sentence
The most common reason a genuinely useful video never gets quoted is that the answer only exists as a demonstration. You show the setting being changed; you never say "the setting is under Advanced, and it is off by default". A sentence that can be lifted out and still make sense is a quotable sentence. A sentence built on "this one" and "like I said" is not.
This costs nothing and helps humans too, which is the test for every item in this section. Viewers who arrive from a search want confirmation in the first fifteen seconds that they are in the right place, and the same sentence does both jobs. Our guide to the first thirty seconds goes further into how to do that without giving away the payoff.
One question per video
A twenty-minute video that answers nine loosely related questions is hard to cite for any of them, and hard to title. Videos that get quoted tend to be narrow: one question, answered completely, with the scope stated. This is the same instinct that makes a video rank in YouTube search, which is not a coincidence — both systems are trying to match a specific need to a specific artefact.
Treat captions as the document, not an accessibility afterthought
Automatic captions are good and still wrong in exactly the places that matter: product names, version numbers, units, jargon, anything said quickly. Those misheard words are the words a retrieval system would have keyed on. Reviewing and correcting the caption track for your top handful of explainer videos is an hour of work with a permanent effect, and it improves the viewing experience for everyone watching without sound. The captions guide covers the mechanics.
Chapter at the answer boundaries
Chapters are how a long video becomes several citable segments. Google has surfaced key moments from video chapters for years, and the mechanics are strict but simple: the first timestamp must be 00:00, there must be at least three chapters, and each must run at least ten seconds. Studio can generate them automatically, and manual timestamps in the description override the automatic set.
The part people get wrong is where to put the boundaries. Chapters named for the structure of your script — intro, background, main part, outro — are worthless to a system trying to find the moment that answers a question. Chapters named for the questions they answer are the point. Chapters and key moments has the full treatment.
Use specifics in speech, not only on screen
Numbers, model names, settings paths, dates and units said aloud end up in the transcript. The same information shown only as on-screen text does not. Creators who edit heavily tend to move all the specifics into graphics, which looks better and makes the video invisible to anything reading the text. Say the number as well as showing it.
A forty-minute retrofit
Take your three highest-traffic explainer videos. For each: correct the caption track where it mangled a product name or number, rewrite the chapters so each one names a question, and add one sentence to the top of the description that answers the title's question in plain prose. That is the whole programme. Anything sold beyond it — schema packages for videos you do not host, "GEO audits", transcript keyword density — is either inapplicable or invented.
Titles that work for a reader with no eyes
A title is now read by two audiences with opposite preferences. The human scanning a feed responds to tension, specificity and a reason to care. The retrieval layer is matching a question to a document, and rewards a title that states the subject unambiguously.
Those are less contradictory than they sound, because the failure modes differ. Curiosity without a subject fails both: "I tried the thing everyone is wrong about" gives a machine nothing and gives a human no reason to trust the click. What works is a concrete subject plus one turn of tension — the subject carries the matching, the tension carries the click. Our titles guide works through the formulas, and the title checker shows where a long title truncates on each surface.
The description deserves one small change and no more. Its first two lines are the only part with real weight, and they should contain a plain-prose answer to the title's question rather than a list of links. Everything else in a description is housekeeping. The SEO guide covers how little the rest of those fields do.
How to measure this without fooling yourself
There is no report anywhere that says "cited in an AI Overview". Expect none. What you have is external traffic by source in the Reach tab, and a small number of honest inferences you can draw from it.
Check the shape rather than the number. Search-driven external traffic builds slowly and persists; feed-driven external traffic spikes and dies. Compare a video's external share against your own channel's normal, not against a benchmark from a blog post, because the mix differs wildly by niche — a repair channel and a vlog channel are in different businesses on this axis.
Then keep the decision small. If external traffic is a couple of per cent of your views, the correct response to a doubling is interest, not a strategy change. The retrofit above is worth doing because it is cheap and helps your existing viewers; it is not worth doing instead of making a better video. A useful annual habit here is the channel audit, where external sources get looked at once properly rather than checked anxiously every week.
What not to spend money on
The gap between a real shift and a sellable service is where most creator money disappears. Four specific things to skip.
- Schema markup for videos you do not host. Structured data describes pages on your own site. If the video lives on YouTube and you have no page of your own, there is nothing to mark up. Embedding your videos in written pages you own is genuinely useful — but that is a publishing project, not a plugin.
- "AI search optimisation" retainers for channels. Ask what the deliverable is. If it is captions, chapters and clearer titles, you now have that list for free. If it is something else, ask which documented mechanism it exploits.
- Transcript stuffing. Reading keyword lists aloud degrades the video for humans and produces exactly the low-effort signature YouTube's monetisation reviews look for. Our monetisation piece covers what that review actually samples.
- Publishing volume aimed at machines. A channel of thin explainers built to be quoted is a channel with no audience, and the surfaces that pay — Home, Suggested, the Shorts feed — are the ones that punish it hardest.
What this changes about packaging, and what it does not
The reasonable forecast is that the text layer keeps growing, keeps converting poorly, and keeps being worth exactly the hygiene it demands. Meanwhile the surfaces that deliver the overwhelming majority of watch time remain visual and get more crowded, not less: a Home feed on a television at four metres, a Shorts feed under a moving thumb, a Discover card between two news stories, a small carousel tile beside machine-written text. Every one of those is decided by whether one image reads instantly.
So the split is clean. The work that makes you quotable is textual, cheap, and finite — say the answer, caption it properly, chapter it honestly, title it clearly. The work that makes you watched is visual and never finishes, because it is competitive: your thumbnail is judged against whatever else is on the screen that second. Anyone telling you the second kind of work is over has confused a new distribution layer with a replacement for the old ones.
Frequently asked questions
Does appearing in an AI Overview bring views?
Sometimes, at a low rate. Pew's measurement had clicks on links inside AI summaries at around 1% of visits. The visitors who do come tend to arrive with a specific question, which makes them good viewers — but the volume is not comparable to a strong ranking in YouTube search.
Do thumbnails matter for AI search?
Not for the text answer itself, which never renders your image. They matter enormously for the video carousels appearing inside those answers, for Google's video results, for Discover cards and for every surface inside YouTube. The practical change is that thumbnails now have to survive being shown very small next to text you did not write.
Should I write blog posts to support my videos?
It is the one genuinely underused tactic here, because it gives you a page you control on a surface where your competition is text-first publishers with no video. Turn a tutorial into a written page with the video embedded and real text around it. That is a publishing commitment rather than a quick win, and it is separate from anything happening on YouTube.
Is YouTube's own AI search a different thing from Google's?
Yes. Ask YouTube, announced at Google I/O in May 2026, answers a natural-language question by assembling material from YouTube's catalogue, and it began with Premium members in the United States. It is a discovery surface inside the platform rather than on Google. Algorithm changes in 2026 covers what is known about it, which is less than most articles imply.
Will being cited by AI hurt my channel by answering the question for people?
For questions with a one-line answer, that traffic was always fragile — a summary answers "what year did X happen" and nobody needed your video. For anything requiring demonstration, judgement or a sequence of steps, text is a poor substitute and the citation is more likely to send someone looking for the video. This is a reason to make videos that show and argue rather than ones that recite facts.
The uncomfortable truth in all of this is how ordinary the advice turns out to be. The single biggest shift in search in a decade, and the response for a creator is: answer one question per video, say the answer out loud, fix your captions, name your chapters after questions. That is not an AI strategy. It is the same craft that made videos findable before any of this existed, applied to a reader who cannot see.
Which leaves the half of the job that never got easier. Once a system has decided your video is eligible to be shown — in a carousel tile, a Discover card, a search result, a Home feed on a television — one image has to do the convincing, at whatever size that surface allows, against everything else on the screen. If producing those images at speed is the step where your process breaks down, Thumblore is built for exactly that: concepts that read at feed size in minutes rather than an evening in an editor.
For the neighbouring pieces: the YouTube SEO guide goes deep on ranking inside YouTube's own search, and what counts as a good CTR explains why the click-through numbers on these surfaces are not comparable with each other.