Key takeaways
- A thumbnail prompt is a brief, not a wish. The models that generate images are optimised to make a good picture; a thumbnail is a different object, judged at 320 pixels wide next to eleven competitors, with a hole in it where your text goes.
- Google's own guidance for its image models says to describe the scene in sentences, give the image a stated purpose, control the camera with photographic language, and say what you want rather than what you do not. Keyword piles — "8k, ultra detailed, masterpiece" — are the habit to drop first.
- Resolution is a setting, not an adjective. In the Gemini API, image size takes 1K, 2K or 4K — roughly 1024, 2048 and 4096 pixels on the longest edge — and defaults to 1K. A 16:9 image at 1K is about 1024×576, which is below the 1280×720 YouTube asks for.
- Text rendering has improved enough to be tempting and not enough to be the plan. Generate the image, composite the type, because a thumbnail's words have to match your channel's font, dodge the duration stamp and survive a shrink no model is measuring.
- Prompt for a set, not a single image. Dynamic thumbnails and Test & Compare both consume three candidates, so the useful unit of work is one prompt skeleton with two or three slots deliberately varied.
- Everything Google's image models produce carries a SynthID watermark you cannot see and did not ask for, and the Gemini app will now tell a stranger it is there.
Ask any current image model for a YouTube thumbnail and you will get something competent and unusable. A face too small, a background too busy, text in a font that does not exist on your channel with a letter missing from the second word, and the whole thing composed as though someone will look at it for four seconds on a 27-inch monitor. It is a good picture. It is not a thumbnail.
The gap is not model quality. It is that a prompt describing a picture and a prompt describing a piece of packaging are different documents. The second one has to carry constraints the model has no way to infer: where the type will sit, which third of the frame must stay empty, what the image has to still read as when it is 168 pixels wide in a sidebar, and what emotion the face is meant to be selling. Those constraints are the job. Everything else is decoration.
This piece is about writing that document. The slots a thumbnail prompt needs, what the model genuinely controls and what it will quietly ignore, how to get a file big enough to upload, when to generate text and when to refuse, how to prompt your own face without it drifting into somebody else's, and how to write one brief that yields three variants worth testing rather than three renders of the same idea.
Why most thumbnail prompts fail
The typical first attempt reads like a search query: "youtube thumbnail, man shocked face, money, bright colours, 4k, hyper realistic, trending". Every term in it is a vote for a vibe, and none of it is an instruction. The model fills the frame because filling the frame makes a better picture. It centres the subject because centred subjects look resolved. It renders the face at the size a portrait wants rather than the size a thumbnail needs, and it treats "bright colours" as licence to use all of them.
A thumbnail has a job description that none of those words touch. It is shown small, often beside other thumbnails competing for the same glance, and it is almost always shown with a title underneath it doing the explaining. So the image does not need to explain. It needs one subject large enough to identify instantly, one piece of information the title does not already carry, enough tonal separation from the background that the subject survives compression, and clear space where words can land. The deep version of that argument is in the piece on thumbnail composition; the point here is that a prompt which does not state those four things has delegated them to a model with no stake in them.
What the model controls, and what it ignores
Prompt-writing gets much easier once you stop asking for the things a model cannot reliably give you. The split, roughly, looks like this.
| Instruction | How reliably it lands | What to do about it |
|---|---|---|
| Subject, expression, wardrobe, props | Strong — this is what the model is for | Be specific; name the emotion rather than "shocked face" |
| Lighting, contrast, colour mood | Strong | Describe the source and direction, not just "dramatic" |
| Camera framing, lens, angle | Strong — photographic vocabulary is understood | "Tight waist-up shot, 50mm, low angle" beats "good composition" |
| Keeping a region of the frame empty | Moderate — it works when you describe it positively | "The right third is flat dark wall with nothing on it" |
| Exact words, in an exact font | Weak to moderate, and worst at small sizes | Composite type afterwards; see below |
| Exact brand colours by hex value | Weak | Name the colour in words, correct it in an editor |
| Pixel dimensions requested in prose | Ignored — this is a parameter, not a prompt | Set aspect ratio and image size in the tool's controls |
That last row is the most common waste of a prompt. Typing "1280x720" into the text box does nothing in most tools. Size and shape are settings, and they have documented values.
Describe a scene; do not list keywords
Google's published best practices for prompting its Gemini image models are unusually direct about this, and they generalise well to every other model worth using. The best-practices documentation tells you to be specific rather than general — its example contrasts "fantasy armor" with "ornate elven plate armor, etched with silver leaf patterns" — to provide context and intent by saying what the image is for, to describe what you want instead of what you do not, to control the camera with photographic and cinematic terms, and to iterate rather than expecting the first render to land. Google's prompting tips for the Gemini app make the same case from the other end: a sentence or two will get you something decent, and the control arrives when you name subject, setting and composition, then lighting and colour, then the camera and its technical character.
Two of those rules do real work on thumbnails specifically. "Say what the image is for" is how you get a model to stop composing like a magazine cover: "a thumbnail for a YouTube video about losing a deposit on a flat" is context it will act on. And "describe what you want, not what you do not" is the fix for the empty-space problem. "No text in the image" is a weak instruction that models routinely violate. "The left half of the frame is a plain, unlit grey wall, completely bare" is a description of a thing, and things get rendered.
The six slots of a thumbnail brief
Everything above collapses into a template. Six slots, in this order, written as prose rather than a list of tags.
- Subject and emotion. Who or what is in frame, how it is lit by expression. Name the emotion precisely — "braced, mouth closed, eyes wide" is a face; "shocked" is a genre.
- One hero object. The single prop that tells the viewer what the video is about and that the title does not already say. One. The second prop costs you more than it adds.
- Background, stated as simple. A described surface, a stated colour, a stated depth of field. Rooms full of detail are where thumbnails go to die.
- Composition and the empty region. Where the subject sits, which part of the frame stays bare, and how tight the crop is.
- Light and separation. The source, its direction, and the contrast between subject and background. This is the single biggest driver of whether the image survives being shrunk.
- Purpose and finish. That it is a YouTube thumbnail, and the look you want — photographic, illustrated, flat-graphic — in a couple of words rather than a stack of quality adjectives.
As a skeleton you can keep in a text file:
A YouTube thumbnail for a video about [topic]. [Subject], [precise expression], [crop and angle], positioned on the [left/right] side of the frame. [One hero object] in [his/her/its] [hands / position]. Background: [one simple described surface], [colour], softly out of focus. The [opposite] third of the frame is empty [surface] with nothing on it. Lighting: [source and direction], high contrast between subject and background. Photographic, clean, no text.
It looks mechanical because it is. The variation that matters happens in the brackets, and keeping the frame identical across a channel is how a set of thumbnails ends up looking like a set — the argument made at length in the piece on thumbnail consistency.
Resolution and shape are settings with documented values
This is where more prompts fail than anywhere else, and it is the easiest thing to fix. In the Gemini API the image configuration exposes two controls that matter: aspect ratio and image size. Aspect ratio accepts 1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9 and 21:9. Image size accepts 1K, 2K and 4K, which correspond to roughly 1024, 2048 and 4096 pixels on the longest edge, and the default when you say nothing is 1K.
Work that through for a thumbnail. 16:9 at 1K is about 1024×576 pixels. YouTube asks for 1280×720. You have generated an image that is below spec before you have touched it, and upscaling a generated image is a worse outcome than generating at size. Ask for 2K and you get roughly 2048×1152, which downsamples to 1280×720 cleanly and leaves headroom for the crop you will inevitably want. 4K is useful when you intend to crop hard into a region, and otherwise it is a large file to carry around — remember the upload ceiling is 2 MB, so the final export is a JPEG decision as much as a resolution one. The mechanics of that export, and the ways a technically correct file still arrives looking soft, are in why thumbnails come out blurry and the thumbnail size guide.
Other models have their own geometry, and it is worth knowing before you fight it.
| Model | Shape control | What that means for a 16:9 thumbnail |
|---|---|---|
| Gemini image models | 16:9 supported directly; size 1K / 2K / 4K | Ask for 16:9 at 2K and you are above spec with room to crop |
| OpenAI's GPT Image | 1024×1024, 1536×1024, 1024×1536 or auto | 1536×1024 is 3:2, so a 16:9 thumbnail means cropping the top and bottom — compose for it |
| Midjourney | --ar parameter, default 1:1 |
Add --ar 16:9 explicitly or you are cropping a square |
Compose for the crop you know is coming
If your model only produces 3:2 or 1:1, say so in the prompt and design around it: ask for the subject to sit in the middle band of the frame, with the top and bottom of the image carrying nothing you need. A 16:9 crop out of a 3:2 render loses roughly a sixth of the height — on a 1536×1024 image, 160 pixels. Models will happily put your hero object exactly where the crop will remove it.
Text: generate the image, composite the type
Text rendering is the capability that has improved most visibly. Google positions Nano Banana Pro, its Gemini 3 Pro image model, specifically on state-of-the-art text rendering and the ability to produce posters, labels and packaging mock-ups with legible written content. OpenAI's GPT Image model is pitched squarely at work where in-image text has to come out readable — menus, labels, diagrams. Midjourney's v7 handles text when you wrap the exact string in quotation marks. So yes, you can ask for words now and often get them.
You should still composite the type yourself, for reasons that have nothing to do with model quality. Your thumbnail text has to be in your channel's typeface, at a weight that holds at 168 pixels, positioned clear of the bottom-right corner where the duration stamp sits, and identical in treatment to the last forty videos. None of that is something a generative model is optimising for, and all of it is trivial in an editor. Generated type also cannot be edited — change one word and you re-roll the whole image, which is an absurd price for fixing a typo.
The practical workflow is to ask the model for a clean image with an explicit empty region, then set the words on top. Our free add text to thumbnail tool does exactly that step, and the question of which typeface survives the shrink is covered in the best fonts for thumbnails. If you do generate text — a sign in the scene, a number on a screen, a label on a box — keep it to two or three words, quote the exact string, and inspect it at full size, because a single malformed letter in otherwise excellent artwork is the tell that reads as carelessness.
Prompting your own face
A face is the strongest attractor a thumbnail has, and most creators want their own. Current models take reference images for this, and the quality of what comes back is determined far more by the references than by the prose. Practitioner guidance converges on the same advice: supply several high-resolution shots from different angles rather than one, state explicitly that facial features must be preserved exactly as in the references, and refine in several passes instead of expecting a single generation to nail the likeness.
Which means the leverage is upstream. A thirty-minute session shooting yourself against a plain wall in even light, with five or six expressions and a couple of angles, is worth more than any prompt you can write — and those frames are directly usable as thumbnail material anyway, which is the subject of how to take photos for thumbnails. Prompt the model to place and light a face it has been shown, not to invent one that resembles you.
Other people's faces are a different question entirely, and the answer is mostly no: likeness, consent and YouTube's own policies all converge there. The policy detail — what counts as realistic altered content, where disclosure applies, and the specific ways AI thumbnails get channels in trouble — is covered in are AI thumbnails allowed on YouTube.
Prompt for three, not for one
The unit of thumbnail work changed this year. Test & Compare consumes up to three images and finds one winner; the dynamic thumbnails announced at Made on YouTube in September 2026 take three options and show a different one to a different segment of your audience. Either way, a single perfect image is the wrong deliverable.
The mistake is to generate three renders of the same prompt and call them variants. Seed noise is not a hypothesis. A useful set varies exactly one meaningful thing per candidate, so that whatever wins tells you something you can reuse:
- Emotion. Same framing, same object, same light: braced versus delighted versus deadpan.
- Distance. Waist-up versus a tight head-and-shoulders crop. This is the variable that most often moves the number, because it changes how much survives the shrink.
- Hero object. The thing versus the outcome — the broken part versus the finished build.
- Colour field. One background family against another, holding everything else constant.
Write the skeleton once, change one bracket per candidate, and keep a note of which bracket you changed. Three months of that and you have a channel-specific answer to a question no general advice can give you. The testing mechanics — how long to run, what counts as a result, where the trap is — are in the Test & Compare guide.
A worked example
Take a video titled "I rebuilt my studio for £400". The first-instinct prompt:
youtube thumbnail, man in studio, shocked, budget setup, money, bright, 4k, hyper detailed, professional
What comes back is a symmetrical shot of a man in the middle of a cluttered room, a vague pile of cash somewhere, the whole frame busy, nothing legible below about 400 pixels wide. Rewritten as a brief:
A YouTube thumbnail for a video about rebuilding a home studio on a tight budget. A man in a plain dark t-shirt, waist-up, slightly low angle, positioned on the right side of the frame, mouth closed and eyebrows raised as if reassessing something. He is holding a single cheap desk microphone up at chest height. Background: a bare light-grey wall, softly out of focus, one warm practical light glowing in the far corner. The left third of the frame is empty wall with nothing on it. Lighting: a single soft key from the left, strong separation between the subject and the wall. Photographic, clean, no text.
Five things changed, and each maps to a slot: the purpose is stated, the emotion is described rather than labelled, there is exactly one prop, the empty region is positively described so the title text has somewhere to go, and the background is specified as simple instead of being left to the model's taste for detail. Then three variants: swap "eyebrows raised as if reassessing" for "half-smiling, pleased"; swap the waist-up crop for "tight head-and-shoulders, filling the right half"; swap the microphone for "a stack of three mismatched second-hand monitors". Same frame, three different arguments.
Iterate; do not re-roll
Google's guidance says to refine with follow-up instructions rather than rewriting from scratch, and this is the habit that separates people who get usable images in four minutes from people who burn forty. If the image is 80% right, say what to change: make the light warmer, move the subject further right, flatten the background, lose the second object, raise the contrast. Conversational editing preserves the parts that worked. Re-rolling the whole prompt throws away a result you were nearly happy with in exchange for a fresh set of problems.
Keep a short list of your own corrections, because they repeat. Most creators find they are typing the same four fixes — too busy, too small, too centred, too dim — which is a sign those four constraints belong in the skeleton rather than in the follow-up.
Things to stop typing
- Quality adjectives. "8k", "masterpiece", "ultra detailed", "award winning". These were prompt-engineering folklore for a generation of models that no longer ships; current models read them as noise at best and as licence to over-render at worst.
- Negations. "No text", "not cluttered", "without a background". Describe the positive state instead.
- Pixel dimensions in prose. Use the aspect-ratio and size controls.
- Named living artists and studios. A legal and policy risk with no upside for a thumbnail, and a fast route to a style that is not yours.
- Everything at once. Three props, two faces, a logo and a caption in a 1280×720 rectangle is four elements more than it holds.
What the generator inside Studio changes
YouTube now does some of this itself. Ask Studio can generate a thumbnail for a long-form video from what the platform already knows — the title, description, transcript and your channel's visual history — and aims to match your usual style rather than produce a generic image. Reporting on the rollout, including TubeBuddy's write-up, describes it as desktop-browser only, arriving in stages, and excluding Official Artist Channels, channels set as made for kids, and creators under 18.
It is worth using and worth understanding the shape of. A first-party generator optimises for plausibility and consistency with what you have already published, which is exactly what you want for a routine upload and exactly not what you want when you are trying to break a pattern that has stopped working. It also gives you fewer handles: you are not writing the brief, so you cannot vary one slot deliberately. The feature set and its limits are covered in the pieces on Ask Studio's AI tools and everything announced at Made on YouTube 2026.
The watermark you did not ask for
One thing to know before you publish a generated thumbnail: Google embeds a SynthID watermark in media its AI models generate or edit. It is invisible to a viewer and designed to be hard to remove without degrading the image, and the Gemini app now checks images for it on request — anyone can upload your thumbnail and ask whether it was made with Google AI. The detection only covers Google's own watermark, so it is not a universal AI detector, but it does mean "nobody can tell" is no longer a safe assumption for anything you generated with Gemini.
This is not a reason to avoid generated thumbnails. It is a reason to be relaxed about them: use them where they are honest, disclose what the rules require, and do not build a workflow that depends on a viewer never finding out. The policy boundaries are worked through in the AI thumbnail rules piece.
The checklist
- Is the purpose stated in the prompt — that this is a thumbnail, for a video about a specific thing?
- Is there exactly one subject and one hero object?
- Is the empty region described as a thing rather than requested as an absence?
- Is the emotion described rather than named?
- Is the background specified as simple, with stated colour and depth of field?
- Is lighting direction and subject-to-background contrast in the prompt?
- Is the aspect ratio set to 16:9 in the controls, and the size at least 2K rather than the default?
- Is the type being composited afterwards rather than generated?
- Do you have three candidates that each vary one deliberate thing?
- Does the final file come in under 2 MB without visible artefacts?
Where this leaves the work
The skill that used to matter in thumbnail production was execution: masking a subject cleanly, balancing a composition, setting type that holds at small sizes. Generation has eaten a chunk of that and left something narrower and more valuable in its place — the ability to specify. Knowing that the left third must be empty, that the crop should be tighter, that one prop is the limit, and being able to say so in a sentence a model will act on.
None of which is really about prompts. It is the same brief a good designer would have asked you for and that most creators have never written down. The models have simply made the cost of not writing it visible, in the form of forty mediocre renders and an afternoon gone.
If you would rather not maintain a prompt skeleton at all, that is the problem Thumblore exists to remove: it takes the video, not the brief, and returns thumbnails built to the constraints above — 16:9, text you can edit, a subject that survives the shrink, and variants designed to be tested rather than admired. Either way, judge the result at the size a viewer will see it, which the thumbnail preview tool will do for free, and start from the 2026 thumbnail playbook if you want the design rules the prompt is trying to encode.