Key takeaways
- A general chat assistant can make a very good image and still not hand you a thumbnail. The gap is not quality; it is shape, type, brand memory and the fact that the useful unit of work is now three variants, not one picture.
- Shape is the first thing to check. OpenAI's API reference lists the GPT image models' standard sizes as 1024×1024, 1536×1024 and 1024×1536 — 3:2 is the widest of those, so a 16:9 thumbnail means cropping or a newer model that accepts custom dimensions. Google's image models take an aspect ratio and a size as explicit parameters, and 16:9 is one of them.
- Both families now stamp what they make. Google puts an invisible SynthID mark in every image and, at the Nano Banana Pro launch, said the visible sparkle stays for free and AI Pro users. OpenAI says images from ChatGPT and its API carry C2PA Content Credentials plus SynthID.
- The thing neither can do is look like you. Channel recognition comes from the same crop, palette and type across sixty uploads, and a conversation has no memory of your other fifty-nine thumbnails.
- Use the chat models for what only they do: invent a scene you cannot photograph, replace a background, relight a bad frame, remove a thing. Then composite the type and build the set somewhere that knows what a thumbnail is.
- YouTube itself now generates channel-matched thumbnails in Studio and will show one of three to different slices of your audience, which changes what the chat route is competing against.
The question arrives in every creator forum in roughly the same words. You already pay for ChatGPT. You already have Gemini on your phone. Why use anything else to make a thumbnail? It is a fair question, and for the first time it has a serious case behind it: the current image models are genuinely good. They hold a likeness across edits, they composite several reference photographs into one scene, they render text that is readable rather than alien, and they do it in the same window where you were already writing your script.
Then you try to ship one. The file comes out 1536 pixels wide and 1024 tall, which is not 16:9. The words are set in a font that exists nowhere else on your channel, and one of them has a letter doing something strange at the baseline. The face is beautifully lit and too small to read at sidebar size. You ask for a second version and get a different person in a different world. And beside your last ten uploads, nothing about it says your channel made it.
This piece is about that decision, not about prompt craft — the deep version of the prompting argument is in the guide to writing AI prompts for thumbnails. Here the question is narrower and more practical: what a general-purpose model does well, exactly where it stops, what it quietly puts in the file you upload, and when the honest answer is that it is the right tool anyway.
What you are actually asking for
A thumbnail is not a picture of your video. It is a piece of packaging with a specification. It has to be 16:9, which in practice means 1280×720 or a clean multiple of it, and it has to arrive under YouTube's file-size ceiling — the numbers are in the thumbnail size guide. It has to survive being shown at around 168 pixels wide in an up-next sidebar, where a face smaller than about a fifth of the frame stops being a face. It has to carry two to four words that your channel always sets the same way. It has to leave the bottom-right corner alone, because YouTube stamps the duration there. And it has to look like it came from the same channel as the fifty before it, because that recognition is most of what a returning viewer uses to find you in a feed.
None of those constraints is something an image model is optimised for. It is optimised to produce a pleasing image at whatever size you asked for. Every one of the problems below is a version of the same mismatch: you are asking a very good general tool to hit a specification it was never given.
The shape problem, and the arithmetic of a crop
Start with the most mechanical failure, because it is the one that quietly ruins compositions. For OpenAI's GPT image models, the documented standard sizes are 1024×1024, 1536×1024 and 1024×1536, plus an automatic option that lets the model choose. The widest of those is 3:2. OpenAI's API reference also describes arbitrary width-by-height sizes for its newer image models, with edges divisible by 16 and aspect ratios anywhere between 1:3 and 3:1, so 16:9 is reachable — but it is something you have to ask for, and third-party guides to the ChatGPT interface all describe the same workaround: state the ratio in the prompt, because there is no documented ratio picker.
Google took the opposite approach. In the Gemini image API, aspect ratio and output size are separate explicit parameters: you pass a ratio such as 16:9 and a size of 1K, 2K or 4K. Google's own documentation table puts 16:9 at 1376×768 for 1K and 2752×1536 for 2K, which clears YouTube's 1280×720 comfortably. Third-party guides list slightly different pixel pairs for the same ratio, so check the table rather than trusting a blog — and note that the earlier Flash image model's 1K 16:9 output landed nearer 1024×576, which is below what YouTube asks for.
Here is why that matters more than it sounds. A 1536×1024 image cropped to 16:9 becomes 1536×864. You have thrown away 160 pixels of height, and you do not get to choose which 160 until after the model has composed the frame for a different shape. A subject the model placed with comfortable headroom for 3:2 ends up with its chin near the bottom edge or its hair cut off. A scene built around a centred vertical gets its breathing room removed. The composition you liked in the preview is not the composition you upload, which is the single most common reason a thumbnail from a chat window looks worse on the watch page than it did in the chat.
Do the crop before you judge the image
If you are generating in a 3:2 shape, crop to 16:9 first and then decide whether you like it. Judge every candidate in the shape it will actually be seen in, at the size it will actually be seen at. The thumbnail preview tool puts a candidate into the real surfaces for nothing, which is a faster verdict than any amount of staring at a full-size render.
What is in the file you upload
Both companies now mark their output, and the marks behave differently.
Google's position is the more visible one. Every image its models produce carries SynthID, an invisible watermark, and the Gemini app will tell anyone who asks whether an image appears to have been made by Google AI. At the Nano Banana Pro launch, Google said the visible sparkle watermark would remain on images from free and Google AI Pro users and be removed for AI Ultra subscribers and inside AI Studio. Third-party guides have since described a media-watermark setting in Gemini's app preferences, with regional limits and work-or-school accounts excluded; that detail is not something I can confirm from Google's own pages, so treat it as unverified and look in your own settings. For a thumbnail the practical point is blunt: a visible mark in a corner of a 1280×720 frame is competing with your own packaging for the viewer's one-third of a second, and it tells a prospective sponsor something about how the asset was made.
OpenAI's marking is quieter. Its help centre says images generated with ChatGPT, Codex and the API include C2PA Content Credentials and an invisible SynthID watermark, and it runs a verification page where an uploaded image can be checked for those signals. OpenAI is also unusually candid about the limits: it has written that metadata is not a silver bullet, that it can easily be removed either accidentally or intentionally, and that resizing, format conversion and screenshots are among the transformations that break it. Which, read from the other direction, means your own pipeline probably strips it. Resize a generated image to 1280×720 and compress it under the file-size ceiling with something like an image compressor, and the credentials may well not survive the trip.
None of this is a reason to avoid AI thumbnails — YouTube's disclosure rules are narrower than most creators assume, and the detail is in whether AI thumbnails are allowed on YouTube. It is a reason to know what is in the file before you upload it rather than after a viewer tells you.
Text is still the line
Text rendering is the capability that changed most between 2024 and now, and it is still the thing to stop asking for. The 2026 comparisons agree on the shape of the problem rather than the rankings: letter-level detail gets less reliable treatment than overall composition, so failures concentrate exactly where thumbnails live — small type, dense overlays, words sitting on a busy background. One test of several current models found a URL rendered with its dot dropped, and the reviewer's warning generalises well: plausible-looking characters are the ones you miss, so copy has to be checked character by character at full size rather than glanced at.
There is a second reason to composite type yourself, and it survives every improvement in model quality. Your channel's words are set in your channel's font, at your channel's size, in your channel's two colours, in the same corner every week. A model that generates type cannot be told "the usual" — and even when it renders flawlessly, it renders a font nobody else on your channel uses. Generate the image, then set the words in something that keeps a font and a palette between sessions, whether that is a design tool, the free add-text-to-thumbnail tool, or Thumblore. The argument for which font, and why most thumbnails use one weight too light, is in the guide to the best fonts for YouTube thumbnails.
Your own face, and the refusal nobody plans for
Most thumbnails on most channels contain the creator's face, which puts the chat route straight into the part of these products governed by policy rather than capability.
The two companies have moved in opposite directions, and both move often. Google's model declines recognisable public figures from a text prompt; reporting in 2025 found that uploading an official portrait could get a likeness generated anyway, and that the gap appeared to be closed within days. OpenAI went the other way: with its GPT-4o image generation it moved from blanket refusals towards allowing public figures with an opt-out for individuals, and journalists stress-testing it found enforcement inconsistent — the same request refused in one phrasing and produced in another.
For a creator, the failure that actually bites is the false positive. There are repeated reports of Gemini refusing to edit a user's own photograph because it decided the person in it was a public figure. It will be tuned, but it is a bad property in a tool you depend on for a weekly deadline: the pipeline works for months and then declines your face on a Thursday for a reason the interface cannot explain. If your packaging depends on your own likeness, keep a photographed fallback.
Someone else's face is a different question with a harder answer, and YouTube now has machinery pointed at it — see YouTube's likeness detection. The short version is that generating a recognisable person who did not consent is the risk that outlives the thumbnail.
The part a chat window cannot do: look like you
This is the real limitation, and it is structural rather than technical. Channel recognition is the product of repetition: the same subject size, the same crop logic, the same two or three colours, the same type treatment in the same position, upload after upload. It is what lets a returning viewer spot you in a feed of eleven competitors before reading a single word, and the mechanics of it are in the piece on thumbnail consistency.
A conversation has no durable idea of your channel. You can paste a reference thumbnail and ask for something in the same style, and a current model will do a creditable job for that one image. Next week, in a new chat, you do it again — and you get a near miss, because "in the same style" is being re-interpreted from scratch every time. Over twenty uploads, a series of near misses is not a style; it is drift. The tools that solve this store the thing rather than describing it: a saved palette, a locked font, a template with slots. That is a product feature, not a prompt, which is why it is the one gap better models do not close.
One image stopped being the unit of work
The other thing that changed underneath this question is what YouTube does with thumbnails. Test & Compare has let creators run three variants against each other since 2024, and YouTube has said creators have used it to run more than 40 million tests on titles and thumbnails. At Made on YouTube 2026 it went further: Studio can generate thumbnails and titles from a video's content matched to the look of your existing channel, dynamic thumbnails show one of three images to different segments of your audience, and — with the creator's permission — Ask Studio can monitor thumbnail performance on older videos and run its own refresh tests in the background. The full account of those announcements is in what Made on YouTube 2026 actually changed.
Three things follow from that. The first is that the deliverable is a set of three deliberately different candidates, not one hero image, and a chat model's natural output is one image plus two re-rolls of the same idea. The second is that the thumbnail decision is increasingly made by a system with your click-through data in front of it, which no chat window has. The third is less obvious and more important: if Studio is going to generate channel-matched candidates for free, the chat route is no longer competing with a designer's day rate. It is competing with a button. How to read what comes back from a test without fooling yourself is in the guide to Test & Compare.
What the chat models are genuinely better at
All of which could read as a case against them, and it is not. There is a category of thumbnail work where a general-purpose image model is the best tool available to a creator at any price, and it is worth being precise about what it is.
- Scenes you cannot photograph. A flooded living room, your kitchen in 1974, a product the size of a building. Nothing else gets you a plausible version of that in a minute.
- Backgrounds and extensions. Taking a usable photograph of yourself and putting it somewhere else, or extending a frame that was shot too tight for 16:9.
- Repair. Relighting a face shot against a window, removing the chair leg behind your head, cleaning up the one frame that had the right expression.
- Composition from references. The current Google model is built to take several reference images and hold their subjects consistent in one output, which is exactly the problem a two-person reaction thumbnail poses.
- Thinking. Giving you six visual ideas for a video about a dull subject is a language task, and a chat assistant is better at it than any template library.
Notice what every item has in common: it produces source material. None of them produces a finished thumbnail, and treating the output as the deliverable rather than the raw material is the mistake that makes this route look worse than it is.
The four routes, compared honestly
| What you need | Chat assistant (ChatGPT, Gemini) | Design tool (Canva, Photoshop) | Studio's built-in generator | Thumbnail tool |
|---|---|---|---|---|
| 16:9 output at upload size | Yes, if you set it — not always the default | Yes, by preset | Yes | Yes |
| Invents a scene that does not exist | Best in class | Only via its own AI features | Limited to your video's content | Varies by tool |
| Type in your font, same place every week | No | Yes | Channel-matched, not controlled by you | Yes |
| Remembers your channel between sessions | No | Yes, if you build the template | Yes, from your back catalogue | Yes |
| Produces three deliberately different variants | Possible, but you drive it | Manual work × three | Yes, that is the point of it | Yes |
| Marks on the file | SynthID; C2PA on OpenAI output; visible sparkle on some Google tiers | Depends on the feature used | Google provenance signals | Depends on the model behind it |
| Sees your click-through data | No | No | Yes | No |
Read down the columns rather than across the rows and the division of labour is obvious. The chat assistant wins one row outright and loses the rows about repetition. Studio wins the rows about data and loses the row about control. A tool built for the job wins the boring rows, which happen to be the rows you hit fifty-two times a year. A fuller like-for-like comparison of the dedicated tools, with current pricing, is in the best thumbnail generator in 2026.
A hybrid workflow that takes about twenty minutes
- Decide the idea before you open anything. One subject, one piece of information the title does not carry, and the third of the frame you are keeping empty for words. If the idea is not decided, no tool helps.
- Generate the scene, not the thumbnail. Ask for 16:9 explicitly, at the largest size the model offers, with no text in the image and clear space where your type will go.
- Crop and judge in the real shape. If the model handed you 3:2, crop to 16:9 now, and reject candidates whose composition did not survive it.
- Composite the type where your font lives. Two to four words, your weight, your colours, your corner, clear of the bottom-right duration stamp.
- Shrink it to about 168 pixels wide and look again. If the subject is unreadable or the words are a grey smear, the fix is a bigger subject and fewer words, not a better model.
- Build the other two variants by changing one thing. One expression, one background, one claim — one variable per variant, or the test teaches you nothing.
- Export under the file-size ceiling and test. Then log what won, because the point of all of this is the next thumbnail, not this one.
Steps two and five are where a general-purpose model earns its place. Steps four and six are where it costs you the most time, and they are also the steps that repeat every week forever.
What this costs, and why not to anchor on it
Per-image prices are public, volatile and not the number that matters. Third-party trackers citing Google's pricing page put its Pro image model at roughly thirteen cents per 1K or 2K image and about twice that at 4K; OpenAI's smaller image tiers have been quoted at well under a cent. Both companies ship a new image model every few months, so a workflow built around today's price is built on sand.
The cost that persists is the twenty minutes. Three test-ready variants a week is roughly an hour of packaging work, every week, forever, and the route you choose mostly decides how much of that hour is spent on decisions and how much on re-cropping, re-setting type and re-explaining your style to a tool that has forgotten it. If you are producing one thumbnail a month, the chat assistant you already pay for is the obvious answer and this entire article is an argument against something you should ignore. At one a week, the arithmetic turns.
Five ways this goes wrong
- Treating the render as the deliverable. The output is source material. The thumbnail is what you make from it.
- Judging candidates full size. Every thumbnail looks good at 1536 pixels wide. The decision happens at 168.
- Letting the model set the type. Even when the letters are perfect, the font is wrong, because it is not the one your channel uses.
- Re-rolling instead of iterating. Three renders of the same idea is one candidate. Three variants that differ in one deliberate way is a test.
- Starting a new chat every week. The style drift creeps in slowly and shows up as a feed that no longer looks like one channel.
Frequently asked questions
Can ChatGPT make a YouTube thumbnail?
It can make the image. You will need to ask for a 16:9 shape explicitly, because the GPT image models' documented standard sizes top out at 3:2 in landscape, and you will want to add the text yourself so it matches the rest of your channel. For a one-off that is perfectly workable. For a weekly upload the type and consistency steps are the ones that wear you down.
Is Gemini or ChatGPT better for thumbnails?
On shape, Google's models are more convenient: aspect ratio and output size are explicit parameters and 16:9 is among them. On marks, OpenAI's output is less conspicuous — Google said at launch that the visible sparkle stays for free and AI Pro tiers. Both are strong at image quality now, and the difference between them matters less than what you do to the output afterwards.
Will YouTube penalise an AI-generated thumbnail?
Not for being AI-generated. The policies that apply are about misleading packaging, other people's likenesses and the usual content rules, which apply to a photograph just as much. The specifics are in are AI thumbnails allowed on YouTube.
Does the invisible watermark hurt my thumbnail?
No. SynthID and C2PA credentials are provenance signals, not quality penalties, and OpenAI itself notes that resizing and format changes can strip metadata anyway. A visible watermark is a different matter: it takes up frame you need.
Should I just use Studio's generator instead?
Try it, because it is free and it has your back catalogue and your click-through data, which nothing else does. What you give up is control — it matches your existing look rather than letting you decide it, and it works from the video's own content, so it cannot invent the scene you did not film.
Sources
Model capabilities and policies here were checked in October 2026. Both companies ship new image models frequently and both have changed their rules on real people more than once, so confirm anything load-bearing against the current documentation.
- OpenAI — Images API reference (supported sizes)
- OpenAI Help Center — C2PA in ChatGPT images
- OpenAI — Advancing content provenance
- Google — Gemini API image generation (aspect ratio and size)
- Google — Nano Banana Pro announcement (watermarking)
- TechCrunch — YouTube's dynamic thumbnails and Studio AI tools
- Digital Camera World — testing Gemini's limits on real people
- CBC — stress-testing ChatGPT's image rules on real people
The bottom line
The chat models are not bad at thumbnails. They are excellent at the half of the job that is making an image and indifferent to the half that is making a piece of packaging — the fixed shape, the fixed font, the fixed corner, the three variants, the fifty uploads that have to look related. That second half is unglamorous and it is where click-through rate actually lives.
So use them for what they are extraordinary at. Invent the scene, fix the frame, put yourself somewhere you have never been. Then take the file somewhere that knows a thumbnail is 1280×720, that remembers your font and your two colours, and that treats three variants as the normal output rather than an unusual request. That is the whole of what Thumblore does, and it is the reason the two tools are not really competitors.
If you want to get better at the generating half, the prompt structure that works is in AI prompts for YouTube thumbnails. If you want to get better at the deciding half, start with thumbnail composition and then run a real test with Test & Compare. The model you used will matter less than either.