For years, every AI image tool worked the same way under the hood. You typed a prompt into a chatbot, the chatbot rewrote your words into a caption, and a completely separate image engine painted something from that caption. The painter never saw what you actually wrote. It saw a summary.
That era is ending, and it ended quietly. Over the past year and a half, both OpenAI and Google moved image generation inside the language model itself. The same model that reads your brief now draws the picture, in the same pass, with your full conversation, your reference images and its general knowledge all still in play.
We make images with these tools every week, for mood boards, pitch frames and fully AI-generated spec films, so this is not an academic distinction for us. It changes how you write a brief, and it finally puts legible Arabic text within reach, which matters more in this region than any benchmark. Here is the plain-English version.
One caveat before the names start. Everything specific in this article, the model names, the version numbers and the dates, is current as of the writing of this article in August 2026. This space moves fast enough that some of those names will be superseded within months. The part worth keeping is the framework: the three architectures below change much more slowly than the products built on them, and once you can recognize which shape a tool is, you will know how to brief it no matter what it is called this quarter.
The three shapes every model fits
Every image model you can name today has one of three architectures, and the architecture decides how you should talk to it.
Shape A: the relay. Your prompt goes to a language model, which rewrites it into a caption and hands that caption to a separate diffusion model. DALL·E 3 worked this way, and so do Imagen 4, Midjourney, Flux and Stable Diffusion. The weakness is the handoff: nuance, spatial logic and conditional instructions get flattened into whatever survives the caption. This is why "keyword soup" prompting existed at all. The caption was the only thing the painter ever received, so people learned to write like a caption.
Shape B: native generation. One multimodal model reads your words and emits the image directly, the same way it emits text. No handoff, no caption, no second engine. OpenAI crossed this line in March 2025 when ChatGPT's image generation moved into the GPT model itself, and Google followed in August 2025 with the Gemini image model the internet quickly nicknamed Nano Banana. Because nothing is lost in translation, these models can edit conversationally ("warm up the sky, keep everything else") and hold a character consistent across images, which shape A could never do reliably.
Shape C: native generation plus reasoning. The newest step, and the one that changes professional work. The model plans before it renders: it thinks through the composition, can consult search for factual grounding, and in some cases drafts and critiques interim versions before committing to the final frame. Google's Nano Banana Pro (built on Gemini 3 Pro, released November 2025) and OpenAI's GPT Image 2 (released April 2026) both work this way. This is why the newest models can lay out an infographic or render a poster's worth of text where earlier models produced confident gibberish.
Who is who right now
The naming in this space is genuinely confusing, so one paragraph of orientation. At OpenAI, DALL·E 3 is legacy, GPT Image 1 and 1.5 were the first native generation, and GPT Image 2 is the current model with reasoning. At Google, the Imagen line is being retired: per Google's own developer documentation, the Imagen 4 API endpoints shut down on 17 August 2026, and Google points users to the Gemini image family instead. That family is what the nicknames refer to: Nano Banana (the 2025 original), Nano Banana 2 (the current everyday workhorse) and Nano Banana Pro (the reasoning flagship).
Two practical warnings from our own bruises. First, the nicknames are marketing labels, not model IDs: "Nano Banana" covers several distinct models across Gemini generations, and the name in the app, the name in the developer console and the API endpoint are three different strings. If a result looks unexpectedly bad, confirm which model you actually used before judging the tool. Second, if any workflow you rely on still calls an Imagen endpoint, it stops working in August 2026, so migrate it now.
The Arabic question
For anyone producing client-facing graphics in this region, text rendering has always been the dealbreaker, and Arabic was reliably the first thing to break. Both flagships now claim serious progress: OpenAI's launch material for GPT Image 2 claims roughly 99 percent character-level text accuracy across scripts including Arabic, and Google grounds Nano Banana Pro's text rendering with its reasoning pass.
Our advice is the same one we apply ourselves: test your own strings before you trust anyone's benchmark. Vendor numbers are measured on their test sets, not on your client's tagline set small over a busy background. Run the actual line, in the actual weight and size you need, and look closely at the letter connections. The improvement is real, but "much better" and "print-ready every time" are different claims, and only your own test tells you which one you are getting today.
How this changes the way you brief
The architecture shift is not trivia. It rewrites the craft of prompting, in three ways.
- Write briefs, not keyword strings. Comma-separated tag lists were a workaround for shape A, where only a caption survived. Shapes B and C read sentences, conditions and intent. Describe the shot the way you would describe it to a crew: what is happening, where things sit in frame, what matters and what must not change. The model actually processes it.
- Keep the camera language. This is the cinematographer's half of the deal. A reasoning model will happily solve your composition its own way if you leave the decision open. Explicit lens, framing and lighting language ("35mm, waist-up, single warm practical source from the left") is how you keep authorship instead of receiving the model's opinion. The vocabulary you already use on set is exactly the vocabulary these models respond to.
- Spend reasoning where it pays. Thinking passes are slower and cost more. For volume work, quick variants and exploration, the fast non-reasoning models are the sensible default. Save the reasoning models for the frames where they earn it: complex layouts, factual infographics, and anything where text must be legible.
Where we stand
We have already produced work end to end on these tools. Never Hide, our fully AI-generated Ray-Ban spec film, was built with imagery from Nano Banana Pro and motion from a video model, then cut and finished in-house like any other production. Spec means our own exercise, not a client commission, and we label it that way on purpose. You can judge the results yourself on our AI samples page, and our broader AI stance, including the option of a fully AI-free process, is spelled out in AI in Video Production: What We Use It For, and What We Don't.
The tools moved fast this year, and everything named above is a snapshot of August 2026, accurate as of the writing of this article and guaranteed to age. The way to stay oriented is not to memorize model names, which may be stale by winter. It is to know the three shapes, ask which one you are holding, and brief it accordingly. That part will still be true when the names have changed.

