Most AI ad imagery fails for a structural reason rather than an artistic one: a single prompt is asked to extract the brief, invent the idea, art-direct the style and render the file. It will do all four, and all four badly.
Key takeaways
- The named framework is One Job Per Prompt: extraction, concept, style and render are four separate prompts run in sequence, with the Operator ruling between them.
- The chain is extract, expand, transform, render. Each stage hands a structured artefact to the next, and no stage is allowed to improve on the stage before it.
- Style is a chosen transformer, and not a mood. Each style carries its own written specification covering composition, typography, colour and layout, so the difference between two styles is legible rather than felt.
- The render step is the last one and the most literal. Every on-image string is passed verbatim, because one hallucinated character in a locked string fails the ad.
- Well-branded creative earns more from the same attention: Dr Karen Nelson-Field's research with VCCP Media and Amplified found that even in ads with very limited view time, well-branded assets were 2.5 times more effective at driving outcomes than weak ones.1 The render is where codes are either carried or lost.
Why does one big prompt produce mediocre ads?
Because the four jobs inside it have different failure modes, and a single output hides which one failed. Extraction can be wrong about the offer, the concept can be weak, the style can fight the message, and the render can misspell a headline. When all four of those happen inside one pass, the result is simply "a bad ad", and there is nothing available to correct except the whole thing.
Separating them makes each failure visible and cheap. A wrong extraction gets caught before any idea is built on it, and a weak concept dies before art direction is spent on it. A style mismatch becomes a swap instead of a regeneration. And a render error stays a render error, not evidence that the idea was poor.
This is the same argument the guide makes about automating inside a phase rather than across phases, applied to a single production step. I wrote the general case in Stop Automating Across Phases, and the specific version, that AI ads look average because marketers write prompts where they should write specifications, in Prompt vs Spec. This chapter is the production chain those arguments imply.
What are the four jobs?
Extract, expand, transform and render, in that order, each with one responsibility and one output format. The chain matters as much as the parts, because every stage receives a structured artefact and is forbidden from improving on it.
- Extract. Read the source material and pull out what an ad needs: product, audience, the problem or desire, the promise, the offer, the call to action, any visual cues, any tone signals. Missing fields are inferred and marked as assumed rather than silently filled. No design advice and no copywriting: extraction only.
- Expand. Turn the extracted brief into a small number of concept statements: a short title and one or two sentences on the idea behind the ad. No headlines, no visual direction. This stage produces ideas, and an idea that arrives already art-directed has skipped a ruling.
- Transform. Take a chosen concept into a chosen style, using that style's written specification. This is where composition, typography, colour and layout are decided, and where the on-image text is finally written.
- Render. Turn the finished concept into the file, at the required dimensions, with every on-image string reproduced exactly as written.
The Operator rules between each stage. That is Findings Are Not Decisions applied to a production line rather than to a strategy document: the chain proposes at every step and advances only on a ruling.
How is style specified?
As a written transformer per style, and not as an adjective. "Make it bold" and "make it feel raw" produce the same median output from any model. A transformer states the composition rules, the typography rules, the colour rules, the layout rules and the required output fields, so two styles differ structurally and not atmospherically.
| Style family | What the specification pins down | The job it does |
|---|---|---|
| Bold typography | Type as the hero, one focal block, purposeful alignment, a small palette with an accent, contrast at accessibility minimums | Instant comprehension of a single statement |
| Hero | Product or person front and centre, a declared asset type, one realism treatment chosen per concept, rotating framing and light across a set | Authority, proof and product presence |
| Native / lifestyle | Real placement, warm muted grade, documentary framing, caption-style overlay | Belonging in the feed rather than interrupting it |
| Meme | Fixed font and stroke treatment, a named layout, all text on-image, medium rotated across a batch | Disarming a problem-aware reader, and carrying a segmented message lightly |
| Ugly | Deliberate misalignment, clashing type, flat saturated fills, legibility preserved | Breaking the pattern that polish itself has become |
Two rules travel with every transformer. Style is chosen from the fit matrix and never picked by taste, because a style that trips a persona's ejection triggers is unavailable in that lane whatever its general strength. And the required output fields are fixed per style, so a concept that arrives without its call to action or its layout note is incomplete rather than minimal.
In practice
Declare the asset situation before generating: product only, person only, person and product, or placeholder. Ambiguity here is where generated imagery invents a product that does not exist, or renders a person the client has never seen. A one-line declaration at the top of the concept removes an entire category of rework.
Watch for
A batch where every concept has quietly converged on the same framing, the same colour temperature, the same proof device. Variety is a specified requirement rather than an emergent property. Rotate framing, light and colour deliberately across a set, and treat two consecutive concepts sharing a treatment as a fault in the brief.
What does the render step have to guarantee?
Text fidelity, dimensions, and the brand codes. The render is the most literal stage in the chain and the one with the least licence. Every on-image string is passed through verbatim, and a single wrong character in a locked string fails the ad rather than requiring a small fix. That is the Text Fidelity rule, and it exists because a model that is 99% accurate on a headline is 100% unusable.
The codes are the other half. Nelson-Field's research is blunt about what weak branding costs: even in ads with very limited view time, well-branded assets were 2.5 times more effective at driving outcomes than weak ones, across more than 20,000 views of 72 digital video ads.1 Professor Jenni Romaniuk's rule for protecting those assets applies directly to a generation pipeline, where drift is one careless prompt away: "Consistency is crucial," and when a change is proposed, "switch your default answer to 'no'" and demand strong evidence first.2
Which is why the deterministic parts of the render are removed from the model's job entirely. What a generator provably cannot do reliably, which is reproduce an exact wordmark, is not asked of it. The generation leaves the sign-off zone clear and the real asset is placed afterwards. That principle has a name in the glossary, Deterministic Steps Stay Deterministic, and generated imagery is where it earns its keep.
Frequently asked questions
Can't a good model handle all four jobs in one prompt?
It will produce something. The problem is diagnostic and not capability: a single output gives you no way to tell whether the brief, the idea, the style or the render failed. Separation is not distrust of the model, it is what makes the pipeline correctable.
Where does the concept actually come from?
From the gated creative pipeline, not from the image chain. Concepts, angles and hooks are ruled upstream through the five gates, and the expand stage here works inside a concept the Operator already locked. An image chain that invents its own concepts has bypassed every creative gate the method has.
Why is the style chosen before the copy is written?
Because on-image text is a design element, and the amount of it a style can carry differs enormously. A bold-typography ad supports a short commanding line, an ugly ad tolerates a dominant headline and a subline, and a meme carries a setup and a punchline. Writing the copy first and choosing the style afterwards produces text that has to be cut to fit, which is where a locked hook gets quietly edited.
How many concepts should a transformer produce?
Enough to give the Operator a genuine field and few enough to judge in one sitting. What matters more is that they are structurally distinct and not numerous, because a set where framing, tone and colour repeat is one concept produced several times, and the count flatters the work without informing it.
Does this apply to imagery that isn't advertising?
The chain does. The transformers do not. Extraction, concept, style and render separate cleanly for any generated visual. The style specifications here are written for direct-response advertising, where legibility at small sizes and instant comprehension outrank aesthetics.
The bottom line
Generated imagery fails at the seams and not in the pixels. Give each prompt one job, hand a structured artefact down the chain, choose style from a written specification rather than a feeling, rule between the stages, and treat the render as the literal step it is: verbatim text, declared assets, codes carried, and the wordmark placed by hand because that is the thing the machine cannot be trusted with. The model is not the variable. The chain is.
Where this connects
Generated imagery executes concepts locked by The Five Gates, carries the codes defined in The Brand Codes Register, obeys the module set out in Static Ads, and passes the checkpoints in The Review Seams. Its hooks answer to The Hook Law and its claims to The Proof Bank. Rulings between stages are governed by The Operator's Laws. The archive arguments behind it are Prompt vs Spec, AI Doesn't Have a Quality Problem, You Have a Brief Problem and Your AI Marketing Has a CGI Problem. Back to the One Brain Method hub.
Part 5 · Creative Production · Chapter 22 of the One Brain Guide
Previous: Video ·
Next: Synthetic Assets ·
All chapters
Sources
- VCCP Media, Dr Karen Nelson-Field and Amplified, 16 May 2025: research across more than 20,000 views of 72 digital video ads found that even in ads with very limited view time, well-branded assets were "2.5x more effective at driving outcomes than weak ones." Amplified. VCCP Media is an agency and Amplified sells attention measurement, so both have a commercial interest in the finding.
- Professor Jenni Romaniuk, Ehrenberg-Bass Institute, The Four Commandments: future proofing a brand's identity: "Consistency is crucial", and "switch your default answer to 'no'".
Every statistic and quotation on this page has been checked against its primary source. Last verified 24 August 2026.
Free guide
Take the method with you. The complete One Brain Method as one PDF: every chapter, the diagrams, and every named framework, ready to hand to whoever runs your marketing. Enter your email and it's yours.
Want generated ads that carry your brand exactly?
Book a 30-minute call → See what it takes to have this installed
By Bruce Marjoribanks, 27 years in marketing, including building, running and selling his own agency. Founder of Untapped Profits and author of the One Brain Method.
Published 24 August 2026 · Last updated 25 August 2026
