A good brief tells people what the work must achieve. An AI production system must also define what stays fixed, what may vary, how outputs are judged, and how the chosen direction survives the next generation.
An AI image prompt is not a creative brief. It is one instruction inside a production system.
A traditional brief can rely on a designer to interpret brand history, visual hierarchy, production feasibility, cultural context, product accuracy, and what “premium” should not mean. An image model does not carry that responsibility. It predicts an output from supplied instructions and references.
If the direction remains implicit, the output fills the gap with familiar patterns. The result looks competent and interchangeable.
Controlled does not mean deterministic. Current image models still vary, drift and misunderstand. The system makes those failures visible and reviewable.
Layer 1: the creative objective
Begin with the human brief: business objective, audience, single message, desired response, placement, mandatory elements, success criteria. This is the same document described in how to brief a designer, and it is still the starting point.
It states what the work must achieve. It does not yet tell the generation system how to preserve it.
Layer 2: invariants
What must stay consistent across the asset family.
Identity: approved palette, logo treatment, typography applied outside the model, product construction, recurring styling codes.
Subject: casting direction, age range where relevant and lawful, wardrobe rules, grooming, attitude, required human presence.
Camera: lens family, camera-height range, acceptable distortion, focus behaviour, crop logic.
Light: direction, hardness, temperature, contrast, shadow behaviour.
Surface: grain, skin treatment, material texture, saturation, background character.
Invariants must be observable. “Editorial” is not an invariant until the team defines which camera, light, styling and print characteristics make it editorial.
Layer 3: permitted variables
Controlled variation prevents repetition. Define what may change: pose, gesture, environment, prop, subject count, negative-space position, dominant scale, motion, camera rotation, format.
Then define combinations that must not occur. A low camera may use a standing or seated pose. A bird’s-eye camera uses floor posing. Extreme fisheye is reserved for one hero family. Red wardrobe and red environment may not appear together.
The system should create range without abandoning the language.
Layer 4: reference roles
Do not feed every reference into one undifferentiated moodboard. Assign each image a role.
Identity reference: what must remain recognisable. Visual-language reference: colour, light, texture, attitude, finish. Composition reference: pose, crop, camera height, negative space, balance. Exclusion reference: what must not appear, such as synthetic skin, generic streetwear, excessive props, distorted product or incorrect brand colour.
The model does not understand reference roles unless the instruction makes them explicit.
Layer 5: prompt schema
Use one schema for every generation, with named sections: purpose, subject, action, environment, camera, light, visual language, invariants, permitted variables, exclusions, format.
This does not need to become one enormous paragraph. Clear sections make review and revision easier, and they make it obvious which section failed when an output is wrong.
Layer 6: generate families, not accidents
Work in small families. Low-angle standing portraits. Bird’s-eye floor compositions. Close beauty details. Seated attitude portraits. Tactile product interaction.
Generate a contact sheet for one family before starting the next. This reveals repetition, subject drift, inconsistent lens behaviour, palette expansion, product mutation and unusable crops.
An isolated image can look strong while weakening the campaign.
Layer 7: score with a fixed rubric
| Criterion | Points |
|---|---|
| Message and attitude | 20 |
| Brand-language fit | 20 |
| Subject or product accuracy | 15 |
| Composition and placement | 15 |
| Camera and light consistency | 10 |
| Material and skin credibility | 10 |
| Rights and production feasibility | 10 |
Set a minimum total, automatic rejection criteria, a named reviewer and a tie-breaking rule. Automatic rejection might include a mutated product, extra limbs or objects, a false logo, an unusable face, an unsupported claim, or a composition that cannot fit the required format.
Layer 8: promote a golden frame
Once a direction is approved, select a golden frame and record the exact output, model and version, references, prompt, approved corrections, rubric score, and reason for selection.
Current platforms support image inputs and iterative editing. OpenAI and Google both recommend reusing approved images or pose references when consistency matters. Those capabilities help. They do not guarantee preservation after a model update or across complex edits.
The golden frame is a target, not a lock.
Layer 9: finish outside the model
Move deterministic elements into controlled production: typography, logo, legal copy, prices, interface details, product specifications, final colour, retouching, export.
Image generation is strongest where interpretation is useful. It is weaker where exact reproduction is mandatory. Which is also why the type scale and contrast rules belong in a design system, not in a prompt.
Layer 10: preserve provenance
Maintain a production record: asset ID, brief version, model and version, prompt, references and permissions, generation date, output ID, edits, reviewer, approval, rejection reason.
C2PA provides a technical specification for signed provenance assertions and editing history. It helps preserve declared origin and changes where supported. It does not prove that the depicted event occurred or that every rights issue has been cleared.
Provenance is evidence about process, not a certificate of truth.
Do not confuse consistency with sameness
A controlled system should preserve identity, attitude, camera family, light logic and material treatment. It should still allow surprise, rhythm, different human behaviour, varied composition and changes in scale.
If every image repeats the same pose and background, the system is stable and creatively dead.
How we run it
The human brief belongs to the brand and the business. Origin translates that brief into a production language. The generation system makes the language repeatable. We approve the result, and Scale runs it at the volume real testing needs.
The deliverables are a creative objective, an invariant sheet, an asset-family matrix, a reference-role board, a prompt schema, a scoring rubric, golden frames and a production log.
One good prompt can produce one good image. A controlled system explains why the next forty-nine still belong together.
If you are generating assets and they are drifting apart, show us the set and we will tell you which invariant is missing.
Sources
OpenAI, GPT Image model documentation.
Google, Gemini image-generation guidance.
C2PA Specification 2.4.
NIST Generative AI Profile.
US Copyright Office AI initiative.
Current to 30 July 2026. This article provides a production framework, not legal advice.