v1 gave you twelve well-written prompts. This is the version that makes them look like they came out of the same camera on the same afternoon — a locked prompt grammar, a fixed reference plate set, one continuity sheet, a quantified motion law, and a finishing pass that unifies whatever the model hands back.
Every difference below exists for one reason: a generative video model re-rolls its entire look on every wording change. Consistency is not a prompt-quality problem — it is a variance problem, and you beat variance by removing the freedom to vary.
If you only adopt one thing from this document, adopt the finishing pass. Twelve mediocre generations put through one shared grade look more like a production than twelve excellent generations that were never touched.
A v2 prompt has exactly two parts. The head is the only place anything changes: shot size, subject, action, and the declared camera move. The tail is five lines of locked text — three of them identical in every asset, two chosen from a closed list. You never rewrite the tail. You never paraphrase it. Paraphrasing the tail is what broke v1.
Text hallucination is the number-one reject cause on office-interior prompts — the model wants to put a sign on a wall and a spreadsheet on a screen. That is why text, letters, numbers lead the negative string rather than sitting at the end.
Every one of the twelve clips takes place in the same building, on the same afternoon, with the same props. Where v1 described a generic "modern office" twelve times, v2 describes this office once and then refers back to it.
A pale oak table against an off-white plaster wall, one muted sage-teal upholstered chair, floor-to-ceiling window camera-left, open-plan office falling away defocused behind. Used for: M0-COVER · M1-OPEN · M1-BODY · M1-MIND · M2-HERO · M2-MATERIAL · M3-OPEN
The same window wall, one room along: a longer pale oak table, four muted sage-teal chairs, a glass partition softly defocused behind. Used for: M0-WELCOME · M1-SCENARIO · M3-LISTEN
One brushed bronze singing bowl, roughly 150mm. One felt mallet with a wooden shaft. One frosted quartz bowl — Slide 25 only. A closed grey linen notebook. A matte black pen. An unbranded stoneware cup. Plain cream paper, unprinted. A shallow bowl of still water on dark slate.
Charcoal, oatmeal, slate blue, warm grey. Soft matte fabrics. No pure white shirts — they clip against a window key and pull the whole grade brighter. No fine patterns — pinstripe and check alias badly at 720p and the model renders the shimmer as motion. No reflective jewellery, no lanyards, no visible branding.
Five still images, made once, that every clip is generated from. This is the highest-leverage half-hour of the whole job: reference-to-video anchors colour, surface, room geometry and wardrobe far more reliably than any amount of adjective tuning, and it works even where the endpoint gives you no seed control.
Make them by whichever route is cheaper for you — photograph them if you have the bowl and a window, or generate them with the same tail lines and lock the ones you like. Either way, freeze them before you generate a single clip and never regenerate a plate mid-shoot.
The brushed bronze bowl and felt mallet on pale oak, lit camera-left, plaster wall behind. The single most reused plate — it carries the bronze that every Module 2 clip has to match.
Wide, no people, no props. Establishes the room geometry, the window falloff, and the wall tone that becomes your negative space.
The meeting room, same treatment. Keeps the three two-person clips in a room that visibly belongs to the same building.
One mid-shot of a professional in the capsule palette, face turned three-quarters away and softly out of focus. Referencing this rather than describing people in words is what stops the cast changing between Module 1 and Module 3.
A macro of the oak grain meeting the plaster wall. Feeds the ambient clip and pins the texture the grade sits on.
If your endpoint exposes a seed, record the one that produced your approved look-dev probe and reuse it. If it does not, the plate set does the same job more robustly — reference conditioning survives model-side updates in a way seeds often do not.
The reason v1 output varies is that it invites you to generate asset one, then asset two, then asset three. A studio never does that. It develops the look, locks it, and only then runs the sequence.
Generate M2-HERO-01 three times and nothing else. It has the most surface, the most specular detail and the least forgiving material, so it fails first and it fails visibly. Pick the one that reads most like a product film and least like a wellness advert. That take is now the reference for the whole deck. Budget: about 45 minutes and $5.
Pull a clean frame from the approved probe as P1 if you generated it. Assemble P2–P5. Put all five in a folder called /plates and do not touch it again.
Generate all twelve in a single session, in asset order, without editing the tail between them. Three takes per asset. Do not stop to review; reviewing mid-run tempts you into rewording, and rewording is the variance.
Run the nine-point reject list against all thirty-six takes on a projector or a large screen, not a laptop. Select one per asset. Anything that fails goes back for takes four to six with the head reworded — never the tail.
Deflicker, upscale, one shared grade, matched grain, mute, loop. This is what turns twelve independent generations into a set.
The whole shoot on one page. Coherence should be visible here before you generate anything — if a row's optics, move or text zone looks like an outlier, that is a decision to justify, not a variation to enjoy.
| Asset | Slides | Set | Plate | Optic | Move | Frame | Text zone | Sec |
|---|
Move vocabulary — nothing outside these four.
Each prompt below is complete and paste-ready. What you see in ink is the head — the only part that differs between assets. What you see in bronze is the tail, reproduced so you can confirm it is identical everywhere. Copy takes the whole thing.
Audio policy is one line and applies to all twelve: request silent, and mute on import regardless. Native audio will not match across twelve generations and will fight both the trainer and the live bowl.
Where video must not go — unchanged from v1, and worth restating: assessment slides 17, 31 and 43, the learning-outcome checkpoints, and the recap diagrams all stay completely static. Motion behind a checkpoint reads as decoration and spends the credibility the rest of the deck is building.
Review on the projector you will present on, at full screen, with the room lit as it will be on the day. A laptop screen hides exactly the failures that a projector amplifies: flicker, banding, and texture boil.
A take that fails gets a reworded head and another three attempts. It never gets an edited tail. If three assets in a row fail the same way, the fault is in the plate set, not the prompts — remake the plate and re-run those assets together.
Twelve approved takes are still twelve independent generations: slightly different exposure, slightly different noise, slightly different sharpness. One shared pass over all of them collapses that difference — the same thing a colourist does to a real shoot, and the reason real shoots look like shoots.
Run it on every clip with identical settings. Resist the urge to tweak per-clip; per-clip correction is how you re-introduce the variance you just spent a day removing.
What each stage buys you: deflicker kills the frame-to-frame luminance wobble generative video always carries. lanczos to 1080p is sharper than PowerPoint's own scaler and stops the projector doing it badly. eq + curves is the shared grade — a 6% contrast pull, 12% desaturation and a black lift to 4.5% gives the whole set one filmic base. noise lays a single matched grain over everything, which is the trick that makes twelve different renders feel like one stock. unsharp at 0.35 restores bite lost to the upscale without going crispy.
Ping-pong is seamless by construction for DRIFT and HOLD, and acceptable for PUSH — a slow push that slowly pulls back reads as breathing. Never ping-pong a RACK: focus visibly un-racking looks like a mistake.
M2-HERO-01_t02_final.mp4
asset id · take number · stage
Keep the rejected takes. When a slide changes six weeks from now you will want take 01 rather than a fresh generation off a model that has since been updated.
In PowerPoint. Insert as background video, send to back, Loop until Stopped and Start Automatically, mute the track. Then lay a solid slate-charcoal shape at 55–70% opacity over the declared text third. That overlay is not a fallback — it is part of the design, and it guarantees text contrast even when a generation comes back a stop brighter than its neighbours.
Costed at the $0.13/second figure from v1. The v2 structure does not reduce spend much — it reduces the number of days, because you are no longer regenerating asset four to make it match asset nine.
For generating additional assets later that still belong to this shoot. Paste the block, then add one line: Generate for: [slide number + what you want to see].