Revised build · consolidated takes

Five continuous takes, cut to 443 frames

One generation per camera setup instead of one per block, then cut internally to the frame spec. Fewer calls, no inconsistency between generations of the same room, and the jump cuts do double duty — they reproduce the source's structure and hide the model's weakest frames.

060120180240300360420

Why this is better

One take per setup, not one per pose

Generating each block separately was wrong, and for a reason worth stating plainly: it manufactures inconsistency the source does not have. Three separate generations of the bathroom produce three slightly different bathrooms and three slightly different women, and no amount of prompt discipline fixes that — the model has no memory between calls.

Every setup that fits in one continuous shot should be one generation, cut internally.

That is also how the original was made. It is footage shot in a handful of locations and then cut down — the jump cuts are not transitions, they are the seams left behind when the boring parts were removed. Reproducing it as five continuous takes and then cutting them is structurally the same process, not an approximation of it.

There is a second benefit that matters more than consistency. The jump cuts let us discard the model's worst frames. Generate ten seconds of continuous performance, keep the four best sub-second windows, and the cuts hide everything thrown away. A mediocre transition costs nothing if it lands between two frames the audience never sees.

The trade is real and worth naming: consolidation swaps *between-take* inconsistency for *within-take* drift. Over ten seconds a model will widen its framing and quietly flatter the subject. But within-take drift happens inside one room instance and one person instance, so it reads as the same shoot — and the source itself changes framing across its internal cuts. Between five separate generations it would read as five different shoots.

What we hand the model

A storyboard, not three isolated references

The second change is what we hand the model.

Passing a character sheet, a room plate and a wardrobe plate as three separate references asks the model to composite — to take a person from one image, a room from another, and clothes from a third, and integrate them. It often works, but it is three chances to drift and it tells the model nothing about what happens next.

A storyboard tells it the whole arc at once. One image, the character already in the room already wearing the clothes, with the poses of that take laid out as panels in order. Identity, setting and wardrobe arrive pre-integrated, and the movement is shown rather than described.

The constraint is which model can receive it. Only `seedance_2_0` accepts `image_references` alongside a start frame — up to nine images, twelve reference files total. `minimax_h3` accepts references but refuses to combine them with a start frame. `kling3_0`, `wan2_7` and `veo3_1` have no reference channel at all: a start frame is the only image they take.

So there are two routes, and they are worth testing against each other rather than picking blind.

Route 1

Storyboard-derived start frame

Recommended

kling3_0 · 2.0 cr/s (std) · 2.5 (pro)

Generate the storyboard, then crop panel one and use it as the start frame. The storyboard never reaches the video model — it is a planning artifact that guarantees the poses are coherent in that specific room, and it gives you a start frame with character, set and wardrobe already integrated.

For Cheapest by a wide margin. 10s at std is 20 credits. Kling has been the reliable engine on this account.

Against The model sees one frame and a paragraph. It is told what happens next rather than shown it.

Route 2

Storyboard as a live reference

seedance_2_0 · 9.0 cr/s at 1080p

Pass the storyboard as an image_reference alongside the start frame, so the model can see every pose in the take while generating it.

For The only route where the model actually sees the intended movement. This is the approach that has worked before on complex character motion.

Against Four and a half times the price. A 10-second take is 90 credits against Kling's 20 — the full film would run past 400 credits on video alone.

Do not pick between them on argument. Setup D is the hardest take in the film — four distinct poses, a locked camera and a body that rotates — so generate it both ways and look. Kling at 10s std is 20 credits; Seedance at 10s is 90. One hundred and ten credits settles whether a live storyboard reference is worth 4.5x, and the answer transfers to every other setup. If Kling holds, the whole film runs on route 1 for 66 credits of video.

Matching the look

Measured optics, not copied rooms

The rooms in this plan are described in prose, which means the model invents a plausible bathroom rather than one that looks like *that* footage. Closing that gap does not require the source's actual rooms — and using them would defeat the measurement, since half of every frame would arrive pre-made rather than reconstructed.

What actually makes those backgrounds read the way they do is not the specific room. It is a set of optical properties that can be measured off the source and specified as numbers: how warm the light is, where it falls off, what is blown out, how deep the shadows sit, how much grain the sensor puts down. Measured across every setup:

SetRoomMean LR:BR:GClipped ShadowFalloffGrainWhat that means
ABathroom1351.401.130.24%12.41.22 · centre brighter5.9Warm and bright with a centre-weighted falloff — the overhead fitting lights the middle of the room and the corners drop away. Almost nothing clipped. The grainiest setup in the film.
BLiving room1361.411.132.63%15.31.00 · flat5.5Same warmth as the bathroom but with a genuinely blown highlight — 2.6% of the frame is clipped, which is the lamp burning out against the wall. Flat corner-to-corner otherwise. Grainy.
CBedroom1351.361.180.02%9.80.84 · edges brighter4.0The cleanest and least warm of the interiors, and the only one where the edges are brighter than the centre — a window off frame-left throwing light in from the side. Deepest shadows in the film, nothing clipped at all.
DGym corridor1181.471.291.66%14.61.13 · centre brighter4.0The warmest setup by a clear margin, and darker overall — an ageing fluorescent tube reads amber, not white. The tube itself clips. Clean, low grain.
ECar interior1081.131.070.00%18.70.65 · edges much brighter5.7The only neutral-daylight setup — R:B of 1.13 against 1.40 indoors. Darkest mean and the most lifted shadows, which is what overcast light through glass does. Edges far brighter than centre because the side window dominates. Grainy again.

The strongest move here is post, not prompting.

Do not fight a model for a colour temperature. Ask it for a plausible room with the right light *direction* and the right *sources*, then grade the output to the measured numbers — white balance, mean luminance, shadow floor, clip point, vignette, and matched grain. The grade is deterministic, free, exactly repeatable, and hits the target every time; a prompt does none of those things.

That splits the job cleanly. The model owns geometry and content — room type, where the light comes from, what furniture is where, how far she stands from the wall. Post owns the look — every number in the table above. Trying to make the model own both is what produces generic rooms that still do not match.

The grade, step by step

StepTarget
White balanceCorrect each take to its target R:B and R:G. Setup E is the odd one out at 1.13 — grade it neutral and leave the other four warm, or the car stops reading as outdoors.
LuminanceLift or pull to the target mean. A and B and C all sit at ~135; D at 118; E at 108. Getting these wrong is what makes cuts between setups feel like different films.
Shadow floorSet the 5th-percentile luminance to target. C sits deepest at 9.8, E most lifted at 18.7 — that difference is most of what separates "indoor with a window" from "overcast through glass".
Highlight clipB and D need a deliberately blown source — 2.6% and 1.7% of frame. If the model returns a clean image, clip it in post rather than re-rolling.
VignetteA and D want a centre-bright falloff; C and E want the opposite, with the edge nearest the window lifted. This is a mask, not a lens effect.
GrainMatch the measured noise floor — A, B and E are visibly grainy at ~5.5–5.9, C and D clean at ~4.0. Adding grain last is what kills the plastic look models leave behind.

Two more things that read as "the same room" and cost nothing to specify, because both are measurable off the source and neither is content:

- Camera height and subject distance. Every setup here is a phone at chest height except the car, which is in her lap angled up. Distance shows in how much room sits behind her — describe it as a crop ("mid-thigh up", "head and shoulders") rather than hoping. - Depth of field. A phone lens at these distances holds everything sharp. If the model returns a blurred background, that alone will make it read as produced rather than filmed — say "everything in focus, no background blur" explicitly.

What this deliberately does not do is rebuild the source's actual rooms. Not primarily a rights question — it is that the reconstruction stops measuring anything. If the bathroom arrives from the source's own frames, the model is being scored on a person against a pre-made backdrop, and the answer to "how far can these models go" covers half the frame. The plan's rule stands: no job takes a frame of the source video as a reference. What is copied is structure and measured optics; what is generated is everything you can see.

Foundation

Two character sheets

Still the only assets whose failure is unrecoverable. Generate the after first and derive the before from it. Gate: stop after S1 and look — if the before is not convincingly heavy, tired and badly lit, nothing downstream rescues it.

S0 S0-after-sheet.png2 cr
References in
  • none — root asset
Used by

Every after keyframe — KC*, KD*, KE*

Prompt — verbatim
Four-view character reference sheet of one woman on a plain light-grey seamless studio background, evenly lit, no text, no labels, no borders. Left to right: front full-body standing, three-quarter turn, profile, and a tight head-and-shoulders portrait. The same woman in all four. 61 years old. Narrow face, strong jaw, deep-set hazel eyes, thin straight nose. A small crescent scar through the outer end of the left eyebrow. Short dark-auburn hair with clear grey at the temples, worn loose with a slight wave, just covering the ears. Lean and wiry rather than bulky — defined deltoids, visible forearm tendons, a flat stomach without a six-pack. Heavily freckled chest, shoulders and forearms; sun damage on the backs of the hands; loose skin at the elbows and knees; laugh lines and a crepey neck. She reads as a real 61-year-old who trains, not as a fitness model. Soft even studio light. Naturalistic and undoctored — no beauty retouching, no skin smoothing, no glamour lighting, no makeup.
S1 S1-before-sheet.png2 cr
References in
  • S0-after-sheet.png
Used by

KA1–KA3, KB1, M1–M6

Prompt — verbatim
The reference is a character sheet of a woman. Generate a NEW four-view sheet of the EXACT SAME WOMAN — same layout, same plain light-grey seamless background, no text, no labels — but as she looked five years ago, before she got in shape. The same woman five years earlier and about thirty pounds heavier. Identical bone structure, identical eye shape and colour, identical nose, identical mouth width, identical ear shape, and the same crescent scar through the left eyebrow. The weight sits around the middle, the upper arms and the upper back: soft undefined arms with loose triceps, a thicker waist, a rounded lower belly, a softer jawline with a full under-chin. Hair is uniformly steel-grey, flat and dull, cut blunt just above the shoulders, unstyled. Skin is pale, dry and unlit, no tan, dark circles under the eyes, puffy eyelids. Posture is slumped — shoulders rolled forward, chin slightly down, weight on one hip. Wearing a plain grey cotton vest and loose black trousers, bare feet. Photographed on a phone under flat overhead light. Completely undoctored and unflattering — no retouching, no smoothing, no makeup, no glamour lighting. She must read as an ordinary unposed 61-year-old woman, not a model.

The build

Five takes

Each take runs one prompt describing the whole arc in order, with explicit timings so the model paces it. The start frame is panel one cropped out of the storyboard, so character, room and wardrobe arrive already integrated rather than composited from three references.

Take A

Bathroom

frames 0–78 · 114–149 · 175–178 · strobe · 116f delivered · 22 cr

Serves blocks 1, 2, 4, 11 + every before slice of the strobe

The source material

Source frame 0
f0
Source frame 10
f10
Source frame 20
f20
Source frame 29
f29
Source frame 30
f30
Source frame 42
f42
Source frame 54
f54
Source frame 66
f66
Source frame 77
f77
Source frame 114
f114
Source frame 126
f126
Source frame 138
f138
Source frame 148
f148

Cut list out of this take

Block 130ffrom ≈1.0s — the folded hold
Block 248ffrom ≈4.2s — already gripping, into the start of the turn
Block 435ffrom ≈8.0s — the settled profile
Block 113ffrom ≈8.5s — the same profile
Strobe5 slicesfrom ≈8.2s — 2–4 frames each

The turn at 6.5s is the only movement asked for anywhere in this take, and it happens between two windows we keep — so if the model executes it clumsily, the jump cut at block 2 hides it.

Performance arc to generate

0.0–2.5Stands square to camera, arms folded across the stomach, flat stare down the lens. Barely moves.
2.5–4.0Unfolds her arms and brings her right hand across to grip her left upper arm.
4.0–6.5Squeezes and holds the arm, looking down at it.
6.5–10.0Turns slowly to her left until she is fully side-on, still gripping, eyes down, and stays there.

Steps

  1. SBABathroom storyboard image2 cr
    References in
    • S1-before-sheet.png
    Output

    sba.png

    Call
    higgsfield generate create nano_banana_pro \
      --aspect_ratio 16:9 --resolution 2k \
      --image S1-before-sheet.png \
      --prompt "<below>"
    Prompt — verbatim
    A four-panel contact sheet, four photographs laid side by side in a horizontal row with thin white gutters between them, no text, no numbers, no labels, no arrows. Each panel is a full photograph that fills its panel edge to edge — the room fills the frame behind her in every one. No plain or empty backdrop anywhere. Every panel shows the SAME woman from the reference in the SAME small ordinary bathroom, shot from the SAME fixed camera position — a phone resting on a shelf at chest height, dead level, framing her from mid-thigh up, her head near the top of the frame. The room is identical in all four: white subway tile to shoulder height, a plain shower curtain rail on the left, a small basin with chrome taps on the right, a single frosted ceiling light directly overhead throwing flat cold light with no fill. She wears a washed-out olive cotton vest top and loose black jersey trousers, bare feet, no makeup, no jewellery. Panel 1: facing the camera square-on, arms folded across her stomach, flat unimpressed expression, looking down the lens. Panel 2: arms unfolded, her right hand coming across her body toward her left upper arm. Panel 3: her right hand gripping and squeezing her left upper arm, looking down at it. Panel 4: turned fully side-on to the camera facing frame right, still gripping the arm, eyes down, weight on one hip. Unedited phone snapshot quality in every panel — flat light, faint sensor noise, no colour grade, no retouching, no smoothing, deliberately unflattering.
  2. VA Generate the take — 10s continuous video20 cr
    Start frame
    • sba-panel1.png cropped from the storyboard
    Output

    take-a.mp4

    Generated / used

    10s generated, 116f delivered across 5 windows

    Call
    higgsfield generate create kling3_0 \
      --aspect_ratio 9:16 --duration 10 --mode std --sound off \
      --start-image sba-panel1.png \
      --prompt "<below>"
    Prompt — verbatim
    Locked-off phone camera resting on a shelf — absolutely no camera movement of any kind, no pan, no drift, no push, no zoom, no handheld shake, for the entire ten seconds. She stands square to the camera with her arms folded across her stomach, breathing, holding the lens with a flat closed-mouth expression, and stays like that for about two and a half seconds. She then unfolds her arms and brings her right hand across her body to grip her left upper arm. She squeezes and holds it, looking down at her arm, for a few seconds. Finally she turns slowly to her left until she is fully side-on to the camera, still gripping her arm, eyes down, and stays there until the end. Her posture stays slumped and her shoulders stay rolled forward throughout. She never smiles. Flat cold overhead bathroom light. Unedited amateur phone video, no colour grade.
Take B

Sofa

frames 78–114 · 36f delivered · 2 cr

Serves blocks 3

The source material

Source frame 78
f78
Source frame 90
f90
Source frame 102
f102
Source frame 113
f113

Cut list out of this take

Block 336fa still with a 5% push

Measured motion in the source is near zero across all 36 frames. A still with a push is indistinguishable and saves a whole generation.

Performance arc to generate

Slumped in the sofa corner in near-total profile, one hand near her face, a small dog asleep against her thigh. Warm dim tungsten from a single lamp.

Steps

  1. SBBSofa still image2 cr
    References in
    • S1-before-sheet.png
    Output

    sbb.png

    Call
    higgsfield generate create nano_banana_pro \
      --aspect_ratio 9:16 --resolution 2k \
      --image S1-before-sheet.png \
      --prompt "<below>"
    Prompt — verbatim
    Amateur vertical phone photo of the woman in the reference, sitting slumped back into the corner of a brown corduroy sofa in a small dim living room. She sits in the left third of the frame in near-total profile facing right, one hand raised near her face mid-gesture. A small wiry brown terrier is curled up asleep against her thigh at the lower right. She wears an oversized oatmeal marl sweatshirt and dark jeans, no shoes, no makeup. Behind her: a low table with a TV remote on it, a beige wall, a dark doorway into a hall on the right. One standard lamp behind the sofa throws a warm blown-out patch up the wall and is the only light source — everything else is dim, warm and slightly yellow. Camera handheld at seated eye level a couple of metres away. Unedited phone snapshot, grain in the shadows, no colour grade, no retouching. It should feel like a photo someone took without her noticing.
Take C

Bedroom

frames 205–269 · strobe · 64f delivered · 18 cr

Serves blocks 22, 23, 24, 25 + every after slice of the strobe

The source material

Source frame 205
f205
Source frame 209
f209
Source frame 213
f213
Source frame 214
f214
Source frame 215
f215
Source frame 216
f216
Source frame 224
f224
Source frame 232
f232
Source frame 240
f240
Source frame 241
f241
Source frame 250
f250
Source frame 259
f259
Source frame 268
f268

Cut list out of this take

Strobe6 slicesfrom ≈0.3s — the tight held pose
Block 229ffrom ≈0.5s — same tight framing
Block 232ffrom ≈1.9s — the blurred hands
Block 2425ffrom ≈3.2s — hands out, bouncing
Block 2528ffrom ≈6.0s — arm behind the head

The framing widens naturally across a handheld take, which is what the source does too — the block 25 window is wider than the block 22 window in the original. Here that comes free instead of needing a separate keyframe.

Performance arc to generate

0.0–1.5Framed tight on head and shoulders, chin lifted, holding the lens with a warm closed-mouth expression. Still.
1.5–2.5Both hands swing up into the bottom of frame toward the lens.
2.5–5.0Hands out, gesturing, bouncing lightly on the spot, grinning.
5.0–8.0One arm sweeps up and behind her head, elbow high, torso leaning, swaying.

Steps

  1. SBCBedroom storyboard image2 cr
    References in
    • S0-after-sheet.png
    Output

    sbc.png

    Call
    higgsfield generate create nano_banana_pro \
      --aspect_ratio 16:9 --resolution 2k \
      --image S0-after-sheet.png \
      --prompt "<below>"
    Prompt — verbatim
    A four-panel contact sheet, four photographs laid side by side in a horizontal row with thin white gutters between them, no text, no numbers, no labels, no arrows. Each panel is a full photograph that fills its panel edge to edge — the room fills the frame behind her in every one. No plain or empty backdrop anywhere. Every panel shows the SAME woman from the reference in the SAME bright bedroom — white panelled wardrobe doors slightly out of focus behind her, a pale wall, bright morning light through a half-open blind from the left. She wears a plain white ribbed cotton vest and a thin silver chain. Handheld phone held by someone standing in front of her at her eye height. Panel 1: framed tight on head and shoulders, the top of her head near the frame edge, chin lifted, warm slightly amused closed-mouth expression straight down the lens. Panel 2: same tight framing, both hands swinging up into the bottom of the frame toward the lens, hands motion-blurred, face sharp. Panel 3: framed a little wider — chest up — both hands out toward the lens mid-gesture, fingers open, grinning. Panel 4: framed wider again — waist up — one arm up and behind her head with the elbow high, torso leaning, head tilted, laughing. Unedited phone snapshot quality in every panel. Real 61-year-old skin — freckles across the chest and shoulders, laugh lines, crepey neck. No retouching, no smoothing, no makeup.
  2. VC Generate the take — 8s continuous video16 cr
    Start frame
    • sbc-panel1.png cropped from the storyboard
    Output

    take-c.mp4

    Generated / used

    8s generated, 64f delivered across 5 windows

    Call
    higgsfield generate create kling3_0 \
      --aspect_ratio 9:16 --duration 8 --mode std --sound off \
      --start-image sbc-panel1.png \
      --prompt "<below>"
    Prompt — verbatim
    Handheld phone held by someone standing in front of her, small natural handheld drift only — no pan, no zoom. She starts framed tight on head and shoulders, chin lifted, holding the lens with a warm amused closed-mouth expression, completely still for about a second and a half. She then swings both hands up into the bottom of the frame toward the lens and gestures with them, bouncing lightly on the spot to a beat, grinning. Finally she sweeps one arm up and behind her head with the elbow high, leans into it and sways, laughing, and holds that until the end. Light, unselfconscious, full of energy, like she is dancing for a friend who is filming her. Warm bright bedroom light. Unedited amateur phone video, no colour grade.
Take D

Gym corridor

frames 269–364 · 95f delivered · 22 cr

Serves blocks 26, 27, 28, 29

The source material

Source frame 269
f269
Source frame 280
f280
Source frame 290
f290
Source frame 300
f300
Source frame 301
f301
Source frame 311
f311
Source frame 321
f321
Source frame 331
f331
Source frame 332
f332
Source frame 338
f338
Source frame 344
f344
Source frame 348
f348
Source frame 349
f349
Source frame 354
f354
Source frame 359
f359
Source frame 363
f363

Cut list out of this take

Block 2632ffrom ≈1.0s — the relaxed stand
Block 2731ffrom ≈3.8s — the held flex
Block 2817ffrom ≈6.0s — hands clasped
Block 2915ffrom ≈8.5s — turned, glancing back

The hardest take in the film and the right one to A/B the two routes on. Background delta across the source's internal cuts measures under 1.0, so the camera must not move at all across ten seconds.

Performance arc to generate

0.0–2.5Standing relaxed, arms at her sides, weight settling, turned very slightly toward the lens.
2.5–5.0Raises both arms into a double-bicep flex and holds it, looking down the lens without smiling.
5.0–7.0Lowers her arms and clasps her hands in front of her waist, square to camera.
7.0–10.0Turns to show her back and shoulder, glancing back over her shoulder at the lens.

Steps

  1. SBDGym corridor storyboard image2 cr
    References in
    • S0-after-sheet.png
    Output

    sbd.png

    Call
    higgsfield generate create nano_banana_pro \
      --aspect_ratio 16:9 --resolution 2k \
      --image S0-after-sheet.png \
      --prompt "<below>"
    Prompt — verbatim
    A four-panel contact sheet, four photographs laid side by side in a horizontal row with thin white gutters between them, no text, no numbers, no labels, no arrows. Each panel is a full photograph that fills its panel edge to edge — the room fills the frame behind her in every one. No plain or empty backdrop anywhere. Every panel shows the SAME woman from the reference in the SAME plain gym corridor, shot from the SAME fixed camera at chest height, framing her from mid-thigh up, at exactly the same distance in all four panels. The corridor is identical in every panel: cream-painted breeze-block walls, a scuffed grey vinyl floor, a plain flush door with a steel push-bar on the left, a bare fluorescent strip across the ceiling throwing harsh flat light from directly overhead. She wears a teal racerback sports bra and plain charcoal high-waisted leggings. Panel 1: standing relaxed, arms at her sides, turned very slightly so her right side is toward camera, chin level, neutral confident expression down the lens. Panel 2: both arms raised into a double-bicep flex and held, elbows out and level with her shoulders, not smiling. Panel 3: arms down, hands clasped loosely in front of her waist, square to camera. Panel 4: turned with her back three-quarters to camera showing her shoulder and upper back, glancing back over her shoulder at the lens. Real muscle definition for 61, not exaggerated. Unedited phone snapshot quality, faint sensor noise, no colour grade. Real skin — sun damage, freckles, loose skin at the elbows. No retouching.
  2. VD Generate the take — 10s continuous video20 cr
    Start frame
    • sbd-panel1.png cropped from the storyboard
    Output

    take-d.mp4

    Generated / used

    10s generated, 95f delivered across 4 windows

    Call
    higgsfield generate create kling3_0 \
      --aspect_ratio 9:16 --duration 10 --mode std --sound off \
      --start-image sbd-panel1.png \
      --prompt "<below>"
    Prompt — verbatim
    Camera fixed on a prop at chest height — absolutely no camera movement, no pan, no drift, no push, no zoom, for the entire ten seconds. Only she moves. She stands relaxed with her arms at her sides, weight settling, turned very slightly toward the lens, for about two and a half seconds. She then raises both arms into a double-bicep flex, elbows out and level with her shoulders, and holds it, looking down the lens without smiling. She lowers her arms and clasps her hands loosely in front of her waist, standing square. Finally she turns to show her back and shoulder, glancing back over her shoulder at the lens, and holds that. Real muscle movement and real weight — the flex is controlled, not theatrical. Harsh flat fluorescent light from directly overhead. Unedited amateur phone video, no colour grade.
Take E

Car interior

frames 364–443 · 79f delivered · 12 cr

Serves blocks 30

The source material

Source frame 364
f364
Source frame 380
f380
Source frame 396
f396
Source frame 412
f412
Source frame 428
f428
Source frame 442
f442

Cut list out of this take

Block 3079ffrom ≈0.3s — continuous, no internal cut

Already a single continuous take in the source — 79 frames with no internal cut, the only genuine performance in the film. Nothing changes here; it was always one generation.

Performance arc to generate

0.0–1.0Right arm extended toward the lens, forearm large and blurred in the foreground.
1.0–2.0Pulls the arm back, brings both hands in to her chest.
2.0–3.0Settles back into the seat, relaxed, looking at the lens.
3.0–5.0Raises one fist into a bicep flex, laughing, and holds a warm closed-mouth smile.

Steps

  1. SBECar storyboard image2 cr
    References in
    • S0-after-sheet.png
    Output

    sbe.png

    Call
    higgsfield generate create nano_banana_pro \
      --aspect_ratio 16:9 --resolution 2k \
      --image S0-after-sheet.png \
      --prompt "<below>"
    Prompt — verbatim
    A four-panel contact sheet, four photographs laid side by side in a horizontal row with thin white gutters between them, no text, no numbers, no labels, no arrows. Each panel is a full photograph that fills its panel edge to edge — the room fills the frame behind her in every one. No plain or empty backdrop anywhere. Every panel shows the SAME woman from the reference in the SAME parked older estate car, shot from the SAME phone held low in her lap and angled up — the pale roof lining and the top of the windscreen fill the upper third of each panel, wide-lens distortion at the edges, bare trees and a flat overcast sky through the side window blown out at the left. She wears a faded black t-shirt with a small worn chest print and a steel watch on her left wrist. Panel 1: right arm extended toward the lens, forearm large in the foreground and slightly blurred, face further back and sharp, mouth open mid-word. Panel 2: arm pulled back, both hands coming in toward her chest. Panel 3: settled back into the seat, relaxed, hands down, looking at the lens. Panel 4: one fist raised in a bicep flex at chest height, arm close to her body, smiling with her mouth closed, head slightly tilted. Unedited phone snapshot quality, no colour grade, no retouching, real 61-year-old skin texture.
  2. VE Generate the take — 5s continuous video10 cr
    Start frame
    • sbe-panel1.png cropped from the storyboard
    Output

    take-e.mp4

    Generated / used

    5s generated, 79f delivered across 1 windows

    Call
    higgsfield generate create kling3_0 \
      --aspect_ratio 9:16 --duration 5 --mode std --sound off \
      --start-image sbe-panel1.png \
      --prompt "<below>"
    Prompt — verbatim
    Phone held low in her lap and angled up, small handheld wobble, no pan and no zoom. She begins with her right arm extended toward the lens, forearm large in the foreground. She pulls the arm back toward herself, brings both hands in to her chest, settles back into the seat, then raises one fist into a bicep flex at chest height, laughing, and holds a warm closed-mouth smile for the rest of the clip. Relaxed and playful. Flat overcast daylight through the car windows. Unedited amateur phone video, no colour grade.

Montage

Six stills, unchanged

Blocks 5–10, frames 149–178. These are photographs in the source, held three to seven frames, so they were never video and consolidation does not touch them.

IDFramesHeldScene line
M1149–1523fStanding in a supermarket aisle beside a trolley loaded with groceries, half-turned toward the camera with a polite closed-mouth smile, wearing a faded olive t-shirt. A long fluorescent-lit aisle receding behind her, other shoppers out of focus.
M2152–1564fStanding in a cramped kitchen behind a hob, a pan of food steaming heavily in front of her, arms folded, looking at the camera with a flat unamused expression, wearing a grey t-shirt. Cluttered worktop, a patterned cloth in the foreground, warm overhead light.
M3156–1604fSitting on the open tailgate of a parked pickup at sunset next to a grey-haired man in a blue shirt, legs dangling, both looking at the camera, wearing an orange t-shirt and jeans. Flat empty farmland and a low orange horizon behind them.
M4160–1644fStanding at a busy outdoor market holding a small item up near her chest, wearing a plain blue t-shirt. A dense crowd and stall tables of goods behind and in front of her, bright flat daylight.
M5164–1684fPropped up in a hospital bed in a pale patterned gown holding a paper cup, mouth open in a tired laugh at the camera. A metal bed rail across the foreground, white blanket, pale institutional walls, a window with a blind behind her.
M6168–1757fWorking as a cashier at a supermarket checkout, wearing a blue polo work shirt, leaning across the conveyor belt to scan a large pack of paper towels, not looking at the camera. Rows of checkout lanes and shelving behind her, hard fluorescent light.

Shared suffix — appended to all six

Candid amateur photograph of the woman in the reference sheet, taken by someone else on a phone, several years old. Slightly soft, slightly badly framed, mixed white balance, no colour grade, no retouching, no beauty smoothing. She is not posing for a photographer — she is just there. Flat steel-grey shoulder-length hair, no makeup, about thirty pounds heavier than her reference weight, soft arms and a rounded middle.

Against the previous plan

What consolidation changes

Block planTake plan
Video generations114
Generations per setupup to 41
Room instances per setupup to 41
Keyframes to generate144 storyboards + 1 still
Model asked tohold a poseperform, then be cut
Weakest framesdelivereddiscarded behind the cuts
Video spend85 cr66 cr
Total145 cr~92 cr

What this costs us

The risks consolidation introduces

Worth being straight about: this trades a set of known small problems for a smaller set of larger ones.

Within-take drift

Over ten seconds a model widens its framing and quietly flatters the subject. Mitigated by the cut structure — the delivered windows are ~1s each and the drift between them reads as the same shoot rather than different ones. Score luminance and subject scale at each window rather than across the whole take.

Ten seconds of locked camera

Setups A and D both require zero camera movement for the full generation. This is the single most likely failure, and it is why setup D is the A/B test rather than something easier.

The turn has to actually happen

Under the old plan every pose was a start frame, so a refused transition cost nothing. Now setup A must execute the rotation and setup D must execute the turn, or the later windows do not exist. Fallback is to split that one take in two, not to abandon consolidation.

Storyboard panels may bleed

Asking for four panels of one person in one room risks the model blending panels or producing a single wide image. If it does, generate the four keyframes separately as before and use panel-free start frames — the take consolidation still holds independently of the storyboard.

No seed

Still true, and it bites harder here: a ten-second take you like cannot be re-derived. Save every generation.

Durations and costs queried live — kling3_0 accepts 3–15s, std 2.0 cr/s, pro 2.5 cr/s · seedance_2_0 is the only video model accepting image_references alongside a start frame · every prompt is written for this build; character, rooms, wardrobe and copy are invented · no job takes a frame of the source video as a reference · source frames shown are analysis references with the caption band masked