One prompt, three different lions — and what locks a shot down.
The full script, word for word — 909 words, about 4 minutes to read. Current as at August 2026. AI tools change quickly; if something looks different when you try it, check the product's own help pages.
Mel - 04 - Making video without a camera
Hi, I'm Mel. A video generator can give you a convincing lion in eight seconds. The harder job begins when you want the next shot to feature the same lion, in the same place, with the same visual language. We ran one prompt three times, added two reference images, then anchored the shot with a finished first frame. The results show three different levels of control.
First, separate two products that share the name AI video. A synthetic presenter turns a script into a person speaking to camera. A footage generator creates a shot that was never filmed: an animal, a landscape, a product or an impossible event. This video is about generated footage. The input may be words, images, existing video, or a combination of them.
Take one gives us a broad-chested lion, left of centre, walking through pale grass. It is plausible, well lit and usable.
Take two moves the lion closer, changes the mane and shifts the light. The action is the same; the animal and camera position are new.
Take three changes them again. The face, body, distance and ground all move. One prompt produced a category of shot, not one repeatable lion.
Text-only generation did not fail. All three clips are convincing, and the prompt constrained them to a male lion, dry savannah, golden light and a low, still camera. It controlled the idea. It did not establish an identity. That difference changes how you plan a sequence. If each shot can contain any lion, text may be enough. If the viewer must recognise one lion across several shots, the generator needs visual evidence.
We supplied two Ingredients: a cutout of the lion as the subject, and a savannah image as the environment. Then we described how each should be used. These are references, not pieces being pasted into a collage. Flow interprets them while creating a new moving scene. Our text asked for dry grass, while the environment image was greener. The model blended the signals rather than copying either one directly.
The first referenced take picks up the acacia-tree horizon, but it still creates a new lion, scale and path through the grass.
The second take shifts the animal forward and changes its proportions. The references influence the shot; they do not freeze it.
The third take changes the distance and foreground again. Ingredients improved direction more than identity in this test.
Then we supplied the finished first frame. The lion, landscape, lighting, scale and camera position were fixed. Flow mainly had to create movement.
Our result is a control ladder. A prompt describes the shot. Ingredients steer the subject, location, product or visual style. A start frame anchors the visible composition. Each step reduces the number of decisions left to the generator, which is why the start-frame clip looks strongest. It is also why a poor start frame stays poor: blurry details, awkward anatomy and a bad join become the foundation of the motion. Use the lightest control that gives the sequence what it needs.
Generate one short action at a time. A lion walks forward. It stops. It looks to the side. Long generations give the model more time to change the face, body or background. For a continuation, save the final frame of the useful clip and use it as the next opening frame. Join the selected shots afterwards. Six controlled seconds repeated is a production method; one perfect minute is a lottery ticket.
Choose the generator by what you need to control. Google Flow combines text, Ingredients, Frames and scene assembly in one workspace. Runway is useful when a prepared image, directed camera movement and the hand-off from one final frame to the next shape the workflow. Kling offers Motion Control when a character needs to follow movement from a reference performance. Luma's Modify Video starts with footage you already have and changes its world, style or elements while preserving the underlying motion to a chosen degree. These routes overlap. The job, source material and control required should make the choice.
Watch the complete clip at full size before it goes into anything public. Check whether the subject remains recognisable, whether limbs and paws stay coherent, whether text and logos survive, whether objects touch and move believably, and whether the final frame still belongs to the opening frame. A polished first second can hide a poor seventh second. Keep the misses as evidence while developing the shot, then remove them from the finished sequence.
Video generation is metered, so the cost shown before confirmation is part of the creative decision. Three variations use three generations even when they arrive from one request. Returned clips can consume the allowance whether or not you choose them. The downloaded files also carry provenance. Every clip in this test contained a signed record saying it was created by Google Generative AI, plus an imperceptible SynthID watermark. That is useful information to preserve. Other platforms and export routes may mark files differently, so check the file you receive rather than assuming.
Start by deciding what must remain consistent. Use words for the idea, Ingredients for direction, and a completed frame when composition has to hold. Generate short shots, inspect the movement and keep the strongest result. AI video is already capable of producing footage worth using. Treat it as a sequence of controlled attempts, with evidence at the start and judgement at the end.