Every video on the RealityDrift account is synthetic. There is no miniature city, no wave tank, no crew in "EFFECTS CREW" shirts, no crocodile and no train to the Moon. What you see is a single still image that a video model was told to move, and the reason it reads as phone footage from a real place is that every stage of the pipeline exists to protect that read.
This is the free lesson, so it has two jobs. It is the map of the whole method, detailed enough that you can follow the shape of a production from concept to post. And it tells you what each of the fifteen lessons after it hands you, so you can decide whether the rest is worth paying for. I have tried to make it useful on its own.
What you are making
The account runs five content formats. Each has its own engagement engine and its own rule about where it may be published. Lesson 4 covers them properly; here is the short version so the rest of this lesson makes sense.
- bts_studio, the core. A fake behind-the-scenes miniature effects studio: a film crew on a bluescreen soundstage films a miniature city being destroyed by a tsunami, a dam break, a volcano, a creature. It looks like real set footage and it is entirely generated. The engine is the "is this real?" argument in the comments.
- capture_pov. A first-person subject walks toward a small-scale phenomenon (a cloud, a miniature sun, a blizzard) and captures it in a preserving jar or a bag. The engine is wonder and satisfaction, not debate.
- transit_window. An ordinary commuter train or bus filmed on a phone, where the window is a portal and the vehicle crosses into one impossible destination through a real transit occlusion such as a tunnel mouth. Nobody on board reacts. The engine is sends and saves.
- beast_kin. A famous person and the creature they would come with, photographed together in a boring real location and played completely straight. Five separate clips, one per person, cut together.
- premiere_carpet. A real famous person restyled as a fictional character on the red carpet at that property's premiere, shot from inside the photographers' pit. Same five-clip shape.
The last two use real likenesses, so they are fan parody: Trial Reels only, never boosted, never sold. All five run through the same production pipeline. Only the assets change.
The six stages
Every video goes through six stages in a fixed order. The lesson numbers point at where each stage is taught.
- Concept. Either you reverse-engineer a video that is already working (the reference flow, lesson 5) or you generate an original concept and score it through four filters that predict whether it can travel (lesson 4). Nothing is generated until it has passed the filters, because generation is the expensive part and a weak premise cannot be fixed downstream.
- First frame. A single still generated with Nano Banana Pro (lesson 6). This is where realism is decided. A video model can only move what it is given; it cannot turn a glossy render into a phone photo.
- Last frame. Optional, used only when the shot has a clean before and after. It is a single edit pass on the first frame, never a fresh generation (lesson 7).
- Video. Image-to-video with the first frame locked as the opening pixels and a written paste driving motion, camera and sound. Three engine families, three prompt dialects that do not mix: Kling (lesson 8), Seedance 2.0 and 2.5 (lesson 9), MiniMax H3 (lesson 10). For shots where the camera path has to be exact, a grey Blender blockout goes in as a reference video and the model follows it (lessons 11 and 12).
- Upscale and deliver. Drafts are generated small and the keeper is upscaled to 1080x1920 with a fixed recipe (lesson 13). Generating at full resolution from the start is the expensive way to make a video you may throw away.
- Package and publish. Caption, hashtags, pinned comments and the AI label, per platform, inside the disclosure rules that bind a creator in the EU (lesson 14).
Lesson 15 is what to do when a take fails. Lesson 16 covers the series formats, where one environment plate is edited five times and the clips are joined in CapCut.
The one rule everything depends on
The still comes first, and the still carries every decision it can.
The image model is cheap, fast and controllable: you regenerate a frame for about thirteen cents and see it in seconds. The video model is expensive and only partly controllable: a ten-second take costs between fifty cents and several dollars, and you will not get exactly what you asked for. So the pipeline pushes every question it can into the frame: camera height, lens character, materials, light direction, the number of people, the text on their shirts, the exact state of the set. The video paste then describes motion only. It never re-describes what the frame already shows, because every re-description is an invitation to change it.
When a shot fails, the diagnosis runs in the same order. Wrong look: fix the still. Wrong motion: fix the paste. Wrong camera path on a shot that needs precision: stop prompting the camera and block it in Blender.
What "hyper-real" means here
The style rules are identical for every genre, and they are close to the opposite of what most AI video looks like.
- Vertical, smartphone documentary look, deep focus, a naturalistic Rec.709 grade. Never a cinematic colour grade, never shallow-focus bokeh, never a drone-advert sweep.
- Physics written in real terms. Name the material, the order in which things fail, and the speed in words a person would use: "snaps instantly", "walking pace", "one sheet of water clears the first row before it breaks".
- Deliberate imperfections. Handheld shake, a droplet or dust on the lens, audio clipping on the loudest impact, a reaction that lands slightly late.
- Diegetic sound always. An ambient bed plus timed effects. On the studio genre, at least one shouted crew line.
- Aspect ratio, resolution and duration never appear inside a prompt. Those are parameters you set in the tool. This sounds pedantic and it is not: a model that reads a ratio or a pixel size in the text will happily render a phone screen showing a video, or letterbox the frame.
Nothing on that list is decoration. Each item exists because its absence is a tell that a viewer catches in the first second, and the first second is where the skip happens.
The artefacts
A RealityDrift video is produced from one markdown file, the prompt file. It holds the concept, the shot strategy, the generation parameters, the still prompts, the master beat table, and one paste per video engine. The Vault sells these files; two of them are free and quoted here so you can see what you are working towards.
TSUNAMI.md is the oldest file in the Vault, the remake of the video that started the genre (a miniature Miami in a wave tank, roughly 700 million views in two days). It predates the house template, so its first frame is described inside a timecoded breakdown and a scene inventory rather than in a labelled block. This is the opening beat and the inventory, verbatim:
0:00-0:02
ACTION: Water in tank is disturbed; a wave begins forming on the right.
SUBJECTS: Crew members scattered along foreground retaining wall, looking at tank. Black t-shirts. Static.
CAMERA: Wide shot, elevated high angle, static.
FOCUS: Deep focus.
KEY DETAILS: Blue screen with red tracking crosses, water tank, miniature pier with Ferris wheel, foreground camera dolly track.
- Massive blue screen backdrop spanning the entire rear wall, marked with evenly spaced red crosses.
- Overhead studio lighting grid with large diffuse panels.
- Large water tank taking up the midground.
- Concrete retaining wall in the foreground.
- Camera dolly tracks laid parallel to the retaining wall.
- Large black telescoping camera crane (Technocrane style).
- Miniature sandy beach.
- Miniature wooden pier with a small Ferris wheel and a blue/yellow structure.
- Miniature coastal road lined with small palm trees.
- Miniature city block of pastel, low-rise Art Deco buildings.
- Miniature city block of modern, high-rise glass and concrete skyscrapers.
- Dozens of crew members wearing casual clothes (mostly black t-shirts, some reading "EFFECTS CREW").
Read that inventory as the shopping list for the first frame: every item on it has to be in the still, at the right position and scale, or the video model has nothing to move. The same file's first video paste, written for Kling, shows how little the motion text repeats the frame:
Vertical video, wide elevated shot, static then panning left. A massive artificial wave builds in a large indoor water tank, rolling left towards a miniature wooden pier with a Ferris wheel. Crew members in black "EFFECTS CREW" shirts watch from behind a concrete retaining wall in the foreground. A massive blue screen with red tracking markers serves as the backdrop. The wave crashes violently into the pier, exploding into a huge splash and obliterating the Ferris wheel. Studio lighting is bright and even. Deep focus, realistic behind-the-scenes documentary style.
AUDIO: Loud mechanical clanking, a shout of "Action," building low rumble of water, massive explosive crash with splintering wood.
One camera move, one primary action, a material named ("splintering wood"), a shouted line, and an audio line at the end. Its negative prompt lists the tells to suppress: CGI look, cinematic colour grading, shallow depth of field, life-size scale, a clean environment, a missing bluescreen, missing crew, a static camera, slow-motion water. Lesson 8 explains why that list is kept short and why "camera shake" is never on it.
The second free file, RAPTOR_GATE_RUN.md, is the modern form: a five-shot creature episode with nine reference image prompts, a master playbook, a Kling multi-shot paste, a Seedance storyboard, a split-generation fallback and thirteen numbered iteration notes. Here is the opening of its Shot 1 composition still, which doubles as the Kling first frame:
Vertical smartphone photo taken inside a large soundstage dressed as a paddock corridor. The phone is held at chest height beside a steel gate unit at the near end of the corridor, lens aimed straight down its length, slightly tilted framing, deep focus, everything perfectly sharp with no motion blur anywhere. At the right edge of frame, close to the camera, the near riveted pillar of the gate unit and its raised grey guillotine slab with a small mesh window, the chain drive on the pillar, and the gate operator in a yellow hard hat and black t-shirt printed "EFFECTS CREW" with one hand on the red lever of his control pedestal.
Notice what the sentence spends its words on: where the phone is, what is nearest the lens, what the shirt says in quotation marks, and the instruction that nothing is blurred. Those are the four things a still gets wrong by default. The prompt runs for another two thousand characters in the same register and closes with a line that appears in every still we make: "It must not look AI-generated, but like a photo taken by a real person on set."
The file also records what went wrong first: a single-frame version that packed the escape gate three metres from the runner, a moderation rejection caused by predator vocabulary, and the discovery that Seedance refuses any reference image containing a photorealistic human face. Those notes are the reason the Vault files are worth more than the prompts alone. The failures are the expensive part, and they are already paid for.
A worked mini-example
Here is one concept run through all six stages at the level of detail you will operate at after lesson 14. The concept is DAM_BREACH_ALPINE from the account's first idea batch: a scaled concrete arch dam above a miniature alpine village fails on cue, and the crew films the flood wave erasing the main street.
Concept (lesson 4). The batch scored it 5/4/5/5, nineteen out of twenty. Debate trigger 5: dam-failure test models are real civil-engineering practice, so belief is defensible. Readability 4: a hairline crack crawling across a dam face with water jetting from one point is legible mid-action. Producibility 5: water is the pipeline's most reliable physics. Brand and safety 5: the studio wrapper removes any news plausibility. The planted tell for eagle-eyed commenters: the reservoir holds slightly too much water for the visible tank depth.
Shot strategy (lesson 5). First frame only, no end frame. Chaotic water does not want an end frame; an end frame turns a burst into a morph. One continuous handheld take, one camera move (a slow push toward the dam face), fifteen seconds, with the burst placed late so the final third has its own driver.
First frame (lesson 6). One still from Nano Banana Pro, generated in Google AI Studio with the image model selected and the vertical ratio and 2K size set as parameters in the tool. The prompt runs to about 2,500 characters and covers, in order: the phone position (chest height, three metres back from the set edge, tilted a few degrees), the dam (grey cast concrete with formwork lines, a hairline crack, one jet of water), the reservoir behind it (dark water, a visible tank wall), the village below (about thirty chalets, a church steeple, a main street running toward the camera), the crew (three people in black "EFFECTS CREW" shirts behind a monitor cart, one pointing), the bluescreen with tracking crosses, overhead softboxes, cable runs and warning tape, and the capture conditions: iPhone original camera aesthetic, mild sharpening, film grain, no motion blur. Cost: $0.134. If the render comes back with six chalets instead of thirty, or the dam twice the size of the village, regenerate. Never carry a wrong frame forward.
Draft (lesson 10). The first moving take runs on MiniMax H3 Max at 480P for ten seconds: fifty cents. You are checking three things only: does the burst happen where the paste says, does the water arrive as a wall rather than a trickle, and does the camera push read as a push. Two or three drafts are normal.
Final (lessons 8 and 9). Once the draft proves the blocking, the same first frame and a paste in the final engine's dialect: Kling 3.0 if you want the negative field and camera controls, Seedance 2.5 if you want timestamped beats and joint audio. Fifteen seconds at 720p on Seedance 2.5 is $3.47; on Kling the cost is membership credits.
Upscale (lesson 13). The keeper is rotated to landscape, upscaled through Topaz so the short edge becomes 1080, rotated back and forced to exactly 1080x1920 with an H.264 re-encode. Frame rate untouched. Two ffmpeg commands and one upscale run.
Package (lesson 14). A three-line caption under about a hundred characters: a hook specific to this video, a one-line AI disclosure that names the tool, and a call to action. Something like: "Alpine unit, dam day." then "AI-generated with Seedance 2.5." then "Full prompts in my vault, link in bio." Then three to five hashtags, up to three pinned comments within minutes of posting, the platform's AI toggle on, and a cross-post to Facebook the same day.
Total generation spend for that video, including two drafts and one wrong frame: roughly $5 to $6.
What one video costs
Prices are the official pay-as-you-go rates at the time of writing and they move. The shape of the budget does not.
| Stage | Tool | Rate |
|---|---|---|
| Frames | Nano Banana Pro (gemini-3-pro-image) | $0.134 per 1K or 2K image, $0.24 per 4K image |
| Frames, cheaper iteration | Nano Banana 2 (gemini-3.1-flash-image) | $0.067 per 1K image, $0.101 per 2K image |
| Drafts | MiniMax H3 Max at 480P | $0.05 per second |
| Drafts | Seedance 2.0 Mini | $0.036 per second |
| Drafts | Kling 2.5 Turbo Pro | about $0.31 per five seconds in credits |
| Finals | MiniMax H3 | $0.08 per second at 768P, $0.13 at 2K |
| Finals | Seedance 2.0 at 720p | $0.15 per second |
| Finals | Seedance 2.5 | $0.103 per second at 480p, $0.231 at 720p, $0.569 at 1080p |
| Finals | Kling 3.0 | membership credits per second, more with native audio |
| Upscale | Topaz through Replicate | billed per run; check the model page for the current rate |
A keeper lands between $3 and $15 in generation spend, with heavy iteration days reaching $25. Posting daily across platforms at scale is a $500 to $2,000 monthly budget. You do not need that to start: $50 of credit takes a concept from still to a finished fifteen-second reel with drafts along the way, and lesson 2 sets the accounts up so that one wrong parameter cannot quietly empty them.
The tools, and where each one is taught
| Tool | Job in the pipeline | Lesson |
|---|---|---|
| Claude Code with the Starter Kit | Runs the house method as commands: scoring concepts, writing prompt files, packaging | 3 |
| Gemini in Google AI Studio | Watches a reference video and returns a structured analysis | 5 |
| Nano Banana Pro and Nano Banana 2 | First frame, last frame edit, reference stills | 6, 7, 16 |
| Kling 3.0, O3, 2.5 Turbo Pro | Image-to-video with negatives, camera presets and cheap drafts | 8 |
| Seedance 2.0 and 2.5 (Dreamina or ModelArk) | Image-to-video with joint audio, references, timestamps on 2.5 | 9 |
| MiniMax H3 and H3 Max | The cheap end of finals and the cheapest reference-video path | 10 |
| Blender with the Blender MCP | Grey blockouts whose camera path the model follows | 11, 12 |
| ffmpeg, yt-dlp | Frame extraction from references, the rotate-upscale-rotate recipe | 5, 13 |
| Topaz through Replicate | The upscale itself | 13 |
| Cloudflare R2 | Public URLs for input frames and reference clips | 2 |
| CapCut | Concatenating series clips, trims, on-screen text | 14, 16 |
How the course is organised
| Lesson | What you leave with |
|---|---|
| 2. Accounts and tools | Every account opened, every program installed on macOS or Windows, keys stored safely, spend limits set |
| 3. Claude Code and the Starter Kit | Claude Code running inside the kit, the six skills understood, four sessions typed out |
| 4. Finding and scoring concepts | The four filters, the genre catalog, a scored idea batch with a first frame prompt per concept |
| 5. Reverse-engineering a reference video | A frame-by-frame analysis of a working video, turned into a prompt file skeleton |
| 6. The first frame with Nano Banana Pro | A still that reads as a phone photo, and the density and gloss failures it avoids |
| 7. The last frame edit pass | The single-pass edit, and knowing when to skip it |
| 8. Kling image-to-video | A moving draft from your frames, with a negative list that holds |
| 9. Seedance 2.0 and 2.5, two dialects | Both pastes from one playbook, and the face rule that rejects inputs outright |
| 10. MiniMax H3 | The prose dialect, the one-paragraph soundscape, and the 480P draft habit |
| 11. Blender previz: build the blockout | The Blender MCP connected to Claude Code, a grey scene with a keyframed camera |
| 12. Blender previz: submit and iterate | The playblast submitted as a reference video, and how to read the takes |
| 13. Upscale and deliver | A 1080x1920 master from a 480p draft, with no interpolation |
| 14. Package, label and publish | A labelled, compliant post on four platforms |
| 15. Troubleshooting failed takes | The diagnosis order, and the fixes for the twenty most common failures |
| 16. Series formats: plate plus edits | Five clips from one environment plate, cut in CapCut, for the celebrity tracks |
The prompt files in the Vault are the worked examples for lessons 6 to 10. Each one is the complete recipe for a published video, including the notes on what failed before it worked. Read the two free files alongside these lessons and the format will make sense fast.
Before lesson 2
You need a Google account, a payment card, and a computer that can run Blender (any laptop from the last five years, macOS or Windows). You do not need to know how to code. You will use a terminal in lessons 2, 3, 11, 12 and 13, and every command you need is written out for both operating systems.
Checkpoint: you can name the six stages in order, you know why the still comes first, and you can look at the two quoted excerpts and say what each sentence is there to prevent.
Next: lesson 2, Accounts and tools, opens every account, installs every program and puts the keys somewhere safe.

Discussion
Members and pack owners can post
No comments yet. Questions about a step, or a note on what changed for you, go here.
Sign in to see whether your account can post here. The discussion is open to Lab members and Previz Pack owners.