From generative video to programmable video

When most people say “AI video” right now, they mean one thing: you type a prompt, a model thinks for a while, and a clip comes back. It’s impressive. A few sentences in, a few seconds of footage out.
But there’s another way to use AI to make video, and it looks more like web development than filmmaking. Instead of asking AI to generate the frames, you ask it to write the code that generates them.
# two ways to make a video
INPUT STEP STEP OUTPUT
generative prompt → model ············ → pixels
code-based prompt → code → frames → video
--------------------------------------------------------
generative ≠ run it again, get something different
code-based = run it again, get the same framesA generated clip is a finished artifact: if it’s not quite right, you ask again and get something different. Code renders the same frames every time, so you can keep working on it, nudging it until it’s exactly right and changing only what needs to change. That predictability is what makes this approach useful for professional work, and AI coding tools are what suddenly make it practical.
What if a video were just a website with a timeline?
Remotion has been doing this for years, letting developers build videos out of the same React components they use for websites. More recently, HeyGen open-sourced HyperFrames, which does the same thing with plain HTML and CSS. Either way, a browser draws each frame, and the frames are stitched together into a video.
Both rest on the same idea: the browser is an extremely capable graphics engine.
It already renders typography, illustration, animation, 3D, images and video. All it’s missing to become a video renderer is a clock it can step through one frame at a time. Give it one, capture what’s on screen at each step, and you have a video.
An exploded view of the capture: a clock and the scene the browser draws for it, both running steadily, and a film strip of empty frames stepping across the scene. Each frame stops over the scene, takes a screenshot of it as the scene reaches that frame's time, and slides on as the next empty frame arrives.
If you’ve written React, you know the mental model. You don’t reach in and change an interface; you describe what it should look like for a given state:
# UI is a function of state
UI = f(state);Video works the same way, with time in place of state:
# frame is a function of time
frame = f(time);You don’t tell a title to fade in. You describe where it should be at any moment: invisible at frame 0, halfway there at frame 15, fully in place by frame 30. Ask twice and you get the same answer.
Rendering, not rolling the dice
That last part is the real benefit. Ask a generative model for the same clip twice and you get two interpretations. Render the same code twice and you get the same frames, down to the pixel.
It sounds like a technical footnote, but it changes how the work goes. You can tune a coded piece the way you’d tune a layout: try something, look, adjust, look again. Every round builds on the last, and nothing you’ve already gotten right drifts while you work on something else. You stop when it’s right, not when you get lucky.
It matters even more once a client is involved, because revisions on a motion piece tend to be mundane:
- Make the logo 20% smaller.
- Hold this frame for another second.
- Ease the title in more gently.
- Replace this headline.
- Make a vertical version.
With a generative model, each request is a new roll of the dice, and there’s a real chance the thing you liked doesn’t survive.
With code, each request is a surgical edit: a size, a duration, a curve, a line of text, an aspect ratio. Everything the client already approved stays exactly where it was. You’re not spinning up a black box again and hoping for the best. You’re changing one thing and rendering again.
One video becomes a hundred
The second benefit follows from the first. If a video is a function, you can give it different inputs:
<Promo title="Poetry in Motion" theme="blue" format="vertical" duration={12} />Change the title, theme, format or data and render again.
That’s where a lot of tedious production work lives: a campaign with a dozen headlines, the same piece laid out properly at 16:9, 1:1 and 9:16, a card for every speaker on the schedule, this month’s numbers animated in the same style as last month’s. None of it is creative work, but someone usually does it by hand. Here it becomes a list of inputs.
Nobody wants to hand-edit 400 timelines. A design system for video turns out to look a lot like a design system for the web.
Why this matters now
If all of this has been possible for years, why does it matter now?
Because the renderer was never the bottleneck. Writing the program was. Programmatic video needed a developer with decent motion instincts, and that’s a small group of people without much spare time.
Programmatic video isn’t new. What’s new is that writing the program has gotten cheap.
With a coding agent, you describe the animation and it writes or changes the code. “Stagger the three cards in from the left, a beat apart, with a slight overshoot” is a sentence for you and a few lines of code for the model. And the medium is familiar territory for these models: HTML, CSS and JavaScript.
This isn’t a replacement for generative video
This isn’t a “code good, generative bad” argument. The two are good at different things.
Code excels at things with structure: typography, UI, diagrams, charts, logos, geometric animation, product graphics, and anything that has to be exactly the same every time.
Generative video excels at what’s hard to describe procedurally: people, environments, cinematic imagery, organic movement, realistic camera work. I see what it makes as an input to the system, not its source of truth.
So stack them:
# the model supplies texture; code supplies structure + control
STEP LAYER BY SUPPLIES
[1/3] generative assets / footage model texture
[2/3] deterministic composition code structure
[3/3] typography + graphics + data code control
-------------------------------------------------------
✓ video renderedLet a model generate the footage, then bring it into a coded composition where the titles, graphics, branding and timing are exact and editable.
The generative part supplies the texture; the code supplies the structure and control.
A small experiment
Arguments are cheap, so I tried it: the stack above in miniature, with AI-generated footage under coded layouts. I gave a coding agent a brief and a folder of assets and had it build the piece as a plain web page, with no video framework at all. A browser stepped through the frames and they were encoded into a video, which took about 90 seconds for a 6-second render.
The brief: three event promos (a disco night, a data science talk and a game jam), each already designed as a 16:9 cover and a 4:5 Instagram post. I wanted one seamless loop of about 6 seconds that cycles through all six. Each graphic floats as a card over a blurred copy of its own background video, and at every change each element moves and scales to its place in the next layout. The inputs were:
- the six layouts, as PNG exports and as HTML source from the design tool
- three AI-generated 5-second background loops
- one screenshot of the floating-card look
Here’s how it went:
- Plan it: “Here are the assets. Build this loop, plus an admin view to review it.” It initially offered to rebuild the layouts from the PNGs, but I pointed it at the original design source instead. It ported all six with their styles intact, and each matched its PNG export on the first render.
- Make it move: The first render held every layout perfectly and fell apart in between. “Disco Night” sits on one line in 16:9 and two in 4:5, so the title cross-dissolved into a smeared double image. Text also ran ahead of the card as it resized and got clipped at its edge. The fix was structural: each title became individual words that fly to their new positions, and every element now moves with the card.
- Swap the footage: When the final background renders arrived, swapping them in meant changing one file path per event and rendering again.
- Adjust the timing: I wanted each layout held longer and the transitions slower, so I did this one myself in the admin view: hold from 0.4 to 1 second, transition from 0.6 to 2, then render. No code changed, just two numbers in a settings file.
- Generate variations: The variations were already built in: one set of templates, three sets of event details, two aspect ratios. A fourth event would be one more entry in a list.
The layouts worked on the first try because the agent had real design source instead of screenshots. The motion took two rounds of fixes. The agent found those problems by rendering and checking still frames from the middle of each transition. The fixes changed the motion without disturbing the layouts that were already right.
The timing change is the one that matters. In a generative tool, “hold it longer” means regenerating the clip and hoping everything else survives. Here it was two numbers. Everything else rendered exactly as before, just spread over 18 seconds instead of 6.
Control is the point
For professional work, this is the part that matters. Client video is rarely a single shot at something impressive. It’s a brief, a first cut, feedback, a second cut, a client who loves it except for the logo, and a vertical version due Friday. That process depends on getting something right and keeping it right.
That’s what determinism buys you. You’re not rolling the dice and hoping for a good take; you’re building a system and tuning it until it’s exactly what you want. When the edits come back, you make them surgically, and everything already approved stays approved. And once the system exists, the low-value production work—resizes, variations, next month’s version—is just a different set of inputs.