24 Hours in Ancient Rome
One Roman worker's wage, one day, 83 scenes. A full test production run through VIA, from script to assembled film.
Same street, same man, all day.
The hardest thing in AI video is the fortieth scene looking like the first. These are three different scenes, drawn hours of story apart.
The brief
The episode follows Marcus, a young man dropped into Ancient Rome with one day's wage for an ordinary worker and twenty-four hours to live on it. It is told in first person, straight to camera, through streets, a lodging house, a bathhouse, a bakery and a wealthy house. It is the kind of video history channels make every week, and the kind AI tools usually break: one lead character on screen for most of it, recurring places, and a narration that names specific things the viewer must see.
Setup: the world, decided once
Before a single scene was drawn, VIA built the production bible from the script. Each recurring character (the lodging-house keeper, the bathhouse attendant, the dice gambler, the doorkeeper of a wealthy house) got a reference sheet with front, side and back views and fixed clothing. Each of the eight locations was drawn once as an empty set and approved with its fixed features. The main street, for example, keeps a tall insula on the left, a stair doorway with a worn threshold, a travertine fountain at the corner, a shrine in the right-hand wall and basalt paving with raised curbs.
Those features are written into every scene set on that street, and every still is judged against them. That is why the fountain and shrine are where they should be in each of the three pictures above.
Scenes: 83 checked shots
The script was split into 83 scenes, each tied to its line of narration and timed to the voice-over. For each one VIA wrote a shot description (who is in frame, where, doing what, holding what, from which camera) and an independent check compared it with the narration before any image was made. When the line mentions bread, the bread is in the shot; when it mentions the wage, the coins are.
Stills were drawn from the approved references and judged for identity, location, props and framing. The chosen still then became the first frame of its clip, with the motion described from that exact picture, so the video could not drift to a different face or room.
Film: judged as a video
The scenes were assembled with the narration on one timeline, 9 minutes 22 seconds across 14 sections, and reviewed as a viewer would watch it. Scenes that did not work were sent back individually; the rest of the film stayed as it was.
What we learned
- Name what you want drawn. "A crowd" draws an empty street; "a porter with a basket, two women at the fountain" draws life.
- One place, one reference. A location drawn fresh for each scene drifts. Views of one approved set do not.
- Approve the still before you pay for motion. Motion inherits every flaw of its first frame.
- Judge the film, not the clip. A shot that looks impressive alone can still break the sequence.
More on each of these in the guides: keeping characters consistent and avoiding AI slop.
More from the production