Scenema Team

How to Make a 6-Minute AI Explainer Video From a Short Prompt

Every AI video tool ships an impressive fifteen-second clip. Past thirty seconds, the wheels come off. Characters drift into strangers by shot five. The narrator’s voice swaps timbre mid-sentence. The hard part is making dozens of separate generations feel like one continuous film.

This tutorial walks you through a working solution. The finished piece is a six-minute Vox-style animated explainer on the financial mechanics of the 2026 FIFA World Cup. You will see how it was built in Scenema from a single 79-word prompt and one style-preset selection, in about fifteen minutes total. All fifty-four shots share the same recurring character, the same narrator voice, and the same paper-collage aesthetic. Nothing was hand-stitched afterward.

The finished six-minute explainer. Audio on.

What you will learn

  • How Scenema turns a short prompt into a long-form explainer.
  • Why the same character stays visually consistent across fifty-four shots.
  • How to reproduce the piece from scratch, in about fifteen minutes.

For readers who prefer video walkthroughs, the full YouTube tutorial covering the same ground is below.

Video walkthrough of the same build, roughly ten minutes.

The two inputs

Total user input for the finished six-minute film was a 79-word content prompt and a single style-preset selection.

The prompt, verbatim

A dry, clinical, and data-driven documentary-style video exploring the financial
mechanics of the FIFA World Cup. The video should present specific statistics
regarding ticket prices relative to minimum wage, FIFA's non-profit status and
tax advantages in Zurich, total revenue projections for the 2023-2026 cycle, and
the multi-billion dollar broadcasting rights deals with networks like Fox,
Telemundo, the BBC, and ITV. The tone should be objective, analytical, and
critical, focusing on the disparity between corporate profits and the average
worker.

The style preset

The style preset selected was Vox Explainer. That preset ships with a canonical palette (aged newsprint tan and charcoal black, punctuated by a single aggressive hot-red accent for key data and warnings) and a canonical motif vocabulary (halftone cutouts, staggered stat numbers with overshoot springs, layered paper dioramas, ALERT WASH transitions on ideologically loaded moments). The preset does not dictate content. Content came entirely from the prompt above.

Click through the finished project

The full project is public. Every shot, every prompt, every reference image, and the finished export are viewable directly. Everything visible below is what Scenema produced from the two inputs above.

The interactive viewer is the honest version of this article. Anything the article claims about character consistency, per-shot prompts, or the shot list can be checked directly against the source project.

What actually held that usually does not

Character consistency across fifty-four shots

A small halftone-cutout figure labelled the New Jersey worker appears in roughly a third of the shots, from the opening tower of $10,990 tickets to the closing bench-under-a-floating-stadium shot. Every appearance keeps the same silhouette, the same tired posture, the same halftone rendering, the same faded newsprint palette. The recurring stadium, the Aramco logo, and the category 1 ticket all hold the same way.

Reference images are generated once when the treatment agent identifies each recurring entity. Every downstream shot that mentions the entity via its @tag receives those reference images as generation inputs. The model is never asked to invent the character from scratch mid-project. It is only ever asked to place a known character into a new composition.

Single voice across the full six minutes

The narrator delivers all 1,100 words of the script in a single continuous voice with no swap, no drift, and no perceptible re-recording seam. This holds because Scenema generates the entire narration in a single pass through one voice model with one voice ID, rather than generating scene by scene and stitching the chunks. Long-form voice generation is a solved problem when the pipeline treats the narration as a single utterance.

Style coherence, injected at the shot level

The paper-collage-on-aged-newsprint aesthetic is a treatment-level constraint. When the style preset is selected up front, the treatment agent commits to concrete palette and motif language in the written treatment. That language is then injected into every single per-shot prompt written by the director agent downstream. Every keyframe prompt starts with the same palette instructions and the same motif vocabulary. There is no room for a shot to reinvent the visual language, because the shot prompt itself cannot be written without inheriting the preset’s language verbatim.

Under the hood

The pipeline runs in one direction. Nothing later informs anything earlier.

1. Narration script

Generated first, because everything downstream is driven by it. The treatment agent reads the input prompt, considers the style preset, and produces a full narration script grounded in the actual FIFA financial numbers cited in the prompt. About 1,100 words across eleven paragraphs.

2. Visual treatment

Informed by both the narration and the style preset. Includes concept, palette, motifs, and an eight-act structure. You have the option to review and edit the treatment before any shot design starts. This is the recommended entry point for creative direction.

3. Shot list

The director agent takes the eight acts of the treatment and the eleven narration paragraphs and decomposes them into fifty-four individual shots. Each shot has a specific piece of narration it is meant to illustrate, a duration derived from the pace of that narration text, and a placement in one of the eight acts.

4. Character reference images

Generated before any per-shot keyframe. The tool identifies four recurring entities (worker, stadium, Aramco logo, category 1 ticket) and generates one canonical reference image per entity. These references become reusable inputs for every downstream keyframe that features the entity.

5. Per-shot keyframes

Every shot receives a director-written keyframe prompt combining subject, framing, lighting, and mood. Any entity appearing in the shot is referenced by its @tag, which attaches the corresponding reference image to the generation.

6. Per-shot videos

Each shot’s video is generated from the corresponding keyframe plus a director-written video prompt describing camera and subject motion.

7. Narration audio

Generated in one continuous pass, using a single voice model and voice ID for the entire six minutes.

8. Final export

All fifty-four video clips, the narration audio, and any transitions from the shot list are muxed into a single mp4.

Reproduce it yourself

You can clone the public template into any Scenema account with one click. Cloning mostly carries over the starting prompt so you do not have to type it from scratch. The agent then re-derives the treatment, shots, and characters, which means your run will produce a video in the same shape and style but with natural variability in specific compositions and characters.

Building from scratch is the same three steps used to produce the finished piece.

  1. Type a short prompt describing what the video is about and what tone it should carry.
  2. Select a style preset from the preset library.
  3. Approve the treatment when it appears, or edit it first.

Everything after that runs automatically until the mp4 lands in the export tab. About fifteen minutes end to end. The Studio shows a live credit estimate before every generation so you always see the cost before you spend.

What broke, and what is still hard

The Vox preset is opinionated

The preset produced a highly coherent aesthetic across all fifty-four shots. However, its canonical motifs (halftone cutouts, hot-red ALERT WASH transitions, staggered stat number entries) do dominate the visual grammar in a way that is instantly recognizable. If you want a more restrained aesthetic, pick a different style preset from the library, or bring your own reference and register it as a custom style in Scenema.

Character identity is silhouette-level, not pixel-perfect

The recurring worker cutout is unmistakably the same character across all his appearances. A viewer who freezes on individual frames will see minor variation in the halftone dot pattern from shot to shot. Fine for stylized aesthetics like this one. Less forgiving for photoreal work, which is why picking a stylized preset matters most when character continuity is critical.

Visual continuity between shots is the real open problem

Narrative continuity is handled. The narration flows, the shot list follows the story, and the shots line up with the beats they illustrate. What is not yet handled is true visual continuity between adjacent shots: the ability for shot N to end on a specific frame and shot N+1 to begin from that same frame, so the camera or subject motion carries across the cut. That requires first-frame-last-frame conditioning in the video model, which is an active area we are exploring for a future Scenema release. In the meantime, edits between shots read as clean cuts rather than continuous motion.

Where to go next

  • New to Scenema? Start at scenema.ai and try the Vox explainer preset with any subject you are curious about.
  • Want the interactive project source for this piece? Open the full template.
  • Interested in the technical architecture behind the treatment agent and the director agent? Further writing on that is coming, along with prompt guides for the other style presets.

The FIFA piece is one worked example. Any subject you can describe in a short prompt will produce a comparable result through the same three steps.