Guide · 2 October 2026 · 7 min read
How to make a faceless YouTube video with AI, step by step
The seven steps from an idea to a published long-form faceless video, what each one needs, and where the time usually goes.
A faceless YouTube video is one with no presenter on camera: a narrated script over illustrations, photos, animation or screen footage. You can make one in seven steps. Done by hand, a 10-minute video takes days. With an AI video generator, most of the steps are automated and you keep the decisions that matter.
1. Pick one idea and one format
Start with a single, specific idea that a stranger would click: a question, a ranking, a story. Then choose a format that suits it: an explainer, a documentary, an escalating ladder ("every level of…"), a business breakdown, a story with one recurring character. The format decides how the script is structured and how it looks, so it is worth choosing deliberately.
2. Write the script in sections, not in one go
Long videos fail when they are written as one long block. Outline first (hook, a handful of sections, an ending), then write each section against the outline. A 10-minute video is roughly 1,400 to 1,600 spoken words. Open with the payoff, not a greeting, and keep a heading for each section: they become your chapters.
For videos about real events, work from a source you can check, and keep every date, number and name tied to it. Invented specifics are the fastest way to lose a viewer who knows the subject.
3. Generate the narration
Pick a voice that fits the topic and set a pace. Documentaries sit around 140 to 155 words a minute; faster narration reads as rushed. Pauses between sections give the viewer a place to breathe, and they are where chapter cards and music changes go. Spell out names and acronyms phonetically in the text sent to the voice if it mispronounces them, but keep the real spelling in the captions.
4. Time the words, then cut scenes
Transcribe the finished narration to get a timestamp for every word. Cut the video into scenes on sentence boundaries, usually 5 to 10 seconds each for explainers and shorter for fast formats. Word timestamps are also what make captions land on the right word.
5. Make the pictures, and keep them varied
Each scene needs a picture. Illustrated formats use an image model with a fixed visual style; documentary formats should use real, properly licensed photos, because image models cannot draw real people accurately. Whatever you use, make sure consecutive scenes show different things: a wide shot, a detail, an object, a different room. A video that repeats one picture feels slow even when the script is good.
If you use a recurring character, give the image model a reference of them rather than a text description. Text alone drifts, and the character looks different in every scene.
6. Review before you render
Look through the storyboard before the final render. Replace a scene, redraw an image, fix a line. This is the cheapest point to catch mistakes: nothing expensive has run yet.
7. Package it and publish
A video is a title, a thumbnail and a description as much as it is a film. Write several titles and thumbnail options and pick one, add chapters from your section headings, and schedule the upload for when your audience is online. Mark the video as containing altered or synthetic content where YouTube asks, and credit any photos that require it.
Tellavid does steps 2 to 7 from one sentence: it writes the script, narrates it, times the words, draws or finds the pictures, pauses for your review, renders with captions and music, makes six thumbnail options and uploads the finished video to your channel on your schedule.
See it applied to a real video.
Every example in the showcase includes the sentence that started it and every stage in between.
Open the showcase