AI video

AI Video Creator

A complete video from scratch: script, AI presenter, B-roll, voice — no footage needed.

AI Video Creator is Film To Story without a film. You save a character once (an AI picture or a photo), choose how much of the video should be the presenter on camera, and your AI writes a script that alternates presenter segments (PR01, PR02…) with B-roll pictures and short AI video clips. The presenter is animated on your PC with free models: motion from a recorded clip or a Grok Imagine clip, lips synced to the voiceover, an AI-placed camera that punches in on key words. Optional source videos can still be added and cut like in the other tools.

  • InputA topic + a saved character · optional source videos
  • OutputYouTube 16:9 and/or Reel 9:16, captions word by word
  • PresenterLivePortrait motion + MuseTalk lips, free and bundled
  • VisualsGemini / ChatGPT / Grok pictures, Grok Imagine clips, your footage
  • ControlPresenter share target, motion library, camera events

Who it is for

  • Faceless channels that want a face — one consistent presenter across every episode.
  • Educators and explainers who have the knowledge but not the footage.
  • Marketers who need a talking spokesperson for product or news videos, fast.
  • Creators who make the same video in several languages with the same character.

What you get

  • A presenter scriptBeats with presenter segments and B-roll, each with its duration, motion (calm talk, explaining with hands, leaning in…) and camera events; the validator rejects scripts that leave the presenter on camera far above your target.
  • Animated presenter segmentsEach segment is rendered on your PC from the character's reference picture: motion from your recorded clip or a Grok Imagine clip, lips synced to that beat's voice, cut exactly where it falls in the beat.
  • B-roll pictures and AI clipsPictures in a chosen style from Gemini, ChatGPT or Grok; 6–10 s clips made by Grok Imagine from the script's prompt; charts drawn from real numbers; your own footage imported where the script asked for it.
  • Voiceover, captions, musicSame engines as Film To Story; word-level captions; music with ducking.
  • Two formats, one projectReel 9:16 with hook title and YouTube 16:9, each rendered and published separately.
  • Characters librarySaved characters with reference pictures, motion clips and a default voice — reused across projects and tools.

How it works, step by step

The Film To Story pages, minus the film: a Character step instead of Import, a Script page with the presenter segments, and a B-roll card that also makes video clips.

  1. 1

    Character

    Pick a saved character or make one: a picture generated by Gemini / ChatGPT / Grok from a description (with reference pictures attached), or your own photo. Add motion clips — record yourself for 10 seconds or let Grok Imagine make a clip — and tag them (calm talk, hands, lean in…).

  2. 2

    Settings

    Topic, tone, duration and language as in Film To Story, plus the presenter share (how much of the video is on camera, e.g. 50 %) and which B-roll kinds the AI may use: AI pictures, AI video clips, real footage.

  3. 3

    Master Prompt

    The AI writes the script and a presenter list: segments with motion, camera events (zoom on a word, anchored on the face or hands) and durations that respect your share target; B-roll pictures and clips fill the rest.

  4. 4

    Script

    Import checks the schema, the share against your target and the motion variety. Each segment shows its beat, voice and motion; Render this segment animates it right away (Grok clip motion, LivePortrait, or still picture with lips), Try picture + motion previews a different combination.

  5. 5

    B-roll

    Pictures: generate with Gemini / ChatGPT / Grok or import. AI video clips: Grok Imagine makes each one from its prompt (6 or 10 s, cut to the beat at render), one by one or all at once; import your own clip instead. Real footage: the script asks, you import.

  6. 6

    Review

    Voice engine and voice per project or from the character; per-beat regeneration; the strip shows presenter segments, pictures and clips in order; music and ducking.

  7. 7

    Render

    The Render page shows which presenter segments are rendered and fresh; Render N segments makes the missing ones in one go. The video is cut without a film: presenter segments where they fall in each beat (lips on the words), clips trimmed exactly, pictures animated, captions word by word.

  8. 8

    Publish

    Per-format cards, the unified row (Render again · Download · Copy link · cloud state · Publish), thumbnail composer and the publish modal.

Features in depth

Everything runs on your PC with free models — the only accounts you need are the AI chats you already use.

One character, every episode

Saved characters with pictures, motion clips and a voice; the same face and style across projects.

Free presenter engine

LivePortrait motion + MuseTalk lip-sync + MediaPipe, bundled (≈ 4.5 GB), no API, no subscription.

Presenter share you decide

Tell the AI 30 % or 70 % on camera; the prompt computes the quota and the validator rejects scripts that overshoot.

AI video clips as B-roll

Grok Imagine text→video from the script's prompt, right on the B-roll card, in the row's aspect.

Pictures from three AIs

Gemini, ChatGPT or Grok, with reference pictures attached, in one visual style.

Camera that follows the words

Script-placed zoom events anchored on the face or hands, eased at render.

Any language

Script and voice in your language; captions word by word; Arabic RTL handled.

Hidden-browser AI

Pictures and clips are made in the background with your connected accounts, several at once.

Cross-PC sync

Characters and projects restore on another PC.

Every setting you control

Nothing is hidden behind a paywall or an "advanced" plan — these are the knobs in the app.

SettingOptionsNotes
CharacterSaved · picture (AI / photo) · motion clips · default voiceReusable across projects.
Presenter share10–90 %Target the AI must respect; validator tolerance +25 %.
B-roll kindsAI pictures · AI video clips · real footageWhich the script may use.
SegmentMotion · camera events · duration · engine (auto / Grok clip / LivePortrait / still)Per segment on the Script page.
AI clipsGrok Imagine 6 / 10 s · aspect from the row · start atCut exactly to the beat at render.
Voice · music · captions · thumbnailAs in Film To Story
FormatReel 9:16 · YouTube 16:9Each rendered separately.

What you need

  • A saved character (any picture works; a clean front-facing portrait works best)
  • A Gemini / ChatGPT / Grok account (hidden browser) for the script, pictures and clips — Grok for AI video clips
  • An ElevenLabs or Gradium key for the voiceover
  • Linked accounts to publish
  • A GPU (the presenter engine runs on your PC; CPU works but is slow)

Tips & best practices

  • Record your own 10-second motion clips once per mood — they make the presenter look like a real person, not a loop.
  • Keep the share around 40–60 %: enough face to feel human, enough B-roll to stay interesting.
  • Let Grok make the AI clips while you review the voice — they run in the background.
  • Ask for at least four motion types in the script; the validator warns when the same one repeats.
  • Render every segment before the final render — the Render page lists the missing ones and does them in one click.

Honest limits

What the tool does not do (yet), so you are not surprised.

  • Presenter quality depends on the reference picture and the motion clip; extreme head turns or hands over the face break the lips.
  • AI video clips need a Grok account with Imagine; pictures work with any of the three AIs.
  • Rendering the presenter is heavy: expect roughly real time on a GPU, much longer on CPU.
  • Arabic voiceover requires ElevenLabs.
  • Windows 10 / 11 only.

Frequently asked questions

Do I need to film anything?
No. One picture is enough; motion can come from a Grok Imagine clip. Recording yourself for ten seconds improves realism but is optional.
Is the presenter engine free?
Yes. LivePortrait, MuseTalk and MediaPipe are bundled with the app and run on your PC — no API key, no credits.
Can the presenter speak my language?
Yes — lips follow the voiceover you generate (ElevenLabs / Gradium), whatever the language.
How do I control how much the presenter is on screen?
The presenter share setting. The Master Prompt turns it into a second quota; the importer rejects scripts far above it and shows the real share.
Can I mix in real footage?
Yes — allow "real footage" in Settings, the script asks for specific clips and you import them on the B-roll card; you can also add source videos like in Video → Documentary.
Where are the AI clips made?
On the B-roll card: Generate · Grok makes a 6 or 10 s clip from the prompt in the hidden browser; it lands on the row and is cut to the beat at render.

Make your first AI video

Free download for Windows. A topic and a picture tonight, a presenter video tomorrow.

Free · No watermark · No clip limits · Runs on your PC