Story video

Film To Story

A full film → one viral recap video with AI voiceover, in Reel and YouTube formats.

Film To Story turns a whole film into one short story video the way recap channels do — but automated. The app analyzes every shot, your AI writes the script beat by beat with the shot IDs to show, an ElevenLabs or Gradium voice reads it, and the renderer cuts the matching shots to the narration with captions, a hook title, music and a copyright shield. Eight pages, one per step, and both formats from the same script.

  • InputFilm file, direct link, YouTube / video page, magnet
  • OutputReel 1080×1920 and/or YouTube 1920×1080 + thumbnail / cover
  • VoiceElevenLabs or Gradium, beat by beat
  • AIGemini / ChatGPT / any AI IDE for the Master Prompt
  • ShieldZoom · mirror (except text) · grade · vignette · short cuts

Who it is for

  • Movie recap and "explained" channels on YouTube, TikTok and Reels.
  • Faceless storytelling channels that want original stories told over film footage.
  • Creators in Arabic, French or English who need voiceover and captions in their language.
  • Anyone who has spent an evening scrubbing through a film for the right shot.

What you get

  • A voiced story videoNarration written by your AI, read by ElevenLabs or Gradium, cut to the exact shots with the rhythm marks the script asked for.
  • Two formats from one scriptReel 9:16 (4:3 picture on a blurred background, hook title, captions) and YouTube 16:9 (clean full frame, captions). Both are kept; render either again any time.
  • Thumbnail & coverAn AI picture generated in the hidden browser (or your own) with your text composed on top — font, size, outline, shadow, gradient, highlight colour — in 1280×720 and 1080×1920.
  • Publish textsYouTube title, description, hashtags and tags; TikTok caption; Facebook / Instagram caption — written by the AI, editable in the publish modal.
  • A shot map you can browseEvery shot with its ID, keyframe, description, dialogue, faces and optional on-screen text; title cards and credits excluded automatically.
  • Cloud snapshotScript, voiceover, settings, poster and renders' cloud mirrors are synced — open the story on another PC and continue; re-download the film with the same link.

How it works, step by step

Eight pages in order. Each one can be revisited; the Render page keeps both formats.

  1. 1

    Import

    From your PC, a direct link, a YouTube / video page link (yt-dlp) or a magnet link (aria2). The download runs in the background with progress on the project card; a TMDB poster becomes the project thumbnail (change it from the header).

  2. 2

    Analyze

    Shot detection, keyframes, dialogue transcription and visual descriptions build the shot map: each shot gets an ID (S0001…), a description, its dialogue, a face count. Studio logos, opening names and the whole credits roll are excluded by a credits filter. Optional Read on-screen text (OCR) adds a TEXT column (signs, letters, chapter cards) — off by default for films.

    • A 2-hour film → roughly 1,000–1,500 usable shots.
    • Analysis runs in the background; you can prepare settings meanwhile.
    • Browse the map afterwards: filter, search, see hidden shots.
  3. 3

    Settings

    Target duration (buckets or custom minutes), story and voice language, and the story style:

    • Faithful recap — the film's story, compressed.
    • Fast & dramatic — high-tempo, punchy.
    • Twist first — open on the twist, then explain.
    • Villain's POV — retold from the antagonist.
    • Mystery — withhold, tease, reveal.
    • Viral fiction — a brand-new story built only from the film's shots that never names the film, its characters or actors; genre picker: horror, true-crime, revenge, betrayal, karma, time loop, or auto.
  4. 4

    Master Prompt

    Copy the prompt (or download the .md) into your AI IDE or chat — or run it with Gemini / ChatGPT in the hidden browser. The AI writes story.json: beats with narration and shot IDs (plus alternates), a hook title, YouTube / TikTok / Meta metadata, a thumbnail prompt, a Suno music prompt, pronunciations, and rhythm marks per beat.

    The prompt lists the project files an AI IDE may open — shot map, transcript, keyframes — so it can look at a shot before picking it.

  5. 5

    Story

    Import the JSON. It is validated strictly against the shot map (IDs must exist, excluded credit shots are stripped with a warning, word budget, schema) and every auto-fix is explained. For Viral fiction any leaked film name is flagged. A Script check lists weak beats (shots unrelated to the narration, one shot for a long narration…) with an "Ask the AI to fix this beat" prompt — paste the tiny JSON back and only that beat changes.

  6. 6

    Review

    Pick the voice engine (ElevenLabs or Gradium), model and voice — library search, filters, previews — estimate credits, then generate the voiceover beat by beat. Listen, edit a narration and regenerate just that beat, reorder or swap shots from the shot strip, set per-beat rhythm marks (pause after, fade to black, echo, slow motion), add a music track (AI Suno prompt or your own file) with automatic ducking under speech.

  7. 7

    Render

    Choose Reel or YouTube, the copyright shield preset (zoom, mirror except shots with readable text, colour grade, vignette, short cuts), captions (font, size, colours, position), hook title (font, box, colours), fade-to-black length, and the three rhythm switches (Transitions, Audio effects, Slow motion). The timeline preview shows which shot plays when. Render again any time — only that format is replaced.

  8. 8

    Publish

    One card per format: preview, then the same row as a Video → Shorts clip — Render again, Download, Copy link, cloud state, Publish — and the thumbnail / cover section: Generate with Gemini / ChatGPT, import your own, compose the text, Upload to cloud. The publish modal prefills YouTube, TikTok and Meta fields; the thumbnail is set on YouTube, sent to Facebook and used as the Instagram Reel cover.

Features in depth

What makes a recap look and sound like a channel with an editor.

Six story styles

From a faithful recap to a Viral fiction that never names the film. The prompt enforces the style; the validator checks the leaks.

Copyright shield

Zoom, mirror (never on shots with readable text — detected per shot), colour grade, vignette and short cuts, chosen as presets. The original audio is never used.

Two voice engines

ElevenLabs (cloud key pool rotated by credits, dozens of languages incl. Arabic) or Gradium (local vault, 45,000 free characters a month, word timings).

Rhythm marks

Pause after a beat, fade to black into it, an echo on the last word that cuts to black just before the word is spoken so it rings out on black, and slow motion — each toggleable at render time.

Thumbnail composer

AI picture + your text: Impact or any font, ≤ 3 lines auto-shrunk, outline, shadow, gradient band, last-word highlight, position and alignment, uppercase toggle.

Both formats

Reel with blurred background and hook title, YouTube full-frame — from the same script and voiceover. Both kept; either re-rendered alone.

Script QA

Advisory notes per beat and a tiny fix prompt for the AI; only the fixed beat is replaced, other voices stay valid.

Cross-PC sync

State and project bundle are synced after each change; restore on another PC, re-download the film, keep rendering.

Languages

Story language and voice language independent; captions follow the voice, right-to-left for Arabic.

Every setting you control

Nothing is hidden behind a paywall or an "advanced" plan — these are the knobs in the app.

SettingOptionsNotes
Target durationBuckets (e.g. 3–5, 5–8, 8–12 min) · custom minutesWord budget derived from the duration.
Story styleFaithful · Fast & dramatic · Twist first · Villain's POV · Mystery · Viral fiction (+ genre)Viral fiction forbids naming the film.
LanguagesStory language · voice languageAny Whisper / voice-engine language.
OCR (on-screen text)Off (default for films) / OnAdds a TEXT column to the shot map.
VoiceElevenLabs or Gradium · model · voice · speedCredits estimate before generating.
MusicAI prompt (Suno) or your file · ducking on/offDips to ~35 % under speech.
Output formatReel 9:16 · YouTube 16:9Both can exist side by side.
Copyright shieldOff · Light · Standard · AggressiveMirror skipped on readable-text shots.
CaptionsOn/Off · font · size · colours · positionWindows or downloaded Google fonts.
Hook titleFont · box · colours · positionReel only.
RhythmTransitions · Audio effects · Slow motion switches · fade 0.2–2 sIgnore the script's marks when off.
EncoderAuto (NVENC → CPU)

What you need

  • The film file or a link (direct, video page, magnet)
  • A Gemini / ChatGPT account (hidden browser) or an AI IDE / chat for the Master Prompt
  • An ElevenLabs or Gradium API key for the voiceover (free tiers work)
  • Linked YouTube / TikTok / Facebook / Instagram accounts to publish
  • A GPU is strongly recommended for analysis of full films

Tips & best practices

  • Turn OCR on for films with chapter cards or letters the story depends on.
  • Estimate credits before generating: a 6-minute script is roughly 5,000–7,000 characters.
  • Use alternates: the script gives each beat alternate shots; the shot strip lets you swap in one click.
  • Keep two or three echo endings per video — the strongest beat ending, not every beat.
  • Compose the thumbnail early; it can be prepared before the render finishes.
  • Upload the thumbnail to the cloud before publishing, or the post goes out without it (the app warns you).

Honest limits

What the tool does not do (yet), so you are not surprised.

  • A full film analysis is heavy on CPU-only machines (hours); an NVIDIA GPU brings it down to minutes.
  • Voiceover needs a voice-engine key; there is no built-in offline voice.
  • Arabic voiceover requires ElevenLabs (Gradium does not offer Arabic yet).
  • The copyright shield reduces automatic matches; it is not a legal guarantee — fair use depends on your country and the platform.
  • The film file itself is never synced to the cloud; re-download it on a second PC.

Frequently asked questions

Does it use the film's original audio?
Never. The video carries your voiceover, optional music and audio effects only.
How long does a 2-hour film take?
Import depends on your connection; analysis takes minutes on a modern NVIDIA GPU and considerably longer on CPU; voiceover a minute or two; render a few minutes per format.
Can I write the script myself?
Yes — edit any narration on the Review page and regenerate only that beat, or paste your own story.json that follows the schema.
What is Viral fiction?
A story style where the AI invents a new plot (horror, true-crime, revenge…) told only with the film's shots, never naming the film, its characters or actors. The validator flags any leaked name.
Reel and YouTube — do I need two scripts?
No. One script and one voiceover; the renderer builds the 9:16 Reel and the 16:9 YouTube video from them.
Where do the thumbnail pictures come from?
From your connected Gemini / ChatGPT in the hidden browser using the prompt the script proposes, or from any picture you import. Your text is composed on top locally.
Can I continue on another PC?
Yes. The story's state (script, voices, settings, poster) is synced to your cloud space; restore it on the other PC and re-download the film with the same link.

Make your first recap

Free download for Windows. Import a film tonight, publish a Reel tomorrow.

Free · No watermark · No clip limits · Runs on your PC