Video summary

RIP Paid Tools: Make LONG AI Videos With Consistency (Free!)

Main summary

Key takeaways

Technology

Tech/Product Focus Summary (AI long-form video workflow)

The video argues that “long-form AI cinematic story videos” are usually ruined by random, inconsistent tool usage (“AI slop”). It then presents a full end-to-end workflow using only free tools to maintain character consistency across frames.


1) Story + timeline generation (free LLM)

  • Tool: Google Gemini
  • Purpose: Generate the core narrative and a structured timeline.

Process

  • Use a “master prompts” document (hosted via Discord, also linked in the video description).
  • Key instruction: copy the first master prompt and paste it into Gemini.
  • Notes:
    • The prompt specifies duration and cinematic style.
    • The creator manually reviews and tweaks the narrative to improve the odds of going viral.

2) Visual planning: chronologically ordered scene prompts

  • Tool: Gemini, using the second master prompt.
  • Output requirement: Gemini is instructed to generate at least ~18 unique visual prompts (adjustable to video length).
  • Constraint: Prompts should be returned in chronological order.

3) Character creation to enforce consistency

The workflow references “Nano Banana” as an inspiration/realism reference, then uses:

Main character + anchor prompts

  • Primary generator: Google Flow (free text-to-image / image generation)
  • Problem highlighted: Default images often fail to match consistent cinematic style or character geometry.
  • Solution:
    • Copy the third master prompt into Gemini to generate highly detailed text-to-image prompts for anchor characters.
    • Create a main protagonist full-face character (the video claims this improves “trust” versus masked avatars).
    • Generate a three-side character sheet (via another prompt).

Consistency enforcement across scenes

  • Use the chosen character name and reuse it as a master reference in later prompts.
  • For each scene:
    • Upload the character-aware prompt into the image/video pipeline.
    • Inject character identity using the @ symbol.
    • Apply manual QC: if anything looks wrong, open the prompt and aggressively regenerate.

4) Animate static images (tool selection by motion complexity)

The workflow generates cinematic motion clips by splitting prompts into motion tiers and selecting tools based on limits/quality.

Upscaling

  • Before downloading, scenes should be natively upscaled.

Camera motion → animation prompts (Gemini)

  • Tool: Gemini
  • Inputs:
    • the final master prompt
    • plus a separate camera movements document uploaded alongside the prompt
  • Gemini outputs precise animation prompts tailored to the camera movement data.
  • The creator manually sorts animation prompts into:
    • High motion
    • Medium motion

High motion (quality-focused, with daily credits/limits)

Option A: Google Omni Flash

  • How (Google Flow):
    • hover image → three dots → Animate
    • select Omni model
    • paste high-motion prompt
  • Cost: 15 credits per generation
  • Claimed limit: ~3–5 clips/day

Option B: Dolla AI

  • Uses “seedance fast model” via an engineered prompt.
  • Claim: “free seedance two generations”
  • Notes:
    • Mentions watermark concern, but claims it will be handled later.
  • Limit: ~5–6 clips/day

Strategy

  • Combining Omni Flash + Dolla AI yields ~10 high-motion clips, enough for 10–15 minute videos.

Medium motion (unlimited)

  • Tool: Meta AI
  • Process: Create → upload image → paste medium-motion prompt.
  • Claim: less motion means the engine is less likely to lag or under-animate.
  • Most of the B-roll is generated inside Meta.

Watermark-free downloads via browser dev tools

To download medium-motion Meta videos without watermarks:

  1. click the video
  2. right-click near the border → inspect element
  3. select the video container in the code
  4. copy the direct URL from the right panel
  5. open the direct URL in a new tab to obtain a watermark-free file

5) Voiceover (text-to-speech with controlled settings)

  • Tool: Google AI Studio
  • Model: Gemini 2.5 Pro voice model
  • Ensures single-speaker audio mode (voice doesn’t randomly change).

Process

  • Use an “initial instruction prompt” at the top.
  • Paste the full generated story text below.
  • Preview and choose a voice character/tone.
  • Download the generated audio file.

6) Background music synced to the story

  • Tool: Gemini’s integrated music creation tool
  • Prompts Gemini to generate background music matching the story’s emotional arc.
  • Download the music track as a high-quality file.

7) Edit/assemble

  • Tool: CapCut
  • Used to assemble the long-form YouTube video.
  • Suggested watermark handling:
    • cover small AI-tool watermarks with the creator’s channel logo to protect brand authority without quality loss.

8) Bonus for complex action scenes (higher-end video model)

If motion/action is too difficult for the other tools, use:

  • Tool: Google Vids
  • Described as official Google access to a top VO3.1 model.

Claims

  • Free daily creations: about 10–12 per day
  • Login uses a standard Google account.

Purpose

  • Generate difficult high-action shots manually and insert them into the automated timeline as a “safety net.”

Main speakers/sources mentioned (at end)

  • Speaker/creator: “I” (the YouTube narrator presenting a step-by-step workflow)

Tools/platforms used in the workflow

  • Google Gemini
  • Discord (master prompt document)
  • Google Flow (image generation + animation)
  • Dolla AI (high-motion video generation)
  • Meta AI (medium-motion video generation)
  • Meta AI / browser Inspect Element (watermark-free downloads)
  • Google AI Studio (text-to-speech)
  • CapCut (editing)
  • Google Vids (bonus high-action video generation)

Original video