Video summary

I Tried Every AI Video Generator So You Don't Have To

Main summary

Key takeaways

Product Review

Product(s) reviewed

The video compares four AI video generators:

  • Cense / “Seance”
  • Cling
  • Grock / “Croc” (referred to as Grock and Croc)
  • Google Omni

An “all-in-one” sponsored access platform, Hicksfield, is mentioned to try multiple models.


Key comparison setup (what was tested)

  • Blind test: multiple outputs generated with the highest settings available.
  • Resolution/settings:
    • Seance: 4K, high bitrate
    • Cling: 4K, high bitrate
    • Grock: 720p
    • Google Omni: native 720p, downloadable/upscalable to 1080p
  • Watermark fairness: the host adds a watermark remover because Google Omni applies a watermark even on paid plans; only Ultra removes it.

Test results (rankings & notable outcomes)

1) Reference image → POV ice cream scene (vegan “dirty Oreo”)

  • Gold (best overall): Video #3 (Seance) — best audio naturalness; scoop looks good.
  • Silver: Video #4 (Google Omni) — better audio than #2.
  • Bronze: Video #2 (Cling) — host had issues with pointing/frame “physics.”
  • Worst/Disliked: Video #1 (Grock) — physics/control issues.

Host’s notes:

  • Host says video #1 had weird physics.
  • Host says the audio is best in #3.

Video ordering (host’s pick): Seance (1st) > Google Omni (2nd) > Cling (3rd) > Grock (4th).


2) Text-to-video: pirate ship “deep drop” prompt

  • Grock: cannot do text-to-video with the tested setup (image input only), so it’s automatically last.
  • Gold (best): Video #4 — better physics/audio; lacking the “deep drop.”
  • Silver: Video #1 — had everything but disliked audio similarity between voices.
  • Bronze: Video #2 — Grock excluded; host says it “wasn’t great.”

Overall feel: strong text-to-video performance from Seance / Cling / Omni, while Grock can’t run text-to-video here.


3) Motion transfer: copy dance/movement from one person to another

  • Grock: again can’t do this (no motion control), so it’s not tested.
  • Host ratings:
    • Gold: Video #1 (fun result; AI replaced some background music)
    • Silver: Video #4 (also “really good”)
    • Bronze: Video #2 (butchered face; weird movement)

4) Dialogue + lip sync: “stay 1 age forever” conversation (fixed camera)

  • Google Omni: couldn’t complete the video due to censorship (“censorship apparently”).
  • Host ratings (overall):
    • Gold: 1st video
    • Silver: 3rd video
    • Bronze: 2nd video
  • Additional breakdown:
    • Best dialogue: Grock
    • Dialogue silver: Seance
    • Dialogue bronze: Cling

5) Complex prompts: multi-shot sequence / interactions

  • Host emphasizes continuity and shot-handling differences.
  • Ratings:
    • Gold: last video (best “did everything right”)
    • Silver: 2nd video (lip sync/voice fixable with an “extra cut”)
    • Bronze: 3rd video (lip sync off)
    • 1st video: “really bad”

For “complex prompt” (follow-up):

  • Seance = Gold
  • Grock = Silver
  • Cling = Bronze

6) References: multi-reference crime scene (9 references)

  • Ratings:
    • Gold (best): Video #2
    • Silver: Video #3
    • Worst picks: Video #1 and #4
      • Host wouldn’t use #1 and #4’s shots; #4 is only slightly better due to fade issues.

“For the references” summary:

  • Seance = Gold
  • Cling = Silver
  • Google Omni = Bronze (close)

7) Cinematic: dragon + girl on a cliff (2 angles)

  • Ratings:
    • Gold: Video #2
    • Silver: Video #1
  • For the remaining outputs, the host refuses to rank Bronze:
    • Video #3: dragon scaling too off
    • Video #4: “too fake/cinematic”

Final blind-test tally / overall

  • Host believes Seance (“Cense”) most often came out Gold.
  • Second place is likely Cling or Grock (host says either).
  • Google Omni is used more selectively.

Strengths & weaknesses (explicit product breakdown)

1) Grock / Croc

Pros

  • Speed: “super fast” generation.
  • Lip sync & emotion: host calls it strong.
  • Less censorship: can generate “sensitive” content more readily.
  • Can do image-based generation (example: inserting a Tom Cruise-like image).

Cons

  • Only 720p output (behind others that do 4K).
  • Expensive: host compares credits and suggests Seance/Cling are better value.
  • Glitches/morphing (example: gun appears/disappears incorrectly).
  • Poor multi-shot scene continuity (characters “teleporting” / inconsistent positions).
  • Not great for complex multi-scene prompts.

When the host would use Grock

  • When you need fast generation, or to experiment with content likely to be censored elsewhere—especially with image/famous-person style inputs.

2) Seance / “Cense” (“Seans/Seedance” in subtitles)

Pros

  • Best for complex tasks overall:
    • references,
    • multi-shot/composite scenes,
    • dialogue/lip sync (often ranked highly),
    • action scenes that follow prompts well.
  • Strong motion transfer (surprisingly tops comparisons).
  • “Consistency” and expressive acting/audio praised across multiple examples.
  • Host expects Cense 2.5 Pro to improve further (expects it to be “best”).

Cons

  • Expensive (one of the most costly generators).
  • Long waiting times.
  • Consistency drops in long/complex prompts:
    • examples include clones/duplicated characters (“Why is this guy’s clone here?”)
  • Sometimes fails with certain graphic/violent content (host cites issues including Google’s failure and implies guardrails can matter).

3) Cling

Pros

  • Inexpensive budget alternative (explicitly positioned vs Seance).
  • 4K and 60 fps (host calls it “crazy”).
  • Supports many references (host: up to ~7 images + possibly video/audio).
  • Often cited as ~80% of Seance quality at ~1/4 the cost (host’s claim).
  • Good camera adherence (e.g., dolly zoom/tracking-style prompts).
  • Strong for simple shots and less complex scenes.

Cons

  • Lip sync degrades as videos get longer (worse after ~10–15 seconds).
  • Voice quality not as strong as Grock/Google/Seance.
  • Long waiting times (host claims sometimes worse than Seance).
  • Not ideal when there’s lots of dialogue.

When the host would use Cling

  • Simple scenes, cost-sensitive production, and quick iterations when Seance is too expensive.

4) Google Omni

Pros

  • Extremely cheap.
  • Strong for video-to-video editing / motion transfer workflows (e.g., outfit change / motion transfer pipelines).
  • In some blind tests, it ranked high—especially audio in test #1.

Cons

  • Not a fair like-for-like comparison (host calls it different from the others).
  • Watermark policy:
    • watermark on free/plus/pro,
    • removed only on Ultra,
    • host finds it annoying/inconsistent.
  • Censorship can block outputs (failed in the dialogue test).
  • Quality issues:
    • “plasticskin” look in examples,
    • host couldn’t reproduce promised “insane” edits from demos.
  • Host’s overall view: quality isn’t there yet compared to Seance/Cling/Grock.

When the host would use Google Omni

  • Mainly for budget video-to-video tasks and frequent editing, despite watermark/censorship/quality issues.

Explicit comparisons / patterns the host emphasizes

  • Seance wins most often for:
    • complex prompts,
    • references,
    • cinematic/multi-shot correctness,
    • action scene adherence,
    • motion transfer.
  • Grock wins on:
    • speed,
    • emotion + lip sync (in some examples),
    • lower censorship / bypassing guardrails.
  • Cling is the value pick:
    • strong for simple shots, camera adherence, and references,
    • better cost/quality ratio than Seance,
    • but lip sync/voice degrade on longer, dialogue-heavy clips.
  • Google Omni is selective:
    • cheap and good for video-to-video editing,
    • but inconsistent execution plus watermark/censorship hassles.

Pros / cons summary (overall)

  • Best overall for complex/serious output: Seance
  • Best speed + lower censorship + emotion/lip sync: Grock
  • Best budget for 4K/simple scenes: Cling
  • Best for budget video-to-video edits (selectively): Google Omni

Overall verdict / recommendation (host’s conclusion)

  • Use Seance for most complex things and high-quality everyday work (if budget/time allow).
  • Use Cling if you don’t have the budget or need quick simple scenes—accepting weaker lip sync/voice for longer dialogue-heavy videos.
  • Use Grock when you need fast generation or content likely to be censored elsewhere.
  • Use Google Omni only sometimes, mainly for budget video-to-video editing; quality and censorship/watermark issues can limit it.

Unique points mentioned (deduplicated list)

  1. Blind test comparing Seance vs Cling vs Grock vs Google Omni.
  2. Highest settings used; differing output resolutions (4K vs 720p).
  3. Watermark fairness: Google Omni watermark removed only on Ultra; watermark remover used.
  4. Test #1 (ice cream POV): Seance best audio/natural voice; Grock physics off.
  5. Test #2 (pirate ship text-to-video): Grock can’t text-to-video (image-only).
  6. Test #3 (motion transfer dance): Grock lacks motion control.
  7. Test #4 (dialogue/lip sync): Google Omni censored; Seance/others ranked; Grock best dialogue.
  8. Test #5 (complex multi-shot): Seance wins complex prompts; lip sync may require extra cuts.
  9. Test #6 (9-reference crime scene): Seance gold; Cling silver; Google bronze.
  10. Test #7 (cinematic dragon): Seance top; others flawed (scale/fakeness).
  11. Grock strengths: speed, emotion, lip sync, less censorship.
  12. Grock cons: 720p only, expensive, glitches/morphing, bad scene continuity (“teleporting”).
  13. Seance strengths: references, consistency/expressiveness, action scene adherence, motion transfer.
  14. Seance cons: expensive, slow waiting time, consistency drops in long complex prompts (clones/duplication).
  15. Cling strengths: cheap, 4K 60fps, many references, good camera adherence, strong simple scenes.
  16. Cling cons: lip sync worsens after ~10 seconds, voice quality lower, long waiting times.
  17. Google Omni strengths: cheap, good for video-to-video editing (motion transfer/editing).
  18. Google Omni cons: watermark policy, censorship blocks outputs, “plasticky” skin, not matching demo-level quality.
  19. Host suggests Hicksfield as an all-in-one platform to test models.

Speakers

  • Single primary speaker/host (no other distinct interview voice indicated in subtitles).

Original video