Video summary

this JUST became the #1 AI model...

Main summary

Key takeaways

Gaming

Storyline

  • Blood Grid (Warhammer-themed “football” board game):

    • No single narrative campaign is described.
    • Matches are driven by the teams you draft/build from various fantasy races (humans, elves, orcs, undead, dwarves, etc.).
    • The game’s injury/death consequences carry forward between turns and matches.
  • Gilded Mansion (LLM social deduction):

    • House-based social deduction inspired by Among Us / Mafia / Werewolf.
    • Players explore the mansion; the goal depends on role:
      • Impostor: avoid detection and “get away with murder”
      • Others: identify who the impostor is
  • Escape-from-Tarkov-style isometric extraction:

    • The player moves around an isometric map, loots/searches, fights enemies, and must extract.
    • If extraction fails, you lose money.

Gameplay highlights & key systems

Blood Grid

  • Turn-based board/boardgame-style combat + sports

    • Teams take turns.
    • A player can move one at a time.
    • Turns continue until a player:
      • fails an action, or
      • runs out of actions, or
      • ends their turn
  • Combat with lasting consequences

    • Fully modeled player state: stats, gear, scars, persistent injuries
    • Players can be injured and can even die, potentially permanently removing them depending on outcomes
  • Dice + RPG-style rules

    • Uses agility/strength tests, rerolls, and outcome resolution reminiscent of D&D-style mechanics
    • A large skills list; success/failure impacts movement, combat, and scoring
  • Autonomous play

    • The game can be played “pretty much on auto” after being built
  • Strategy focus

    • Different team archetypes have different optimal approaches (some want to advance the ball, others prioritize fighting)
    • Balance is tied to coaching “knobs” and how you configure the team

Subway-style demo

  • A separate smaller game concept inspired by underground subway games
  • Includes a “Matrix-style” subway bullet-time twist
  • Features music, voice lines, sound effects, and “fully voiced” production

Extraction shooter prototype (isometric Tarkov-like)

  • Autonomous exploration + extraction

    • The agent navigates, searches objects, encounters enemies, and attempts extraction
    • Extraction failure costs money
  • Manual override possible

    • Player can take control, aim a field-of-vision cone, and search with a keypress
    • Certain actions (like stopping auto-searching) can heal

Gilded Mansion (LLM-based social deduction)

  • Played with different LLM tiers

    • Cheap/low-tier models: weaker at deduction and deception
    • Mid-tier models: more competent
    • Frontier/expensive models: much stronger social deduction performance
  • Outcome example shown: impostors winning

  • Leaderboard/ELO tracking

    • Current top ELO mentioned: DeepSeek V3.2
    • At the time of recording, Grok 4.5 follows

Strategies / key tips emphasized

  • Change your workflow to use Fable 5.1 effectively

    • It’s not treated as an incremental upgrade; it requires rethinking how you build and iterate
  • Effort/reasoning settings

    • Recommendation: start lower
      • Low: quick checks and many build steps
      • Medium: deeper thinking/research
      • High: tasks that truly require thorough “execute after thinking everything through”
    • Avoid reliance on max/extra/ultra
      • The creator claims they were not used
      • Some users report issues (e.g., too many sub-agents with extra/max)
  • Don’t switch modes in the middle of work (with other models)

    • With Claude 5 + Fable 5.1, context loss can happen on older setups when switching reasoning levels
    • Their workflow uses a shortcut to switch on the fly without losing context
  • Design safeguards against “rule-breaking” unintended consequences

    • When adding a rule, anticipate downstream knock-on effects
    • Example: if an AI learns an overly dominant behavior (e.g., sending everyone to the ball), it can become near-guaranteed, reducing the value of skills/defense and ruining tension/balance
  • Use headless simulation for balance

    • Build a headless simulator to run thousands/tens of thousands of games quickly
    • Use a “Balance Bench” that runs multiple instances and applies a Bradley-Terry-style strength rating approach to evaluate odds/tiers
  • LLM benchmark design tip

    • One-turn-at-a-time “play-by-play” decision-making was too expensive/unwieldy for LLMs
    • They replaced it with a play-calling system:
      • LLM selects a preset formation/play approach per turn
      • The game resolves the play, reducing token/call cost and enabling feasible benchmarking
  • Iterate toward “next obvious feature”

    • Example response to “biggest improvement”: adding a replay log

Notable sources / gamers featured (mentioned at the end)

  • Bjan Bowen

  • Gamers/models mentioned in the Gilded Mansion tiers / leaderboard discussion:

    • DeepSeek V3.2
    • Grok 4.5
    • Claude (Claude 5 referenced generally)
    • GPT-4o (mentioned indirectly via “GPT minis” / “GBT minis”)
    • Mr. Medium, Claw, Haiku, Gemini 2.5 Flash, Gemini 2.5 (as mentioned)
    • Opus 4.5 (mentioned as “Clot Opus 4.5”)
    • “Grok 5.6” (mentioned)
    • “six soul” (mentioned as part of the model tier list; exact spoken name unclear)

Original video