Video summary
this JUST became the #1 AI model...
Main summary
Key takeaways
Storyline
-
Blood Grid (Warhammer-themed “football” board game):
- No single narrative campaign is described.
- Matches are driven by the teams you draft/build from various fantasy races (humans, elves, orcs, undead, dwarves, etc.).
- The game’s injury/death consequences carry forward between turns and matches.
-
Gilded Mansion (LLM social deduction):
- House-based social deduction inspired by Among Us / Mafia / Werewolf.
- Players explore the mansion; the goal depends on role:
- Impostor: avoid detection and “get away with murder”
- Others: identify who the impostor is
-
Escape-from-Tarkov-style isometric extraction:
- The player moves around an isometric map, loots/searches, fights enemies, and must extract.
- If extraction fails, you lose money.
Gameplay highlights & key systems
Blood Grid
-
Turn-based board/boardgame-style combat + sports
- Teams take turns.
- A player can move one at a time.
- Turns continue until a player:
- fails an action, or
- runs out of actions, or
- ends their turn
-
Combat with lasting consequences
- Fully modeled player state: stats, gear, scars, persistent injuries
- Players can be injured and can even die, potentially permanently removing them depending on outcomes
-
Dice + RPG-style rules
- Uses agility/strength tests, rerolls, and outcome resolution reminiscent of D&D-style mechanics
- A large skills list; success/failure impacts movement, combat, and scoring
-
Autonomous play
- The game can be played “pretty much on auto” after being built
-
Strategy focus
- Different team archetypes have different optimal approaches (some want to advance the ball, others prioritize fighting)
- Balance is tied to coaching “knobs” and how you configure the team
Subway-style demo
- A separate smaller game concept inspired by underground subway games
- Includes a “Matrix-style” subway bullet-time twist
- Features music, voice lines, sound effects, and “fully voiced” production
Extraction shooter prototype (isometric Tarkov-like)
-
Autonomous exploration + extraction
- The agent navigates, searches objects, encounters enemies, and attempts extraction
- Extraction failure costs money
-
Manual override possible
- Player can take control, aim a field-of-vision cone, and search with a keypress
- Certain actions (like stopping auto-searching) can heal
Gilded Mansion (LLM-based social deduction)
-
Played with different LLM tiers
- Cheap/low-tier models: weaker at deduction and deception
- Mid-tier models: more competent
- Frontier/expensive models: much stronger social deduction performance
-
Outcome example shown: impostors winning
-
Leaderboard/ELO tracking
- Current top ELO mentioned: DeepSeek V3.2
- At the time of recording, Grok 4.5 follows
Strategies / key tips emphasized
-
Change your workflow to use Fable 5.1 effectively
- It’s not treated as an incremental upgrade; it requires rethinking how you build and iterate
-
Effort/reasoning settings
- Recommendation: start lower
- Low: quick checks and many build steps
- Medium: deeper thinking/research
- High: tasks that truly require thorough “execute after thinking everything through”
- Avoid reliance on max/extra/ultra
- The creator claims they were not used
- Some users report issues (e.g., too many sub-agents with extra/max)
- Recommendation: start lower
-
Don’t switch modes in the middle of work (with other models)
- With Claude 5 + Fable 5.1, context loss can happen on older setups when switching reasoning levels
- Their workflow uses a shortcut to switch on the fly without losing context
-
Design safeguards against “rule-breaking” unintended consequences
- When adding a rule, anticipate downstream knock-on effects
- Example: if an AI learns an overly dominant behavior (e.g., sending everyone to the ball), it can become near-guaranteed, reducing the value of skills/defense and ruining tension/balance
-
Use headless simulation for balance
- Build a headless simulator to run thousands/tens of thousands of games quickly
- Use a “Balance Bench” that runs multiple instances and applies a Bradley-Terry-style strength rating approach to evaluate odds/tiers
-
LLM benchmark design tip
- One-turn-at-a-time “play-by-play” decision-making was too expensive/unwieldy for LLMs
- They replaced it with a play-calling system:
- LLM selects a preset formation/play approach per turn
- The game resolves the play, reducing token/call cost and enabling feasible benchmarking
-
Iterate toward “next obvious feature”
- Example response to “biggest improvement”: adding a replay log
Notable sources / gamers featured (mentioned at the end)
-
Bjan Bowen
-
Gamers/models mentioned in the Gilded Mansion tiers / leaderboard discussion:
- DeepSeek V3.2
- Grok 4.5
- Claude (Claude 5 referenced generally)
- GPT-4o (mentioned indirectly via “GPT minis” / “GBT minis”)
- Mr. Medium, Claw, Haiku, Gemini 2.5 Flash, Gemini 2.5 (as mentioned)
- Opus 4.5 (mentioned as “Clot Opus 4.5”)
- “Grok 5.6” (mentioned)
- “six soul” (mentioned as part of the model tier list; exact spoken name unclear)