Video summary

Club TWiT: AI User Group #18 - Hacking Your Workflow

Main summary

Key takeaways

Technology

Key technological concepts & products discussed

1) Local LLM orchestration on Mac + “outsourcing” models

  • Darren describes a main Mac “work machine” paired with a NVIDIA DGX Spark setup where models are managed via an arbiter/queue system.
  • The Spark UI reflects queued/asynchronous execution: only one model actively uses heavy resources at a time.
  • He runs local models via Ollama, and mentions switching from “Llama” to Ollama for practical reasons (e.g., model switching via CLI), while noting a preference for Llama branding/feel.
  • They discuss model selection and performance tuning, including native vs other runtime speed comparisons.

2) Realtime / near-realtime voice generation + “Kakoro”

  • The group demos Kakoro, described as faster and closer to real-time voice generation than alternatives.
  • Voice training is framed as:
    • Overnight training / fine-tuning of a voice using transcripts to approximate a real speaker/voice.
    • A tradeoff: not as high quality as “Quinn”, but favored for speed and practicality.
  • They play an example transcript-based voice attempt and discuss limitations and latency.

3) “One-sheets” and advertiser competitive analysis via Hermes skills

  • A workflow called “one sheets” (backgrounders for potential advertisers) is automated using an AI “skill” (likely Claude-based, older model).
  • Inputs include:
    • Internal spreadsheets/manual information
    • Targeting constraints such as:
      • Market/competitor targeting
      • Recommended Twitter audience segments
    • Generated materials such as talking points for advertisers plus info for internal vetting
  • Outputs include:
    • Credibility checks (e.g., whether the advertiser competes with existing ones)
    • Historical and current advertiser lists and comparisons to prior/current advertisers
    • Where the company advertises (pods, creator/YouTube, other channels) and community reputation
  • They mention policy-style exclusions:
    • Certain categories (e.g., crypto / regulated categories) may be automatically rejected (“no one-sheet” for those types).

4) Show prep automation with Obsidian + MCP + Hermes “skills”

  • They use Obsidian as a home for generated notes and dossiers.
  • They maintain an “AI folder” containing outputs such as:
    • show prep
    • research
    • formatting artifacts
  • A show-prep skill is built from prior work and iterated on:
    • Example: research a guest (e.g., Jeffrey Cannell) to produce angles/questions and show-relevant context
    • Addressed concern: hallucinations
      • mitigation depends on available context and source signals
  • They also discuss a rule book / organization policy so agents can consistently place and name generated content.

5) Financial/fitness summaries via AI + data “bridges” (SimpleFin)

  • Darren describes:
    • Daily financial summaries pushed into Obsidian
    • Using SimpleFin as a “bridge” so AI can access account data through a read-only/API-style integration
  • Benefit: avoids direct login issues some services have (e.g., Monarch Money trouble vs SimpleFin working better).
  • He also mentions fitness summaries and structured schemas in Obsidian.

6) Agentic “memory/workspaces” and model-agnostic future-proofing

  • The group emphasizes long-term survivability of notes/config:
    • Store plans/knowledge as Markdown/flat files in Obsidian so future model/harness changes don’t break workflows.
  • They mention vector-based memory systems (e.g., Nemo/Honcho), but Darren prefers a more future-oriented approach he describes as smarter/design-for-the-future vs slower vector-DB style patterns.

7) Cloudflare-hosted personal site + “private garden” access controls

  • A personal site is generated and published using Cloudflare Pages (a free-ish approach with good API/permissions).
  • Private content is gated via Cloudflare Access so internal stakeholders (e.g., IT/CISO) can view human-readable pages.

8) Legacy-code-to-modern-system generation (“code is the spec”)

  • A major demonstration: building an ad sales / continuity backend by using AI to translate legacy code into a new architecture.
  • Approach:
    • Analyze older system code and treat it as a concrete spec
    • The AI produces:
      • A structured spec in chunks (A/B/C/D style)
      • A coding plan handed off to an editor/coding model (mentions Opus / coding agent flow)
      • Chat/iteration loops for review and consensus
  • Implemented features include:
    • Insertion orders with schedule/rotation support
    • Rundowns for producers
    • Audit trails of changes
    • Fair rotation logic to rotate advertiser slots across a year, replacing a manual process
  • Bottleneck described: human review and UI/approval friction, not AI correctness.

9) Workflow tooling: deterministic engines vs “free-running agents”

  • They distinguish:
    • Agent workflows: can “forget” steps; often token-heavy and slower
    • Deterministic workflow engines/state machines: repeatable, robust, trackable
  • Examples discussed:
    • N8/Nondo-style workflows for dashboard/config tasks
    • Temporal for robust enterprise retries/recovery
  • Key takeaway: frequent recurring processes often justify a workflow engine.

10) Office Hours: clip/show assembly by transcript chunking

  • Craig demos a system for Office Hours, a media production Q&A show.
  • Key features:
    • Schedule orchestration for a volunteer-heavy operation
    • A searchable question archive where clicking a keyword shows segments
    • Transcript-based chunking to find topic-related content (basic versions may have imperfect topic keyword matches)
    • Autoplay to compile a “show” from earlier segments
  • Scale: tens of thousands of lines of code and thousands of recorded questions enabling navigation/filtering.

11) Feedback/ticket system with context + AI categorization

  • Craig describes an embedded feedback loop:
    • Users submit “bugs/ideas” with context (page, screenshots, current state)
    • AI categorizes/prioritizes:
      • P0/P1
      • cosmetic vs pain-in-the-ass
    • Comments close the loop so submitters see status changes
  • Darren suggests extending it so AI could also:
    • trigger fixes
    • route work to agents via integrations

12) Hermes integrations + social monitoring workflows

  • Anthony/Craig discuss Hermes:
    • recurring workflows (cron-like tasks)
    • integrations with stream tooling (e.g., Reream) to automate:
      • posting to Discord/Slack at go-live
      • generating checklists
      • cleanup of social posts afterward
      • future automation such as downloading assets and creating AI transcriptions/editor notes
  • Another workflow:
    • Social monitoring for YouTube comments
      • periodically fetch comments
      • run each comment through an LLM classifier
      • they mention preferring deterministic/classifier-style approaches for cost and stability

Reviews / guides / tutorial-style elements

  • Practical guides to running models locally
    • Using Ollama to manage multiple LLMs
    • Model switching approach
    • Local vs Spark offload comparison
  • Voice model usage guide (conceptual)
    • Kakoro training approach (overnight fine-tune)
    • Emphasis on speed vs quality
    • How transcript conditioning is used
  • Workflow engineering guidance
    • Preference for deterministic workflow engines for robust recurring steps
  • Show-prep automation guide
    • How to structure an Obsidian vault
    • How to use MCP permissions so agents can reliably research and format dossiers

Main speakers / sources (as referenced in the subtitles)

  • Leo Laporte (Twit host / moderator)
  • Darren
  • Anthony
  • Craig
  • Alakazip
  • Dano
  • Jeff Jarvis (referenced in the guest/voice demo context)
  • Steve (mentioned in chat about GRC/domain issue)

Original video