Video summary

Agent Frameworks Considered Harmful — Rémi Louf, .txt

Main summary

Key takeaways

Technology

Context / Motivation (Rémi Louf)

After agents became dramatically better—attributed to “opus 4.6.x / 4.6 six”—Rémi Louf (CEO of a small AI company) took two weeks off to “scratch his own itch.”

He automated his repetitive morning workflow into an automated daily brief delivered to Slack:

  • Market/news reviews
  • Updating systems like CRM / Jira / Linear
  • Processing long voice notes recorded during walks

Problem with Existing “Agent” UX

Even when tools and apps are available, many agent systems feel transitional and operator-dependent.

  • His analogy: it can feel like using remote control—similar to having to stay “on” to operate an automated mower/tractor.
  • Example pain point: you may still need to re-steer the agent from a phone, rather than letting it operate autonomously.

Build Approach: “Agents as Events” (Avoid Edge/Graph Maintenance)

He initially tried existing agent frameworks, but found he spent much of the time editing prompts.

Instead, he preferred an “edit-a-file, drop it in a folder, it appears at runtime” workflow—avoiding heavy YAML/code-heavy setup.

Core Runtime Idea

Agents subscribe to events and publish events.

  • Minimal “edge” management
  • Simplified event-driven orchestration

Example Event Flow

  1. Voice note arrives → emits a new event
  2. Voice note agent:
    • accepts the voice note
    • transcribes it
    • converts it into durable notes
    • emits a voice note processed event
  3. Daily brief agent:
    • consumes the processed outputs
    • posts a Slack message

Event-Driven vs Cron/Jobs

He contrasts this with cron/jobs:

  • Cron is time-based
  • His emphasis is event-based triggers tied to:
    • new email
    • CRM updates
    • PR open/merge, etc.

Observability and Reliability: Learning Through Failures

Early versions had real issues, including:

  • duplicate Slack posts
  • missing or “vanished” voice notes
  • prompt-iteration mistakes

Each failure led to runtime improvements:

  • A persistent append-only event log (“systems memory”) to:
    • prevent losing data
    • enable debugging and audit
  • Fixes for retries/queueing by adding proper handling of:
    • attempts
    • state

Content-Addressed “Graph” Representation for Prompt Auditability

He argues that many agent tools create a misleading “live chat” illusion: you often can’t tell what the model truly saw, due to internal handling such as:

  • compaction
  • provider behavior
  • hidden traces

Solution: Hashable Prompt Components

Represent prompts as a list of hashed components, such as:

  • system prompt
  • tool descriptions
  • skill descriptions
  • user message
  • and other relevant parts

By storing prompt components and answers via hash/content addressing, he enables:

  • Auditability: trace exactly which context led to an output
  • Easier compaction/context management (operate on a graph rather than raw strings)
  • Diffs between runs:
    • which components changed (user message vs tools/skills/system parts)

Replay Capability

Because the prompt/graph is recorded, he can:

  • reconstruct prior requests
  • replay and evaluate them by:
    • re-running with different models
    • resending the same structured request deterministically
      • useful for debugging and cost control

Kernel/Runtime Boundaries to Prevent “Bad Actions”

He emphasizes a “kernel” approach: certain mistakes should be impossible, not merely unlikely.

Two key boundaries:

  1. Typed tool calls
    • agents can only call tools that exist and match expected types
  2. Typed events between agents
    • non-negotiable typed contracts reduce errors between agents

Structured Outputs as a Specialty

His company focuses on structured outputs, and this was a “dogfooding” project because a prior provider was “terrible” at structured outputs.

  • Result: roughly 20% of events were wrong or rejected
  • The runtime and typed interfaces are designed to constrain actions and outputs to valid schemas

Claims / Takeaways from Deployment

After internal deployment:

  • He reports deploying ~20 agents, not only by technical staff
  • He claims background agents can feel “magical”:
    • mornings end with an inbox daily brief that’s even better than manual processing

Advice

  • Open-source models are good enough for his use case (including local models)
  • The “agent framework” ecosystem is still unsettled
    • build before you buy
  • Framework builders should eat their own dog food
  • Immerse/experiment seriously:
    • his two-week dive changed the company’s trajectory

Additional Note

  • The project code isn’t being sold
  • He encourages reading the blog / stealing the code

Main Speaker / Sources

Speaker

  • Rémi Louf (CEO of Text; mentions CTO and internal deployment at his company)

Sources referenced (not primary speakers)

  • OpenAI models / Anthropic Claude
    • referenced in relation to hidden reasoning/traces behavior
  • Prior “codeex”/framework tools and general “cron jobs” concept
    • no other named speakers mentioned

Original video