Video summary
Harness Engineering on Rails - Joël Quenneville
Main summary
Key takeaways
Summary
Joël Quenneville argues that AI-assisted coding can produce a working first draft quickly, but the time spent fixing its quality problems may erase the apparent productivity gain. The answer, he says, is not simply better prompting: developers should treat an agent as a loop and improve the tools and feedback around it.
Agents as iterative systems
An LLM on its own takes text in and produces text. An agent harness adds code that lets the model interact with its environment—for example, read and write files, run tests, and respond to errors. In an agent loop, the system can take an action, inspect the result, and make further model calls until it reaches an exit condition. This lets an agent correct an imperfect first attempt rather than needing to generate perfect code immediately.
Quenneville illustrates the idea with FizzBuzz: a first draft can contain a logic error, such as checking divisibility by three before divisibility by fifteen. Running tests and feeding the failure back through the loop can guide the agent toward a correct version. He says commercial coding tools such as Cursor, Claude Code, and Codex use generalized versions of this loop.
Move rules into tools that enforce them
A long AGENTS.md or similar instruction file can help, but it is still a prompt: the agent may not follow every rule, especially as its context grows. Quenneville recommends moving repeatable judgments into mechanisms that enforce or check them more reliably:
- Linters and custom RuboCop cops: Run RuboCop on changed Ruby files, and consider custom cops for recurring application-specific mistakes, such as incorrect scoping in a multi-tenant app.
- Hooks: Use lifecycle hooks to run checks before an agent exits. If a check fails, the hook can prevent completion and prompt the agent to fix the issue.
- Rails generators: Let generators handle conventions such as migration filenames instead of asking the agent to reproduce the rules from prose. A hook or interceptor can block direct writes to a migrations directory and tell the agent to use the generator instead.
- Generators and templates for project patterns: Encode conventions such as controller inheritance or associated authorization policies in the tools that create those files.
His broader point is to make the correct path the easy path for an agent. Instructions can still be useful, but he favors keeping AGENTS.md relatively thin—more like a routing layer than a repository for every coding rule.
Use focused reviews for less mechanical judgments
Not every quality concern can be captured with a linter or generator. For more subjective review, Quenneville recommends narrow review skills rather than a broad file of general programming advice. His example uses three stages: a structural review of whether the solution fits the problem, a more methodical review, and a final pass for readability and documentation.
These reviews can run through subagents and be partially automated. Quenneville distinguishes this from asking the agent to follow heuristics before it writes anything: reviewing a specific, generated program for violations is often easier than steering the model across the whole space of possible programs.
A framework for harness controls
The talk presents a two-by-two model of agent controls:
- Direct, inferential controls: Prompts and
AGENTS.mdinstructions. - Direct, deterministic controls: Generators, configuration scripts, and templates.
- Feedback-based, deterministic controls: Tests, RuboCop, other code-quality tools, and security scanners.
- Feedback-based, inferential controls: Agent-driven reviews and analyses.
Quenneville recommends shifting suitable judgments out of prompts and into the other categories, while warning against over-constraining agents. Citing Goodhart’s law, he cautions that when a metric becomes a rigid target, an agent may optimize the score without producing the result developers actually want.
Domain-specific scripts and continuous improvement
For a document-processing pipeline, Quenneville found that asking an LLM to inspect production data and assorted code produced vague or conflicting explanations. He wrote a domain-specific script to classify documents by how far they progressed through the pipeline. The resulting categories gave the model stable comparisons and a more useful vocabulary for investigation; the data could also be visualized to reveal where processing was failing.
He notes that these scripts, generators, and checks can improve the experience for human developers as well as agents. His suggested improvement loop is to examine where a person had to intervene, what signals or tools were missing, what information had to be copied manually, and whether recurring errors can be prevented. He also describes a “retro” skill that reviews a long AI-coding session and suggests possible automations.
Practical guidance and references
- Don’t send an unpolished AI-generated change to colleagues and pass the cleanup burden to them.
- When an agent makes a recurring mistake, look for a way to change the system so it can detect or avoid that mistake next time.
- Use deterministic checks for enforceable rules; reserve prompts for guidance that cannot be encoded reliably.
- Consider domain-specific scripts when agents need clearer categories, consistent definitions, or better visibility into a workflow.
- Quenneville mentions his
joelq/skillsGitHub repository, including the retro skill, and points to a separate “retro” skill by Matt Pocock. - He recommends watching a forthcoming recording of a talk by Rachel about Rails generators. He also refers to Kenzie’s earlier talk about using source code as a control mechanism.
Reviews, guides, or tutorials: This is a technical presentation with examples and practical implementation guidance, rather than a product review. It discusses how to combine agent loops, tests, hooks, linters, generators, review skills, and domain-specific scripts to improve AI-generated Rails code.
Main speaker/source: Joël Quenneville, lead developer at thoughtbot and co-host of the Bike Shed podcast. Other sources mentioned include Matt Pocock, and talks by Rachel and Kenzie.
Rate this summary
Your feedback will help improve summaries.
Improve this summary
Reprocess with a stronger model when the summary feels incomplete or inaccurate.
Translate summary in another language
Ask questions to this video
Chat for follow-up questions, clarifications, and source-backed answers.