Video summary
Research Insights Made Simple #26: как строить работающие evals для AI-агентов
Main summary
Key takeaways
Main ideas, concepts, and lessons
-
Evals shouldn’t just test “did it pass?”—they should evaluate the whole AI system and its business/production behavior.
- Instead of evaluating only a model (e.g., on benchmark-style success/failure), evaluate:
- the task outcome
- the result
- the execution trace (how the agent/system arrived there)
- operations questions like stability and cost when integrated into a process
- Instead of evaluating only a model (e.g., on benchmark-style success/failure), evaluate:
-
Work with “reproducible scenarios” and a “gate” for rollout.
- Similar to engineering testing:
- define a starting state
- define a contract (what the agent can do)
- run the agent on reproducible scenarios
- use judgement/grading to decide whether to roll out or not
- Similar to engineering testing:
-
Use an evaluation loop that mirrors production reality (offline ↔ online loops).
- Offline loop (pre-rollout):
- collect user-like traces/corner cases
- improve prompts, graders, and evaluation datasets
- Online loop (post-rollout):
- in production, collect traces from real users
- annotate corner cases
- expand the evaluation dataset
- Offline loop (pre-rollout):
-
LM-as-judge and deterministic checks are complementary.
- Some evaluations can be deterministic and fast (regex/blacklists, code-level checks, mathematical checks).
- Others require LM judgment (quality/rubric-based evaluation).
- LM-as-judge may require calibration (e.g., using reasoning, comparing to reference preferences, or asking other models).
-
Eval design should support non-determinism and agent behavior.
- Agents are not just single outputs—they involve:
- planning
- tool usage
- multiple steps/turns
- context management
- trajectories that can go off-track
- Therefore, evals should measure trajectory features, not only final answers (e.g., number of steps/turns, tool calls, whether the agent stayed within guardrails).
- Agents are not just single outputs—they involve:
-
Different “types of evals” for business/product settings.
- The speaker frames evals into multiple categories (explained in the talk), emphasizing how they map to business needs.
Methodology / framework presented (detailed bullet points)
1) Build evals as part of product logic (not just a benchmark)
- Treat evals/prompts/graders as business-logic components that must be tested like any production system.
- Don’t assume third-party accuracy will remain stable—models and toolchains change frequently.
2) Use a “two-circuit” evaluation approach
-
Circuit A: offline (development + rollout preparation)
- prepare/improve:
- prompts
- context assembly strategy
- evaluation datasets
- graders/judges
- use reproducible scenarios and run them repeatedly before deployment
- prepare/improve:
-
Circuit B: online (production feedback loop)
- real users generate traces
- you splice/merge/manual-markup corner cases
- expand the dataset and refine evals (and prompts/graders) over time
3) Choose the appropriate grader/eval type (four categories described)
-
Type 1: Deterministic fast checks
- binary pass/fail where possible
- implemented via business rules:
- rules for prohibited content/PII
- forbidden words via regex/blacklists
- code-level or math-based checks
-
Type 2: “LM as a judge”
- use an LLM to score outcomes
- can be binary or graded/ranged
- should include/encourage reasoning or justification (with care, since reflection/explanations can be gamed)
- may require hacks like:
- asking other models for consistency checks
- improving prompts to get better judge behavior
-
Type 3: Subject-matter expert evaluation (“subject meta-experts”)
- human experts prepare/label datasets:
- build representative scenarios
- do markup on examples (confirm what’s wrong, what’s missing)
- can start small internally (even ~100 cases as a starter dataset)
- scaling considerations:
- internal trusted annotators first
- external expert platforms may be needed for larger rollouts or higher risk
- human experts prepare/label datasets:
-
Type 4: User-facing/feedback-loop signals (“online metrics” and product analytics)
- incorporate user feedback:
- thumbs up/down, ratings
- behavioral traces from the product (click paths, user actions)
- use analytics baselines and compare variants (especially in B2C/internal products with high feedback volume)
- incorporate user feedback:
4) Build evals using “task → graders → metrics → outcome/trajectory”
A system-level view was described (inspired by an external “anatomy” of eval systems):
-
Task definition
- includes expected input/output
- works for both prompt-based and agent-based flows
-
Graders
- combine deterministic checks + LM judging
-
Metrics
- not always obvious; derive them from the task and operational constraints
- examples discussed:
- latency/cost impact (long checks affect performance)
- PI/secret leakage detection via in-request checks
- success path efficiency (how many steps were needed)
-
Outcome / trajectory measurement
- for agents:
- measure how the agent moved toward the goal
- count steps/turns/tool calls
- detect early stops vs excessive wandering
- correlate mistakes with trace segments
- for agents:
5) Make evaluation instrumentation part of the system (tracing/telemetry)
- Collect execution traces and product analytics together:
- technical traces (e.g., via OpenTelemetry/DataBricks-like tooling)
- user journey/product interactions (what the user clicked, which buttons, etc.)
- Use these traces to:
- build evaluation sets
- replay realistic behavior in offline evals
- treat real interactions as positive/negative examples
- detect failure modes that pure final-answer scoring misses
6) Use evals to counter model/tool “cheating” and regressions
- Models can exploit eval loopholes (e.g., creating dummy files or placeholders that satisfy surface checks).
- Evals must evolve with:
- new models (behavior changes)
- new tool providers / pricing / availability
- prompt changes becoming insufficient for newer models
- Without evals, teams fall into:
- “trusting the demo” or “it’s the vendor—so it’s fine” behavior
7) Roll out progressively using the maturity pyramid idea (L0 → higher levels)
- Start with prototypes / one-shot prompts + small datasets:
- “demo-level” correctness
- hackathons are mentioned as an easy starting point
- Then integrate into loops and production traffic:
- use telemetry + feedback to improve prompts/graders
- Over time:
- prompts become business logic
- you aim for closed-loop improvements (offline + online)
8) “Skill” / tool evaluation as a control mechanism (measurability)
- Many agent “skills”/tools should provide their own evaluation signals.
- Measure whether turning a skill on/off affects task quality.
- Potentially deactivate skills if they are harmful or unnecessary.
- Guidance summary: testing/measurability pushes closer to standard engineering/testing maturity.
Key examples mentioned (how evals apply)
-
Localization automation
- correct behavior across languages and models
- evaluation can include expert validation and feedback loops
-
Virtual assistant / regulated domains
- some sections should not be edited (risk of breaking rules/constraints)
-
Agent trajectories
- measure step counts and tool-call behavior (too many turns implies issues)
-
PII leakage / policy checks
- run checks during request/response cycles to detect unwanted data
- tie metrics to operational cost/latency
-
Code-related examples
- discussion of “cheating” in benchmarks and the need for robust checks
Speakers / sources featured (identified)
Speakers
- Sasha (host; name not fully captured in subtitles)
- Zhenya (Evgeny/engineering director) — Engineering Director at Flo (as stated), discussing:
- eval development
- product analytics/tracing approach
Organizations / tools / sources mentioned
- Flo
- MLflow
- DataBricks (tracing integration concept)
- OpenTelemetry (telemetry/tracing standards)
- Toloka / Mechanical Turk (examples of crowdsourcing/external labeling platforms)
- TSSR (tool mentioned for generating datasets)
- .gitignore (mentioned in the context of cheating/benchmarks; “Git ignore” implied)
- Harness (mentioned as a benchmark/evaluation platform)
- AI providers/models mentioned (in passing):
- Anthropic (Claude, Opus)
- Gemini
- OpenAI
- Hugging Face
- MCP (mentioned in a tooling compatibility context)
- Meta (mentioned in context of a paper/benchmark discussion; exact paper title not given)
- A third-party referenced “anatomy of systems / eval building” framework (author not specified in subtitles)