Video summary
FinOps for AI Agents: Who Spent All the Tokens? — Tisha Chawla & Susheem Koul, Microsoft
Main summary
Key takeaways
Summary (FinOps for AI Agents: “Who Spent All the Tokens?”)
The talk explores why AI agent token spending is often hard to trace and control—due to runaway loops, exploding context, and unbounded tool/model calls—and proposes a shift from token maxing to value maxing through token cost governance.
Core problem and why it’s hard
- Token cost is the unit of cost (billing is tied to tokens).
- Cost is generated at the LLM/model call boundary, but existing “token ops” tooling typically operates at the request level (e.g., model gateways with routing/downgrades, hard caps, halting).
- In agent systems, what’s missing is the ability to control at the agent run / tool / sub-agent / context growth boundaries—i.e., the call path between agents and tools.
First-principles approach to cost control
- Treat tokens as both the cost unit and the value metric to optimize.
- Track where cost is created (at the model call boundary) and enable attribution.
- Use attribution to enforce policies that stop the root causes:
- mitigate loops
- prevent excessive context growth
- reduce irrelevant tool outputs
- Use hard budget halting only as a last resort (budget cap as a circuit breaker).
Proposed solution: “TokenOps” platform (run-level governance)
The proposed platform includes:
- Cumulative budgets across attributed runs, not just per-request caps.
- Enforcement on the call path (run-level), enabling actions that steer agents before killing them.
- An architecture with an out-of-band (outbound) control plane that does not interfere with application code.
Architecture (out-of-band plane) and modules
An outbound control plane is described with three major modules:
-
Instrumentation
- Observability using OpenTelemetry
- Cost telemetry and enrichment
- Attribution: identify which agent run / user dimension caused spend
-
Accounting
- Ledger-style accumulation of total runs/spend
-
Enforced layer
- Steering and policy execution
- Halt as final protection when budgets are exhausted
Demo/code-level integration concepts
Key integration mechanism: a boundary annotation.
- Applied to methods (intended to work across frameworks such as LangChain).
- Tracks input/output
- Forwards traces to the control plane
- Writes ledger entries tied to run IDs and attributes
- Provides a channel for the control plane to push actions back into agent behavior
Additional integration notes:
- The control plane runs inside the tenant (claimed to reduce data leak concerns).
- A governor component uses developer configs defining which actions are allowed (to prevent arbitrary or untrusted changes to agent behavior).
Control plane: segments, budgets, actions, policies
Inside the control plane:
-
Segmentation
- Group runs by “dimensions” (e.g., cohort tags like user cohorts)
- Apply budgets at different granularities (agent-level, run-level, cohort-level rollups)
-
Budgets
- Static thresholds over time windows for segments/runs
-
Actions (two flavors)
- Halt actions: kill/stop the agent if budget is exceeded
- Steer actions: modify behavior to fit within budgets (avoid immediate termination)
-
Policies
- Combine budgets + actions + conditions
- Execute enforcement for targeted segments/runs
Policy catalog / failure modes covered (examples)
The approach benchmarks against policy coverage intended to address:
- Spend management
- budget/guardrails
- Context management
- context compaction
- tool output reduction
- Loop detection and progress detection
- Additional run control mechanisms for “runaway token” failure modes
Steering examples (behavioral changes)
A “cost guard” mechanism can:
- Consider both:
- budget consumed so far, and
- velocity (rate of token consumption)
- If it predicts running out by the end of the run, it injects guidance into system instructions, e.g.:
- “be more succinct”
- “summarize more”
Goal: reduce future token usage rather than killing immediately.
Demo scenarios and results
Scenario 1: Preview mode
- Policies are evaluated, but enforcement actions are not executed.
- Dashboard shows which policies would run; run completes without failures.
Scenario 2: Governance enabled (halt behavior)
- If the agent exceeds the cost cap, it gets killed immediately (circuit breaker).
Scenario 3: Steering behavior (cost guard)
- Budget is higher than the halt scenario but still insufficient for normal completion.
- Instead of killing, the system steers generation to fit the budget.
Benchmark claims
- Tested on open source repos / stress tests (examples named: Browser-use and MetaGPT).
- With the full policy suite enabled:
- Average spend reduced ~78%
- Completion rate improved from ~67% to ~96% compared to simple throttling (which kills runs)
Speaker / source identification
- Tisha Chawla (presenter)
- Susheem Koul / Sashim (co-presenter; demo presenter)
- Source context: Microsoft event/talk (per video title)