Video summary
Preferences Over Benchmarks: Model Routing — Archana Kamath & Tyler Gillam, DigitalOcean
Main summary
Key takeaways
Main topic
The talk argues that model routing should be driven by request-specific “preferences” (cost/latency/quality/rules/task context) rather than by always selecting the top benchmark model. It positions model routing as a key capability in DigitalOcean’s inference engine.
Why “one best model” is the wrong approach (3 reasons)
- Cost explosion
- Inference spending is rapidly increasing; even large companies (examples mentioned: Walmart, Uber, Microsoft) cap usage to control inference bills.
- Fit / overkill
- Using one frontier / best-in-benchmark model for every task is usually inefficient because many tasks can be handled by smaller or cheaper models.
- Risk / reliability
- Relying on a single model creates a failure mode (no failover if the model degrades/outages).
- Model orchestration is presented as the “new phase” for production robustness.
Core premise: there is no single “best model,” only the right model per request
The “right model” depends on factors that benchmarks can’t fully capture:
- Task/request type (e.g., classification vs. code completion vs. code review/security)
- System prompts and tools used with the model
- Cost budget and willingness to trade off
- Latency requirements (not all use cases need the same responsiveness)
- End-user preference
Product/architecture: DigitalOcean Inference Router (live demo)
Key issues with prior “auto routing” approaches
- Earlier attempts felt like a black box; if routing produced poor results, builders couldn’t easily improve it.
How this router is designed differently
- Requests go through:
- an Open Proxy and
- a purpose-built routing model (both stated as open source)
- No vendor lock-in emphasized (open components; ability to keep control).
Router customization inputs
Users provide what matters for their workload:
- cost, latency, quality
- preferred models
- hard rules
- task description in natural language
Router execution includes:
- presets
- rule/decision-tree style controls
- ability to change options “in a single line of code”
Users validate using their own evaluations, not just public leaderboards.
Routing model performance claims
- < 200 ms routing decision time
- “Costs customers nothing extra” (router itself is free/included)
- In evaluations, routing is claimed to outperform frontier models on routing tasks with fractional latency.
Features highlighted in the demo
UI configuration
Demonstrated router presets such as:
- Software engineering
- General writing
- Knowledge bases
- Document intelligence
Example custom “Software engineering” tasks:
- bug fixing
- code generation
- test writing
Multi-model pools per task were shown.
Two routing strategies shown
- Manual ranking / preferred model with failover
- Example: for code generation, always try GLM 5.2; if it fails, fail over to GPT 5.2
- Fastest/recency-based selection
- Example: for bug fixing, choose the model that’s been fastest in the last ~30 minutes from a pool.
Demo results (what was measured)
Playground side-by-side tests
- Compared routing vs. always using a single premium model (e.g., “Opus” mentioned).
- Examples of tasks:
- Fibonacci function → router chooses cheaper/faster model based on code-snippet intent
- “Optimize my function” → routes to GPT 5.2
- “Write some unit tests” → routes to a model for test writing / code verification
- The pattern: faster + cheaper while still maintaining quality.
Evaluations
- An evaluation compared:
- Opus vs. routed approach
- Reported that correctness/quality scores were close (within judge margin of error), while routing used fewer tokens and was significantly faster.
Real workflow / observability with OpenCode
- Two terminals compared:
- Left: single-model approach (always “Opus”)
- Right: OpenCode configured to use the software engineering router
- Live observability shown:
- token usage in real time
- which models were selected
- task-to-model mappings
- cost accumulating live
- latency per step
Across session steps, reported large cost reductions, e.g.:
- One session step: router ~8 cents vs. Opus ~25 cents (~3x cheaper)
- Another: total router 14 vs Opus 44 (units implied as cents/total session cost)
Claimed latency is improved because routing selects appropriate models per step.
“Routing is a foundation” — what’s built on top
The talk lists three higher-level layers that leverage routing:
- Eval: validate the right model for your specific use case with your tests
- Caching: avoid paying repeatedly for the same answers
- Personalization: router learns what works for your team over time
It emphasizes a continuous improvement loop: route → evaluate → adjust.
Quick factual claims (explicit “quick facts” section)
- Routing decision time: under 200 ms per request
- 0 application code changes needed to adopt
- Free/included (no separate router cost or rollout burden)
- Router model is open sourced (referred to as via “Plano”)
- Router sits in the inference engine layer (DigitalOcean stack context described earlier)
Main speakers / sources
- Archana (Archa) Kamath — VP of Engineering for Inference Engine & AI Infrastructure, DigitalOcean
- Tyler Gillam — built parts of the router; performed the live demo