Route each prompt to the cheapest Gemini model that can answer it, and prove it with numbers. Scored against a 300-prompt golden set with an LLM judge.
Log-scale cost (X-axis) vs LLM Judge Quality Score (Y-axis, 1-5 scale). The ideal strategy appears in the top-left quadrant (high quality, lowest cost).
| Strategy | Prompts | Judge Score (1-5) | Quality Retention | Cost / 1k Req | Cost vs Pro | Latency (p50 / p95) | Tier Mix (L / S / P) | Routing Acc | Escalation |
|---|
| Category | Evaluated Prompts | Mean Score | Cost / 1k Requests |
|---|