The 300 hand-written prompts behind every number on the benchmark report. Each one carries a task category, the Gemini tier a reviewer expects to be sufficient, the rubric the LLM judge scores against, and a reference answer when one exists.
Every prompt is written by hand and carries its own source material, so nothing depends on a document the model cannot see. The set is generated from a prompt bank in the repository, so it is reproducible and diffable.
One archetype per category and tier, each with a written justification for why that tier should be enough.
300 distinct prompts, each with the rubric the judge grades against. Rubrics, handbook excerpts, code, and passages are inline.
Schema and distribution checks enforce a 40 / 35 / 25 split across lite, standard, and pro and an even spread across categories.
Prompts with one right answer carry a reference the judge can check against. Open-ended prompts are graded on the rubric alone.
| ID | Category | Expected Tier | Prompt | Judge Rubric | Reference |
|---|---|---|---|---|---|
| Loading golden set… | |||||