Evaluation Golden Set

Golden Set Explorer

The 300 hand-written prompts behind every number on the benchmark report. Each one carries a task category, the Gemini tier a reviewer expects to be sufficient, the rubric the LLM judge scores against, and a reference answer when one exists.

How the set was built

Every prompt is written by hand and carries its own source material, so nothing depends on a document the model cannot see. The set is generated from a prompt bank in the repository, so it is reproducible and diffable.

01

24 seed archetypes

One archetype per category and tier, each with a written justification for why that tier should be enough.

02

Written by hand

300 distinct prompts, each with the rubric the judge grades against. Rubrics, handbook excerpts, code, and passages are inline.

03

Validated

Schema and distribution checks enforce a 40 / 35 / 25 split across lite, standard, and pro and an even spread across categories.

04

Reference answers

Prompts with one right answer carry a reference the judge can check against. Open-ended prompts are graded on the rubric alone.

Category
Expected tier
ID Category Expected Tier Prompt Judge Rubric Reference
Loading golden set…