SKILL.md
Eval Rubric Designer Skill
You can't improve what you can't score. The hard part of evaluating AI output isn't running the judge — it's defining dimensions that are specific, observable, and independent, with anchors concrete enough that two people (or two judge runs) agree. This skill turns "is the output good?" into a rubric and a judge prompt you can run today.
Working from a brief
Given just "I need to eval my summariser", produce the full rubric anyway — infer the task, the output type, and the dimensions that matter for it, and label inferred choices. Never hand back a list of dimension names with no anchors; the anchors are where the rubric earns its keep.
Required Inputs
Ask for these only if they aren't already provided (else infer and label):
- The task — what the AI is supposed to produce, and for whom.
- A sample output (or two) — ideally one good and one weak, to calibrate anchors.
- What "good" means here — the quality bar and any non-negotiables (e.g. must be grounded, must follow format).
- How it'll be scored — human review, LLM-as-judge, or both; and whether you need a single score or per-dimension.
Output Format
Eval Rubric: [task]
1. Dimensions — 3–6 independent dimensions, each with a one-line definition and a weight. Default set, tailored to the task: structure, completeness, correctness/grounding, usefulness, safety/tone.
2. Anchors — for each dimension, concrete descriptions at 1, 3, and 5 (what a poor / acceptable / excellent answer looks like for this task). Anchors must be observable, not "feels good".
| Dimension (weight) | 1 — poor | 3 — acceptable | 5 — excellent |
|---|---|---|---|
| Grounding (×2) | invents facts not in the source | mostly grounded, minor drift | every claim traceable to the source |
