SKILL.md
Context Intelligence Evaluation Methodology
Mode-only skill for how to measure. It complements — and never restates —
context-intelligence-eval-design (which owns scenario mechanics and the two-layer
structural/behavioral structure) and digital-twin-universe (which owns the DTU machinery).
Scope
In scope
- Metric design. Choose metrics across three axes: quality (did it detect what the user means?), efficiency (token/tool cost to detect), efficacy (does detection drive the right outcome?).
- Measure the precursor, not only the failure. Prefer leading indicators (e.g. bounded vs climbing context growth) over lagging ones (e.g. a final timeout). The precursor is testable in a short window; the full failure often is not.
- A/B + statistical-N discipline. A single green run is not proof of a behavioral change. Compare a control arm against a treatment arm; require N independent trials and report the pass rate, not a single anecdote.
- Test data fidelity. Validate on real sessions where available, or on faithfully modeled synthetic corpora that preserve the distribution that matters (e.g. heavy-tail file sizes), never on data hand-shaped to pass.
Out of scope (point elsewhere — do not restate)
- Scenario mechanics, the two-layer structure, DTU profile templates / Gitea URL rewrite →
context-intelligence-eval-design. - DTU lifecycle, launch/exec/profiles → the
digital-twin-universeskill. - "DTU-as-default" rationale and the "artifact-as-success" anti-pattern → already authoritative
in
context-intelligence-eval-designandcontext-intelligence:context/context-intelligence-primitives-reference.md. Reference them; do not repeat them.
