guide

Llm judge metrics

A practical answer grounded in a runnable, published harness rather than opinion. Links to Can you trust the model that grades your content? Measuring when an AI judge waves through broken work.

Status
Answer to a real buyer question · grounded in a runnable harness · auto-published

LLM judge metrics come down to one question. Does the AI actually help students learn, or does it just sound plausible? I lead with education-domain judgment, then build runnable AI-quality harnesses that make that judgment measurable. This page walks through the method and links to one you can run against your own system.

This is the method behind Can you trust the model that grades your content? Measuring when an AI judge waves through broken work. You can read it and run it yourself.

1. Make the failure observable.

Pick the specific way this can go wrong, and build the smallest input that triggers it. If you can’t make it fail on purpose, you can’t prove it works. For a worked example, see Measuring feedback integrity: a blind-solver that catches AI explanations leaking the answer.

2. Measure against a baseline, not a vibe.

Compare to a neutral control so the number means something. A score with nothing to compare it to is theater. For a worked example, see Same answer, different grade: measuring when an AI grader can’t hold a verdict.

3. Check it a second, independent way.

Re-run with a different model family or a held-out set. Agreement across independent checks is the only verdict worth trusting. For a worked example, see Measuring feedback integrity: a blind-solver that catches AI explanations leaking the answer.

My work sits at the overlap: deep classroom, curriculum, and assessment judgment, paired with AI-eval tooling that makes that judgment measurable. Every claim on this page maps to a public, reproducible harness. It’s not a slide. If you want this run against your own system, the method transfers directly.