LLM Eval Harness — Live Demo

Score model answers against a gold set with per-question match rules. This is a browser demo of my open-source llm-eval-harness (Python, tested) — edit either box and re-run. Read the write-up →

Match rules:
exactnormalized answer equals gold
containsgold appears inside the answer
numericfirst numbers within tol