Score model answers against a gold set with per-question match rules. This is a browser demo of my open-source llm-eval-harness (Python, tested) — edit either box and re-run. Read the write-up →
| exact | normalized answer equals gold |
| contains | gold appears inside the answer |
| numeric | first numbers within tol |