Encyclopedia · 176 concepts

MLOps & Infrastructure · intermediate · concept 141 of 176

LLM Evaluation & Benchmarks

Measuring what language models can actually do, via static benchmarks, head-to-head human preference arenas, and LLM-as-judge scoring. Benchmark saturation and training-data contamination make honest evaluation one of the field's hardest open problems: leaderboard gains do not always translate to real-world capability. In practice evals are the CI of AI products: rubric-scored suites run on every prompt or model change, increasingly with agent trajectories under test.

Key terms

MMLUSWE-benchLLM-as-judgeContaminationElo arenaAgent evalsPrompt evals

Where you meet it in the real world

Model selection, regression testing, procurement decisions, safety assessment

Courses, papers, and more