MLOps & Infrastructure · intermediate · concept 141 of 176
LLM Evaluation & Benchmarks
Measuring what language models can actually do, via static benchmarks, head-to-head human preference arenas, and LLM-as-judge scoring. Benchmark saturation and training-data contamination make honest evaluation one of the field's hardest open problems: leaderboard gains do not always translate to real-world capability. In practice evals are the CI of AI products: rubric-scored suites run on every prompt or model change, increasingly with agent trajectories under test.
Key terms
Learn these first
Where you meet it in the real world
Model selection, regression testing, procurement decisions, safety assessment
Videos
Stanford Online · YouTube
IBM Technology · YouTube
Guides and articles
Courses, papers, and more