A New Benchmark for Evaluating Reward Models and LLM Judges
Contributors:
Evan Frick
Tianle Li
Connor Chen
Wei-Lin Chiang
Anastasios N. Angelopoulos
Jiantao Jiao
Banghua Zhu
Joseph Gonzalez
Ion Stoica
Introduction
We introduce a benchmark to solve this problem: PPE, a collection of 16,038 labeled human preference pairs from Chatbot Arena containing responses from 20 different top LLMs and over 121 languages as well as a dataset of 2,555 prompts, each with 32 different sampled response options, totaling 81,760 responses across 4 different models, all grounded with verifiable correctness labels. PPE evaluates reward models on 12 different metrics and 12 different domains, such as their accuracy in selecting human-preferred or verifiably correct responses.
To summarize:
- We curate high quality ground truth preference pairs from Chatbot Arena battles as well as existing verifiable correctness benchmarks.
- We experimentally correlate metrics on each benchmark to downstream RLHF-ed LLMs.
- We fully open-source PPE, the resulting comprehensive benchmark for reward models with metrics directly linked to downstream RLHF outcomes.















