Nov 29, 2023
Elo Uncovered: Robustness and Best Practices in Language Model Evaluation
Elo has proven effective for dynamic games like chess and has recently seen widespread use for evaluating LLMs. But how reliable is it for evaluating static-skill entities like LLMs? We find scenarios where it’s not!

Authors
Meriem Boubdir, Edward Kim, Beyza Ermis, Sara Hooker, Marzieh Fadaee
Abstract
Related works

Research
When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs
Read

Research
CIRCLE: A Framework for Evaluating AI from a Real-World Lens
Read

Research
Findings of the WMT25 Multilingual Instruction Shared Task: Persistent Hurdles in Reasoning, Generation, and Evaluation
Read






