Holistic Evaluation of Language Models (HELM)
Beyond the Imitation Game collaborative benchmark for measuring and extrapolating the capabilities of language models - google/BIG-bench
Measuring Massive Multitask Language Understanding | ICLR 2021 - hendrycks/test
LiveBench AI模型实时评测平台,追踪和比较各模型性能变化
SWE-bench Leaderboards