Humanity's Last Exam

Humanity's Last Exam (HLE) is a challenging AI benchmark of 2,500 questions released in January 2025 by the Center for AI Safety and Scale AI. It was the brainchild of Dan Hendrycks, who was inspired after Elon Musk said existing benchmarks like MMLU were too easy. The questions were crowdsourced from subject-matter experts, filtered by top AI models, and reviewed by humans, with a $500,000 prize pool rewarding the best submissions. Requiring graduate-level expertise, the test spans subjects like mathematics (41%), physics, and biology, with 14% of questions needing multimodal understanding. However, a July 2025 investigation by FutureHouse suggested that roughly 30% of text-only chemistry and biology answers might be wrong, leading the team to launch a continuously updated version called HLE-Rolling.