Agents Last Exam
Study / Research
A benchmark created by UC Berkeley to test AI models on thousands of expert-curated tasks with verifiable outcomes across 55 industries, focusing on economically valuable tasks and mastering industrial software.
Mentioned in 1 video
