Agents Last Exam

Study / Research

A benchmark created by UC Berkeley to test AI models on thousands of expert-curated tasks with verifiable outcomes across 55 industries, focusing on economically valuable tasks and mastering industrial software.

Mentioned in 1 video