SWEBench
Study / Research
A software benchmark mentioned to illustrate that high performance on benchmarks (like Gemini on SWEBench) doesn't always translate to practical utility.
Mentioned in 3 videos
Save the 3 videos on SWEBench to your own pod.
Sign up free to keep building your knowledge base on SWEBench as more episodes are added.
Videos Mentioning SWEBench

Recursive Language Models — Alex Zhang, MIT PhD
Latent Space
A benchmark for code generation where LLMs navigate codebases, used to illustrate the limitations of single LLM calls.

🔬 "The Most Innovative Diffusion Research Is Happening in Drug Discovery, Not Image Generation"
Latent Space
A software benchmark mentioned to illustrate that high performance on benchmarks (like Gemini on SWEBench) doesn't always translate to practical utility.

Going In Deep On Data | YC Paper Club
Y Combinator
A benchmark for coding agents, the original SWEBench team from Princeton collaborated on senior SWEBench.