GSM8K
A dataset that first showed signs of chain-of-thought reasoning capabilities in large language models.
Save the 5 videos on GSM8K to your own pod.
Sign up free to keep building your knowledge base on GSM8K as more episodes are added.
Videos Mentioning GSM8K

Stanford CS329A Self-Improving AI Agents | Part 1 | Course Overview
Stanford Online
A dataset that first showed signs of chain-of-thought reasoning capabilities in large language models.

Stanford CS329A Self-Improving AI Agents | Part 6 | Train Time Scaling/Scaling RL
Stanford Online
A dataset of grade-school math problems used to test the STAR approach.

Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling
Stanford Online
A dataset used to illustrate the generation verification gap, showing that majority voting works better for simpler problems compared to more complex ones.

Stanford CS329A Self-Improving AI Agents | Part 5 | Planning and Multi-Step Reasoning
Stanford Online
A dataset of math problems used in the Swirl experiments to train and test multi-step reasoning and tool use capabilities.

Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification
Stanford Online
A dataset of 8,500 grade school math problems designed to test multi-step reasoning in LLMs, introduced in the first paper.