Transformer

Concept

Neural network architecture visualized and explained as the core model type for LLMs.

Mentioned in 42 videos

Build a research pod on Transformer.

42 expert discussions. Save them all to your own pod, ask any question, get cited answers.

Get Started Free

Videos Mentioning Transformer

Transformers Explained: The Discovery That Changed AI Forever

Transformers Explained: The Discovery That Changed AI Forever

Y Combinator

A neural network architecture that uses self-attention to model relationships in data and generate outputs, forming the basis for many state-of-the-art AI systems.

Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample

Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample

Latent Space

A neural network architecture that is a core component of many modern AI models, including those discussed for audio processing.

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 1 - Diffusion

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 1 - Diffusion

Stanford Online

Current convergence point for image generation architectures, moving towards transformer-based designs like the Diffusion Transformer.

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 4 - Latent Space & Guidance

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 4 - Latent Space & Guidance

Stanford Online

An encoder-decoder architecture centered on attention, introduced in 2017, foundational for most modern language and many vision models.

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 5 - Architectures

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 5 - Architectures

Stanford Online

An architecture that revolutionized the NLP field in 2017, based on the concept of self-attention mechanisms, which can also be applied to images.

Stanford CS153 Frontier Systems | Amit Jain from Luma AI on Unified Intelligence Systems

Stanford CS153 Frontier Systems | Amit Jain from Luma AI on Unified Intelligence Systems

Stanford Online

A highly effective architecture that Luma uses and believes is key to future AI models due to its ability to handle various data types.

Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Infrastructure, Capstone Case

Stanford MS&E435 Economics of the AI Supercycle | Spring 2026 | Infrastructure, Capstone Case

Stanford Online

A type of neural network architecture that is fundamental to current AI models. The discussion touches on their compute requirements and the potential for future reinventing or replacing them.

What Happens After A 1,000,000x AI Compute Leap? | Jeff Dean

What Happens After A 1,000,000x AI Compute Leap? | Jeff Dean

Two Minute Papers

A pivotal model architecture in NLP that preceded current large language models. Mentioned as a comparison point for advancements.

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 8 - Trending Topics

Stanford CME296 Diffusion & Large Vision Models | Spring 2026 | Lecture 8 - Trending Topics

Stanford Online

An architecture initially designed for natural language processing tasks like translation, but adapted for vision tasks due to its scalability benefits, forming the basis of Diffusion Transformers and other generative models.

How Exa is Building the Perfect Search Engine | Deep Dives with a16z

How Exa is Building the Perfect Search Engine | Deep Dives with a16z

a16z Deep Dives

A type of neural network architecture that became very good around 2021, enabling the possibility of building better search engines than Google.

Stanford CS25: Transformers United V6 I From Language Models to Native Multimodal Intelligence

Stanford CS25: Transformers United V6 I From Language Models to Native Multimodal Intelligence

Stanford Online

The core architecture underlying modern large language models, responsible for next token prediction.

Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩  and @swyxtv

Podcast Crossover: AIE, AGI, frontier lab strategy with ​ ⁨@matthew_berman⁩ and @swyxtv

Latent Space

A foundational AI architecture that newer custom chips are designed to optimize for, unlike older designs.

Poolside’s Model Factory, Laguna S, Open Models, and the Race to AGI — Eiso Kant, Poolside AI

Poolside’s Model Factory, Laguna S, Open Models, and the Race to AGI — Eiso Kant, Poolside AI

Latent Space

A neural network architecture that outpaced earlier models, mentioned in the context of the speaker's early work on language models.

What Big Tech Missed And How Startups Can Still Win

What Big Tech Missed And How Startups Can Still Win

Y Combinator

The theoretical idea behind LLMs, with papers existing since 2018 but not fully utilized until later by companies like OpenAI.

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Waymo Co-CEO Dmitri Dolgov: The Demo Is Only 1% Of The Work

Y Combinator

A type of AI breakthrough leveraged by Waymo around 2017 for perception and behavior prediction.

Using AI to Increase Your Intelligence & Enrich Humanity | Dr. Fei-Fei Li

Using AI to Increase Your Intelligence & Enrich Humanity | Dr. Fei-Fei Li

Andrew Huberman

A powerful neural network algorithm that quickly showed greater capabilities than earlier ImageNet AlexNet algorithms, particularly in natural language processing.

Accelerating Your AI Career

Accelerating Your AI Career

DeepLearningAI

Ali Miller is excited about transformers and large language models, noting the recent advancements in these areas.

DeepLearning.AI NLP Learner Community ft.Rishit Dholakia

DeepLearning.AI NLP Learner Community ft.Rishit Dholakia

DeepLearningAI

A type of neural network architecture fundamental to modern NLP, discussed as a key area for skill development.

PreviousPage 2 of 2