Key Moments

Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 1 - Transformers

Stanford OnlineStanford Online
Education5 min read105 min video
Sep 28, 2026|162,761 views|2,718|68
Save to Pod
TL;DR

The Transformer architecture, introduced in 2017, revolutionized NLP by enabling models to scale dramatically with more data and compute, powering today's LLMs and chatbots.

Key Insights

1

The Transformer architecture, introduced in the 2017 paper 'Attention Is All You Need,' is a scalable breakthrough that underpins modern Large Language Models (LLMs).

2

Text is converted into numerical representations through tokenization, with subword tokenization (like Byte Pair Encoding or BPE) being the most common method today due to its balance of vocabulary size and sequence length.

3

Word2vec (CBOW and Skip-gram) were early methods to learn meaningful word embeddings, but they struggled with polysemy (words with multiple meanings) and did not consider word order.

4

Recurrent Neural Networks (RNNs) and their variant Long Short-Term Memory (LSTM) addressed word order but suffered from long-range dependency issues and slow sequential processing.

5

The core innovation of the Transformer is the self-attention mechanism, allowing each token to directly attend to all other tokens in the sequence, overcoming RNN limitations.

6

The Transformer consists of an encoder to process the source text and a decoder to generate the target text, with positional encodings added to token embeddings to account for word order.

Course Overview and Logistics

The course CME295: Transformers & LLMs at Stanford aims to demystify Large Language Models (LLMs), explaining their underlying architecture, training, usage, and limitations. Targeting a broad audience, from aspiring researchers to professionals, the course emphasizes the growing importance of AI literacy. Prerequisites include basic linear algebra and machine learning concepts. Lectures are held weekly and are also recorded. The course includes a midterm and a final exam, with past exams available on the website. A continuously updated summary document and a textbook are provided as resources.

Evolution of NLP Models: From RNNs to Attention

Before Transformers, Natural Language Processing (NLP) models typically used a single model for each specific task, with Recurrent Neural Networks (RNNs) being dominant in the early 2010s. These models processed text sequentially, leading to limitations. The year 2017 marked a turning point with the publication of 'Attention Is All You Need,' which introduced the Transformer architecture. This architecture demonstrated remarkable scalability, meaning performance significantly improved with more data, computational power, and parameters. This scalability paved the way for modern LLMs capable of generating text and code with high efficiency. The subsequent release of models like ChatGPT in late 2022 further solidified the impact of these advancements, making interactions with conversational AI commonplace.

Tokenization: Converting Text to Numbers

Models understand numbers, not text, so the first step in processing text is tokenization. This involves breaking down text into smaller, indivisible units called tokens. The choice of tokenization method impacts the vocabulary size and sequence length. Word-level tokenization, where spaces separate tokens, is simple but can lead to large vocabularies due to word variations (e.g., 'bear' vs. 'bears') and struggles to represent similar words consistently. Character-level tokenization results in a small, robust vocabulary but creates very long sequences, increasing computational complexity. Subword tokenization, such as Byte Pair Encoding (BPE), strikes a balance by breaking words into common subword units, leveraging word roots and reducing vocabulary size while mitigating the risk of out-of-vocabulary words. BPE is currently the most prevalent method.

Early Embedding Techniques: Word2vec and its Limitations

Early approaches like Word2vec (using CBOW and Skip-gram models) aimed to learn meaningful numerical representations (embeddings) for words. These methods trained models on intermediate tasks, like predicting a word from its context or vice versa. The learned embeddings captured semantic similarities, allowing for analogies like 'Paris is to France as Berlin is to Germany.' However, Word2vec had significant limitations: it assigned a single embedding to each word, failing to capture polysemy (words with multiple meanings like 'bank'), and it did not inherently account for word order, treating sentences as bags of words. This meant 'the dog chased the cat' and 'the cat chased the dog' would have similar representations.

Recurrent Neural Networks (RNNs) and LSTMs

To address the word order limitation, Recurrent Neural Networks (RNNs) were developed. RNNs process sequences by maintaining a hidden state that summarizes the information seen so far. This allows them to consider word order. However, standard RNNs suffer from the vanishing gradient problem, making it difficult to capture long-range dependencies – information from many steps back in the sequence. Long Short-Term Memory (LSTM) networks were introduced to mitigate this by incorporating a cell state designed to retain information over longer periods. Despite improvements, LSTMs still faced challenges with very long sequences and sequential processing was inherently slow.

The Attention Mechanism: Direct Connections

The attention mechanism revolutionized NLP by allowing models to directly link a current processing step to relevant parts of the input sequence, bypassing the limitations of sequential processing and fixed hidden states. Instead of relying solely on an RNN's hidden state, attention enables the model to 'look back' at specific input tokens. This is crucial for tasks like translation, where a word in the target language might strongly correspond to a specific word in the source language, regardless of their positions. The paper 'Attention Is All You Need' introduced self-attention, a powerful form of attention where each token attends to all other tokens in the same sequence.

Self-Attention: The Core of the Transformer

Self-attention calculates a representation for each token by considering its relationship with all other tokens in the sequence. This is achieved using Query (Q), Key (K), and Value (V) vectors. Each token generates a Q, K, and V vector. The Q vector of a token is compared against the K vectors of all tokens (including itself) to determine attention weights (how much focus to put on each token). These weights are then used to compute a weighted sum of the V vectors, producing the new representation for the token. This mechanism allows the model to weigh the importance of different tokens dynamically based on the context, overcoming the long-range dependency issues of RNNs and enabling parallel computation across tokens.

The Transformer Architecture: Encoder-Decoder Structure

The Transformer architecture, originally proposed for machine translation, consists of an encoder and a decoder. The encoder processes the input sequence (e.g., English sentence) to create context-aware representations. It comprises multiple layers, each containing a self-attention sub-layer and a feed-forward neural network sub-layer. Positional encodings are added to the input embeddings to inject information about the position of each token, as self-attention itself is permutation-invariant. The decoder generates the output sequence (e.g., French translation) autoregressively. It also uses self-attention, but it's 'masked' to prevent attending to future tokens. Additionally, it incorporates a cross-attention layer that allows it to attend to the encoder's output, enabling it to align the source and target sequences. Both encoder and decoder layers often include residual connections and layer normalization to aid training.

Common Questions

يهدف مقرر CME 295 إلى شرح كيفية عمل نماذج اللغة الكبيرة (LLMs)، وفهم بنيتها الأساسية، وكيفية تدريبها واستخدامها، بالإضافة إلى نقاط قوتها وضعفها. يهدف إلى تزويد الطلاب بفهم شامل لهذه التقنيات.

Topics

Mentioned in this video

More from Stanford Online

View all 142 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free