Key Moments

Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 2 - Large Language Models

Stanford OnlineStanford Online
Education5 min read104 min video
Oct 9, 2026|5,762 views|117|5
Save to Pod
TL;DR

Large language models are evolving rapidly, with decoder-only architectures and Mixture-of-Experts (MoE) becoming dominant for efficiency and performance. However, challenges remain in explainability, training stability, and context window limitations.

Key Insights

1

The Transformer architecture, initially for machine translation, has evolved into encoder-only (BERT) and decoder-only (GPT) variants, with decoder-only architectures now dominating LLMs.

2

Mixture-of-Experts (MoE) models, particularly sparse MoE, allow for significantly larger model sizes without proportional increases in computational cost during inference by selectively activating 'experts'.

3

Rotary Position Embeddings (RoPE) have become the standard for encoding positional information in LLMs, offering better generalization across sequence lengths compared to learned or fixed positional embeddings.

4

Inference techniques like beam search and sampling (top-p, temperature scaling) are used to generate text, with non-deterministic methods preferred for more human-like and varied outputs.

5

Techniques like Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) are employed to reduce memory bandwidth requirements during inference by sharing key/value projections across attention heads.

6

Despite theoretical determinism, practical LLM inference can exhibit non-determinism due to floating-point arithmetic and batching, a phenomenon explored in Horace He's blog and important for reproducible research.

From Transformers to Decoder-Only Architectures

The lecture begins by revisiting the Transformer architecture, a pivotal development that moved away from Recurrent Neural Networks (RNNs) by utilizing self-attention. Initially designed for machine translation with an encoder-decoder structure, the Transformer's components have been adapted for other purposes. Encoder-only models, like BERT, leverage bi-directional self-attention for tasks requiring deep understanding of input text. Conversely, decoder-only models, exemplified by GPT, are designed for autoregressive text generation, predicting tokens sequentially. This decoder-only approach has become the dominant architecture for modern Large Language Models (LLMs) due to its simplicity and effectiveness in generative tasks, treating all problems as text-to-text transformations.

The Rise of Mixture-of-Experts (MoE) for Scalability

A significant challenge in scaling LLMs is the computational cost associated with their massive parameter counts. Traditional dense models activate all parameters for every input, leading to prohibitive inference costs. Mixture-of-Experts (MoE) models offer a solution by employing a 'router' network that directs input tokens to specific 'expert' sub-networks (often feed-forward layers). Sparse MoE, where only a subset (e.g., top-K) of experts are activated, allows models to have billions or even trillions of parameters while keeping inference costs manageable. This selective activation means that while the total parameter count is huge, the active parameter count per token is much smaller. Training MoE models is complex, requiring auxiliary loss functions like 'load balancing loss' to prevent 'routing collapse,' where the router over-specializes on a few experts, leading to inefficient utilization of the total model capacity. Research indicates that effectively scaling MoE involves carefully balancing the number of experts and the number of activated experts per token.

Positional Encoding: From Additive to Rotational Embeddings

The self-attention mechanism in Transformers inherently lacks information about the position of tokens within a sequence. Early solutions, like learned positional embeddings added to token embeddings, struggled with generalization to sequence lengths unseen during training. The original Transformer paper also proposed fixed, sinusoidal positional encodings based on frequency decomposition, analogous to clock hands, which offered better generalization. However, current state-of-the-art LLMs predominantly use Rotary Position Embeddings (RoPE). RoPE integrates positional information by rotating query and key vectors in the attention mechanism based on their position. This approach results in an attention score that is a function of the difference between token positions, offering excellent generalization and avoiding the need to learn or add separate positional encodings. This method is applied within the attention layers themselves, directly influencing the interaction between tokens.

Attention Variants and Efficiency

The quadratic computational complexity of self-attention (O(n^2) with sequence length 'n') motivates research into more efficient attention mechanisms. Sliding window attention, or local attention, limits each token's attention to a fixed window of surrounding tokens, reducing complexity. While this might seem to reintroduce RNN-like limitations, stacking multiple layers effectively increases the receptive field. For reducing memory bandwidth during inference, techniques like Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) are employed. These methods share key and value projection matrices across multiple attention heads, significantly reducing the number of parameters that need to be loaded from memory, especially crucial for autoregressive decoding where past keys and values must be cached.

Inference Strategies and Output Generation

Generating text from LLMs involves predicting the next token autoregressively. While simply picking the most probable token (greedy decoding) is deterministic, it can lead to repetitive or suboptimal sequences. Beam search explores multiple high-probability sequences concurrently but remains deterministic. To introduce variability and more natural-sounding outputs, sampling methods are used. Temperature scaling adjusts the probability distribution, making it flatter (higher temperature) for more diverse outputs or sharper (lower temperature) for more focused outputs. Top-p (nucleus) sampling selects from the smallest set of tokens whose cumulative probability exceeds a threshold 'p'. Determinism can be useful for specific tasks requiring reproducibility, like summarization or using LLMs as evaluators, but non-deterministic sampling is generally preferred for creative or conversational applications.

Prompting Techniques and Context Window Limitations

Interacting with LLMs involves crafting prompts. The context window defines the maximum number of tokens (input prompt + generated output) the model can process. While context windows have expanded significantly (e.g., to 1 million tokens), they are not infinite. "Contextual erosion" or "needle in a haystack" phenomena demonstrate that as context windows grow, models can struggle to retrieve information presented early in the prompt. This is partly due to the inherent difficulty of filtering signal from noise and practical implementation details. Techniques like 'chain-of-thought' prompting improve reasoning by having the model generate intermediate steps, enhancing performance on complex tasks. Other methods like 'few-shot learning' and 'prompt tuning' optimize input presentation without modifying model weights.

Normalization and Training Stability

Training stability in Transformers is enhanced through normalization layers. The original Transformer used post-layer normalization (LayerNorm applied after the residual connection and sub-layer computation). Modern practice often favors pre-layer normalization, where normalization is applied before the residual connection and sub-layer, which can improve gradient flow. RMS Norm is another popular, simpler normalization technique that scales activations and uses a single learnable parameter, offering similar training benefits to LayerNorm with less computational overhead.

Common Questions

الخطوات الأساسية لمعالجة النصوص تتضمن الترميز (Tokenization) لتقسيم النص إلى وحدات، ثم حساب التضمينات (Embeddings) لهذه الرموز. هذه التضمينات تحول الكلمات إلى تمثيلات رقمية يمكن للنموذج فهمها ومعالجتها.

Topics

Mentioned in this video

Concepts
Alibi

طريقة لإضافة انحياز خطي إلى طبقة الانتباه يعتمد على المسافة بين الرموز.

Routing Collapse

مشكلة تحدث عند تدريب نماذج MoE حيث يميل النموذج للاعتماد بشكل مفرط على مجموعة صغيرة من الخبراء.

self-consistency

تقنية لزيادة استقرار المخرجات عن طريق إنشاء مسارات استنتاج متعددة والتصويت بالأغلبية للوصول إلى إجابة.

embeddings

تمثيلات عددية ذات معنى للرموز أو الكلمات، يتم حسابها بعد الترميز.

CS 295

دورة أكاديمية يتم تقديم هذه المحاضرة كجزء منها في جامعة ستانفورد.

Backpropagation

آلية تعلم الشبكات العصبية التي تسمح للموجه في نماذج MoE بمعرفة متى يجب تنشيط كل خبير.

RoPE

طريقة شائعة لترميز الموضع في المحولات عن طريق تدوير متجهات الاستعلام والمفتاح كدالة في موضعها.

Multi-head attention

آلية الانتباه المستخدمة في المحولات، حيث يتم إجراء عملية الانتباه عدة مرات بشكل متوازٍ.

Causal Attention

آلية انتباه تسمح للرمز بالتفاعل فقط مع نفسه والرموز التي تسبقه، وتستخدم في نماذج فك التشفير فقط.

Dense MoE

نوع من خليط الخبراء ينشط جميع الخبراء ولكنه يعطي أوزانًا أكبر للخبراء الأكثر صلة.

Positional Embeddings

طريقة لإضافة معلومات الموقع إلى تضمينات الرموز في المحولات، وقد تكون متعلمة أو حتمية.

Post-norm

طريقة لتطبيق التطبيع بعد الطبقة الفرعية وإضافتها إلى المدخلات في بنية المحولات الأصلية.

Beam Search

تقنية لتوليد النصوص تحافظ على K من المسارات الأكثر احتمالًا في كل خطوة لتعزيز الاحتمالية المشتركة الكلية للتسلسل.

Prompting

عملية إرسال طلب (prompt) إلى نموذج لغوي كبير لتوجيه سلوكه وتوليد المخرجات المطلوبة.

Transformer

هيكلية نماذج التعلم العميق التي تعتمد على آلية الانتباه، والتي هي أساس نماذج اللغة الكبيرة.

Next Sentence Prediction

مهمة تدريب بديلة أخرى استخدمها مؤلفو BERT لتدريب النموذج على فهم كيفية تسلسل الجمل.

Sliding Window Attention

تقنية انتباه تحد من نطاق التفاعل بين الرموز إلى نافذة معينة، مما يقلل من المتطلبات الحسابية.

Mixture-of-Experts

بنية نماذج تسمح بتفعيل جزء فقط من النموذج لكل مدخل، مما يقلل من تكاليف الحوسبة مع زيادة حجم النموذج.

Chain of Thought

ورقة بحثية بارزة تنقل التعلم ضمن السياق إلى مرحلة متقدمة من خلال تعليم النموذج كيفية الوصول إلى النتيجة خطوة بخطوة.

RMS Norm

طريقة بديلة لتطبيع الطبقة لا تتطلب إعادة تمركز المدخلات، وتوفر في المعاملات المتعلمة.

Prompt Tuning

تقنية تحول الأمثلة إلى تعليمات في الموجه، وعادة ما تتم صياغتها يدويًا.

Auto-regressive Decoding

عملية التنبؤ بالرموز واحدًا تلو الآخر، وهي الطريقة التي تعمل بها نماذج GPT.

Grouped Query Attention

طريقة لتحسين كفاءة الذاكرة في طبقات الانتباه عن طريق مشاركة مصفوفات إسقاط المفتاح والقيمة عبر مجموعات من الرؤوس.

Top P

تقنية لأخذ العينات من توزيع الاحتمالات لتوليد الرموز، حيث يتم اختيار الرموز التي لها توزيع احتمالي تراكمي معين.

context window

الحد الأقصى لعدد الرموز التي يمكن للنموذج معالجتها كمدخل في مرة واحدة.

Subword Tokenization

خوارزمية ترميز شائعة توفر توازنًا جيدًا بين حجم المفردات وطول تسلسل النص المدخل.

Recurrent Neural Networks

فئة من الشبكات العصبية كانت شائعة قبل ظهور المحولات، ولكنها كانت تعاني من مشكلة نسيان المعلومات القديمة.

Sparse MoE

نوع من خليط الخبراء يختار فقط أفضل K من الخبراء لتفعيلهم، مما يوفر في الحوسبة.

Pre-norm

طريقة حديثة لتطبيق التطبيع قبل إدخال المدخلات إلى الطبقة، مما يحسن تدفق التدرج.

Few-shot Learning

استخدام عدد قليل من الأمثلة في الموجه لتعليم النموذج تنسيق وسلوك معين.

Encoder-Decoder

بنية المحول الأصلي التي تتكون من مشفر (Encoder) لترميز المدخلات ومفكك تشفير (Decoder) لتوليد المخرجات.

Auxiliary Load Balancing Loss

جزء من دالة الخسارة يهدف إلى منع انهيار التوجيه من خلال تشجيع توزيع الرموز الموجهة بالتساوي عبر الخبراء.

softmax

دالة تستخدم في طبقة الانتباه لتحويل الدرجات إلى احتمالات.

In-Context Learning

تقنية لتحسين أداء النموذج دون تغيير الأوزان، عن طريق تقديم أمثلة وحلول مباشرة في الموجه (prompt).

Masked Language Modeling

مهمة تدريب بديلة تستخدم لتدريب نماذج مثل BERT، حيث يتم إخفاء بعض الرموز في المدخلات ويتوقع النموذج الرموز الأصلية.

Thinking Tokens

رموز تستخدم في نماذج الاستنتاج لتمثيل عملية التفكير أو الاستدلال.

More from Stanford Online

View all 148 summaries

Ask anything from this episode.

Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.

Get Started Free