Key Moments
Stanford CME295 Transformers & LLMs | Autumn 2026 | Lecture 2 - Large Language Models
Key Moments
Large language models are evolving rapidly, with decoder-only architectures and Mixture-of-Experts (MoE) becoming dominant for efficiency and performance. However, challenges remain in explainability, training stability, and context window limitations.
Key Insights
The Transformer architecture, initially for machine translation, has evolved into encoder-only (BERT) and decoder-only (GPT) variants, with decoder-only architectures now dominating LLMs.
Mixture-of-Experts (MoE) models, particularly sparse MoE, allow for significantly larger model sizes without proportional increases in computational cost during inference by selectively activating 'experts'.
Rotary Position Embeddings (RoPE) have become the standard for encoding positional information in LLMs, offering better generalization across sequence lengths compared to learned or fixed positional embeddings.
Inference techniques like beam search and sampling (top-p, temperature scaling) are used to generate text, with non-deterministic methods preferred for more human-like and varied outputs.
Techniques like Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) are employed to reduce memory bandwidth requirements during inference by sharing key/value projections across attention heads.
Despite theoretical determinism, practical LLM inference can exhibit non-determinism due to floating-point arithmetic and batching, a phenomenon explored in Horace He's blog and important for reproducible research.
From Transformers to Decoder-Only Architectures
The lecture begins by revisiting the Transformer architecture, a pivotal development that moved away from Recurrent Neural Networks (RNNs) by utilizing self-attention. Initially designed for machine translation with an encoder-decoder structure, the Transformer's components have been adapted for other purposes. Encoder-only models, like BERT, leverage bi-directional self-attention for tasks requiring deep understanding of input text. Conversely, decoder-only models, exemplified by GPT, are designed for autoregressive text generation, predicting tokens sequentially. This decoder-only approach has become the dominant architecture for modern Large Language Models (LLMs) due to its simplicity and effectiveness in generative tasks, treating all problems as text-to-text transformations.
The Rise of Mixture-of-Experts (MoE) for Scalability
A significant challenge in scaling LLMs is the computational cost associated with their massive parameter counts. Traditional dense models activate all parameters for every input, leading to prohibitive inference costs. Mixture-of-Experts (MoE) models offer a solution by employing a 'router' network that directs input tokens to specific 'expert' sub-networks (often feed-forward layers). Sparse MoE, where only a subset (e.g., top-K) of experts are activated, allows models to have billions or even trillions of parameters while keeping inference costs manageable. This selective activation means that while the total parameter count is huge, the active parameter count per token is much smaller. Training MoE models is complex, requiring auxiliary loss functions like 'load balancing loss' to prevent 'routing collapse,' where the router over-specializes on a few experts, leading to inefficient utilization of the total model capacity. Research indicates that effectively scaling MoE involves carefully balancing the number of experts and the number of activated experts per token.
Positional Encoding: From Additive to Rotational Embeddings
The self-attention mechanism in Transformers inherently lacks information about the position of tokens within a sequence. Early solutions, like learned positional embeddings added to token embeddings, struggled with generalization to sequence lengths unseen during training. The original Transformer paper also proposed fixed, sinusoidal positional encodings based on frequency decomposition, analogous to clock hands, which offered better generalization. However, current state-of-the-art LLMs predominantly use Rotary Position Embeddings (RoPE). RoPE integrates positional information by rotating query and key vectors in the attention mechanism based on their position. This approach results in an attention score that is a function of the difference between token positions, offering excellent generalization and avoiding the need to learn or add separate positional encodings. This method is applied within the attention layers themselves, directly influencing the interaction between tokens.
Attention Variants and Efficiency
The quadratic computational complexity of self-attention (O(n^2) with sequence length 'n') motivates research into more efficient attention mechanisms. Sliding window attention, or local attention, limits each token's attention to a fixed window of surrounding tokens, reducing complexity. While this might seem to reintroduce RNN-like limitations, stacking multiple layers effectively increases the receptive field. For reducing memory bandwidth during inference, techniques like Grouped-Query Attention (GQA) and Multi-Query Attention (MQA) are employed. These methods share key and value projection matrices across multiple attention heads, significantly reducing the number of parameters that need to be loaded from memory, especially crucial for autoregressive decoding where past keys and values must be cached.
Inference Strategies and Output Generation
Generating text from LLMs involves predicting the next token autoregressively. While simply picking the most probable token (greedy decoding) is deterministic, it can lead to repetitive or suboptimal sequences. Beam search explores multiple high-probability sequences concurrently but remains deterministic. To introduce variability and more natural-sounding outputs, sampling methods are used. Temperature scaling adjusts the probability distribution, making it flatter (higher temperature) for more diverse outputs or sharper (lower temperature) for more focused outputs. Top-p (nucleus) sampling selects from the smallest set of tokens whose cumulative probability exceeds a threshold 'p'. Determinism can be useful for specific tasks requiring reproducibility, like summarization or using LLMs as evaluators, but non-deterministic sampling is generally preferred for creative or conversational applications.
Prompting Techniques and Context Window Limitations
Interacting with LLMs involves crafting prompts. The context window defines the maximum number of tokens (input prompt + generated output) the model can process. While context windows have expanded significantly (e.g., to 1 million tokens), they are not infinite. "Contextual erosion" or "needle in a haystack" phenomena demonstrate that as context windows grow, models can struggle to retrieve information presented early in the prompt. This is partly due to the inherent difficulty of filtering signal from noise and practical implementation details. Techniques like 'chain-of-thought' prompting improve reasoning by having the model generate intermediate steps, enhancing performance on complex tasks. Other methods like 'few-shot learning' and 'prompt tuning' optimize input presentation without modifying model weights.
Normalization and Training Stability
Training stability in Transformers is enhanced through normalization layers. The original Transformer used post-layer normalization (LayerNorm applied after the residual connection and sub-layer computation). Modern practice often favors pre-layer normalization, where normalization is applied before the residual connection and sub-layer, which can improve gradient flow. RMS Norm is another popular, simpler normalization technique that scales activations and uses a single learnable parameter, offering similar training benefits to LayerNorm with less computational overhead.
Mentioned in This Episode
●Software & Apps
●Companies
●Books
●Concepts
Common Questions
الخطوات الأساسية لمعالجة النصوص تتضمن الترميز (Tokenization) لتقسيم النص إلى وحدات، ثم حساب التضمينات (Embeddings) لهذه الرموز. هذه التضمينات تحول الكلمات إلى تمثيلات رقمية يمكن للنموذج فهمها ومعالجتها.
Topics
Mentioned in this video
شركة ذُكرت في سياق ورقة بحثية حول نموذج Mixtral of Experts.
نموذج سابق لحساب التضمينات لم يكن يدرك السياق.
نموذج لغوي يعتمد على فك التشفير فقط من المحولات، ولا يتطلب ضبطًا دقيقًا بالضرورة لحالات الاستخدام المختلفة.
نموذج لغوي يعتمد على المشفر فقط من المحولات، ويستخدم لمهام مثل استخراج المشاعر.
نموذج لغوي يعتمد على معمارية Mixture of Experts، ويظهر توزيعًا متنوعًا لتفعيل الخبراء.
طريقة لإضافة انحياز خطي إلى طبقة الانتباه يعتمد على المسافة بين الرموز.
مشكلة تحدث عند تدريب نماذج MoE حيث يميل النموذج للاعتماد بشكل مفرط على مجموعة صغيرة من الخبراء.
تقنية لزيادة استقرار المخرجات عن طريق إنشاء مسارات استنتاج متعددة والتصويت بالأغلبية للوصول إلى إجابة.
تمثيلات عددية ذات معنى للرموز أو الكلمات، يتم حسابها بعد الترميز.
دورة أكاديمية يتم تقديم هذه المحاضرة كجزء منها في جامعة ستانفورد.
آلية تعلم الشبكات العصبية التي تسمح للموجه في نماذج MoE بمعرفة متى يجب تنشيط كل خبير.
طريقة شائعة لترميز الموضع في المحولات عن طريق تدوير متجهات الاستعلام والمفتاح كدالة في موضعها.
آلية الانتباه المستخدمة في المحولات، حيث يتم إجراء عملية الانتباه عدة مرات بشكل متوازٍ.
آلية انتباه تسمح للرمز بالتفاعل فقط مع نفسه والرموز التي تسبقه، وتستخدم في نماذج فك التشفير فقط.
نوع من خليط الخبراء ينشط جميع الخبراء ولكنه يعطي أوزانًا أكبر للخبراء الأكثر صلة.
طريقة لإضافة معلومات الموقع إلى تضمينات الرموز في المحولات، وقد تكون متعلمة أو حتمية.
طريقة لتطبيق التطبيع بعد الطبقة الفرعية وإضافتها إلى المدخلات في بنية المحولات الأصلية.
تقنية لتوليد النصوص تحافظ على K من المسارات الأكثر احتمالًا في كل خطوة لتعزيز الاحتمالية المشتركة الكلية للتسلسل.
عملية إرسال طلب (prompt) إلى نموذج لغوي كبير لتوجيه سلوكه وتوليد المخرجات المطلوبة.
هيكلية نماذج التعلم العميق التي تعتمد على آلية الانتباه، والتي هي أساس نماذج اللغة الكبيرة.
مهمة تدريب بديلة أخرى استخدمها مؤلفو BERT لتدريب النموذج على فهم كيفية تسلسل الجمل.
تقنية انتباه تحد من نطاق التفاعل بين الرموز إلى نافذة معينة، مما يقلل من المتطلبات الحسابية.
بنية نماذج تسمح بتفعيل جزء فقط من النموذج لكل مدخل، مما يقلل من تكاليف الحوسبة مع زيادة حجم النموذج.
ورقة بحثية بارزة تنقل التعلم ضمن السياق إلى مرحلة متقدمة من خلال تعليم النموذج كيفية الوصول إلى النتيجة خطوة بخطوة.
طريقة بديلة لتطبيع الطبقة لا تتطلب إعادة تمركز المدخلات، وتوفر في المعاملات المتعلمة.
تقنية تحول الأمثلة إلى تعليمات في الموجه، وعادة ما تتم صياغتها يدويًا.
عملية التنبؤ بالرموز واحدًا تلو الآخر، وهي الطريقة التي تعمل بها نماذج GPT.
طريقة لتحسين كفاءة الذاكرة في طبقات الانتباه عن طريق مشاركة مصفوفات إسقاط المفتاح والقيمة عبر مجموعات من الرؤوس.
تقنية لأخذ العينات من توزيع الاحتمالات لتوليد الرموز، حيث يتم اختيار الرموز التي لها توزيع احتمالي تراكمي معين.
الحد الأقصى لعدد الرموز التي يمكن للنموذج معالجتها كمدخل في مرة واحدة.
خوارزمية ترميز شائعة توفر توازنًا جيدًا بين حجم المفردات وطول تسلسل النص المدخل.
فئة من الشبكات العصبية كانت شائعة قبل ظهور المحولات، ولكنها كانت تعاني من مشكلة نسيان المعلومات القديمة.
نوع من خليط الخبراء يختار فقط أفضل K من الخبراء لتفعيلهم، مما يوفر في الحوسبة.
طريقة حديثة لتطبيق التطبيع قبل إدخال المدخلات إلى الطبقة، مما يحسن تدفق التدرج.
استخدام عدد قليل من الأمثلة في الموجه لتعليم النموذج تنسيق وسلوك معين.
بنية المحول الأصلي التي تتكون من مشفر (Encoder) لترميز المدخلات ومفكك تشفير (Decoder) لتوليد المخرجات.
جزء من دالة الخسارة يهدف إلى منع انهيار التوجيه من خلال تشجيع توزيع الرموز الموجهة بالتساوي عبر الخبراء.
دالة تستخدم في طبقة الانتباه لتحويل الدرجات إلى احتمالات.
تقنية لتحسين أداء النموذج دون تغيير الأوزان، عن طريق تقديم أمثلة وحلول مباشرة في الموجه (prompt).
مهمة تدريب بديلة تستخدم لتدريب نماذج مثل BERT، حيث يتم إخفاء بعض الرموز في المدخلات ويتوقع النموذج الرموز الأصلية.
رموز تستخدم في نماذج الاستنتاج لتمثيل عملية التفكير أو الاستدلال.
More from Stanford Online
View all 148 summaries
76 minStanford AA274A Principles of Robotic Autonomy | Autumn 2019 | Overview, Mobile Robot Kinematics
75 minStanford AA274A Principles of Robotic Autonomy | Autumn 2019 | Trajectory Optimization
78 minStanford AA274A Principles of Robotic Autonomy | Autumn 2019 | Advanced Trajectory Optimization
79 minStanford AA274A Principles of Robotic Autonomy | Autumn 2019 | Trajectory Tracking
Ask anything from this episode.
Save it, chat with it, and connect it to Claude or ChatGPT. Get cited answers from the actual content — and build your own knowledge base of every podcast and video you care about.
Get Started Free