NYU CS-GY 9223 L: Language Models, Reinforcement Learning, Reasoning (Fall 2026)

Professor Pavel Izmailov

Thursday 2:00-4:30pm, Jacobs Hall, 6 Metrotech, Room 204
Virtual lectures: Zoom.

Pavel Izmailov
Professor
Alex N. Wang
Course Assistant

Course Description

This is a graduate seminar-style course on modern large language models (LLMs), reinforcement learning (RL), and reasoning. The goal is to bring students to the frontier of the field: we will read and discuss contemporary research papers and technical reports, and work through the engineering and scientific decisions behind today's state-of-the-art systems.

Evaluation is primarily project-based. There are no written exams.

Professor Contact

Email: pi390@nyu.edu

Course Assistants

Alex N. Wang

Links

Lectures

Date Topic Main Reading Additional Reading
09/10/2026
Slides
Recording
• LLM basics • Attention Is All You Need • GPT-2: Language Models are Unsupervised Multitask Learners
• GPT-3: Language Models are Few-Shot Learners
09/17/2026
Slides
Recording
• RL foundations: MDPs
• Policy gradients
• PPO
• Sutton & Barto — Reinforcement Learning: An Introduction (Chapters 1, 3, 4, 5) • Lilian Weng — A (Long) Peek into Reinforcement Learning
• OpenAI Spinning Up — Part 1: Key Concepts in RL, Part 2: Kinds of RL Algorithms, Part 3: Intro to Policy Optimization
• David Silver — Reinforcement Learning lecture series
09/24/2026
Slides
Demo
Recording
• Scaling laws • Kaplan et al. — Scaling Laws for Neural Language Models • Chinchilla — Training Compute-Optimal Large Language Models
• Scaling Data-Constrained Language Models
• The Art of Scaling Reinforcement Learning Compute for LLMs
• Understanding Reasoning from Pretraining to Post-Training
10/01/2026
Slides
Demo
Recording
• AlphaGo / AlphaZero • AlphaZero — Mastering Chess and Shogi by Self-Play • MuZero — Mastering Atari, Go, Chess and Shogi with a Learned Model
10/08/2026 • Attention variants and MoE • DeepSeek-V3 (architecture sections only) • GQA — Grouped-Query Attention
• DeepSeekMoE
10/15/2026 • Positional encoding and long context • RoPE — RoFormer
• YaRN — Efficient Context Window Extension
• RULER — What's the Real Context Size of Your Long-Context Language Models?
10/22/2026 • Optimizers and training dynamics • Muon is Scalable for LLM Training (Moonlight) • muP — Tensor Programs V
• Keller Jordan — Muon (blog post)
10/29/2026 • RLHF and preference optimization • InstructGPT — Training Language Models to Follow Instructions
• DPO — Direct Preference Optimization
• Constitutional AI
• Tülu 3
11/05/2026 • DeepSeek-R1
• GRPO
• RLVR
• DeepSeek-R1 • DeepSeekMath — GRPO derivation
• Kimi k1.5
• Chain-of-Thought Prompting (background)
11/12/2026 • Reward hacking and over-optimization • Gao, Schulman & Hilton — Scaling Laws for Reward Model Overoptimization • Baker et al. — Monitoring Reasoning Models and the Risks of Promoting Obfuscation
• Concrete Problems in AI Safety
11/19/2026 • Weak-to-strong generalization • Burns et al. — Weak-to-Strong Generalization • Turpin et al. — Language Models Don't Always Say What They Think
• AI Safety via Debate
12/03/2026 • Kimi K3, full technical report • Kimi K3 technical report • Kimi Linear (KDA)
• Kimi K2
12/10/2026 • World models • V-JEPA 2 • DreamerV3
• GameNGen — Diffusion Models Are Real-Time Game Engines
• Genie 3 (DeepMind blog post)