NYU CS-GY 9223 L: Language Models, Reinforcement Learning, Reasoning (Fall 2026)

Professor Pavel Izmailov

Thursday 2:00-4:30pm, Jacobs Hall, 6 Metrotech, Room 204

Pavel Izmailov
Professor
Alex N. Wang
Course Assistant

Course Description

This is a graduate seminar-style course on modern large language models (LLMs), reinforcement learning (RL), and reasoning. The goal is to bring students to the frontier of the field: we will read and discuss contemporary research papers and technical reports, and work through the engineering and scientific decisions behind today's state-of-the-art systems.

Evaluation is primarily project-based. There are no written exams.

Professor Contact

Email: pi390@nyu.edu

Course Assistants

Alex N. Wang

Links

Lectures

Date Topic Main Reading Additional Reading
09/10/2026
Slides
Recording
• LLM basics Attention Is All You Need GPT-2: Language Models are Unsupervised Multitask Learners
GPT-3: Language Models are Few-Shot Learners
09/17/2026 • RL foundations: MDPs
• Policy gradients
• PPO
Sutton & Barto — Reinforcement Learning: An Introduction PPO — Proximal Policy Optimization Algorithms
DQN — Playing Atari with Deep Reinforcement Learning
GAE — Generalized Advantage Estimation
09/24/2026 • Scaling laws Kaplan et al. — Scaling Laws for Neural Language Models
Chinchilla — Training Compute-Optimal Large Language Models
Scaling Data-Constrained Language Models
10/01/2026 • AlphaGo / AlphaZero AlphaZero — Mastering Chess and Shogi by Self-Play MuZero — Mastering Atari, Go, Chess and Shogi with a Learned Model
10/08/2026 • Attention variants and MoE DeepSeek-V3 (architecture sections only) GQA — Grouped-Query Attention
DeepSeekMoE
10/15/2026 • Positional encoding and long context RoPE — RoFormer
YaRN — Efficient Context Window Extension
RULER — What's the Real Context Size of Your Long-Context Language Models?
10/22/2026 • Optimizers and training dynamics Muon is Scalable for LLM Training (Moonlight) muP — Tensor Programs V
Keller Jordan — Muon (blog post)
10/29/2026 • RLHF and preference optimization InstructGPT — Training Language Models to Follow Instructions
DPO — Direct Preference Optimization
Constitutional AI
Tülu 3
11/05/2026 • DeepSeek-R1
• GRPO
• RLVR
DeepSeek-R1 DeepSeekMath — GRPO derivation
Kimi k1.5
Chain-of-Thought Prompting (background)
11/12/2026 • Reward hacking and over-optimization Gao, Schulman & Hilton — Scaling Laws for Reward Model Overoptimization Baker et al. — Monitoring Reasoning Models and the Risks of Promoting Obfuscation
Concrete Problems in AI Safety
11/19/2026 • Weak-to-strong generalization Burns et al. — Weak-to-Strong Generalization Turpin et al. — Language Models Don't Always Say What They Think
AI Safety via Debate
12/03/2026 • Kimi K3, full technical report Kimi K3 technical report Kimi Linear (KDA)
Kimi K2
12/10/2026 • World models V-JEPA 2 DreamerV3
GameNGen — Diffusion Models Are Real-Time Game Engines
Genie 3 (DeepMind blog post)