CentraleSupélec — Reinforcement Learning
Reinforcement Learning — CentraleSupélec
A rework of the PARL course, structured as four lectures and four notebook tutorials, each tutorial picking up the topic of the lecture before it.
Material in preparation. The slide decks and notebooks are not published yet; they will be added here as they are released.
Lectures
| # | Session | Topic |
|---|---|---|
| 1 | Introduction and MDPs | Markov decision processes, return, value functions, Bellman equations |
| 2 | Dynamic programming | Policy evaluation, policy iteration, value iteration |
| 3 | Model-free prediction | Monte-Carlo methods, temporal differences TD(0) |
| 4 | Model-free control | Monte-Carlo control, SARSA, Q-learning |
Tutorials
Four notebooks mirroring the four lectures: dynamic programming, value iteration, Monte-Carlo, temporal differences. They share a common library of environments and plotting helpers.
See also the Deep Reinforcement Learning training I run for companies — same fundamentals, geared towards hands-on practice with Stable-Baselines3 and Gym.