<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Companies | Etienne Chassaing</title><link>http://etiennechassaing.com/teaching/companies/</link><atom:link href="http://etiennechassaing.com/teaching/companies/index.xml" rel="self" type="application/rss+xml"/><description>Companies</description><generator>Hugo Blox Builder (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sat, 01 Feb 2025 00:00:00 +0000</lastBuildDate><image><url>http://etiennechassaing.com/media/sharing.jpg</url><title>Companies</title><link>http://etiennechassaing.com/teaching/companies/</link></image><item><title>PID and MPC control training</title><link>http://etiennechassaing.com/teaching/companies/pid-mpc/</link><pubDate>Sat, 01 Feb 2025 00:00:00 +0000</pubDate><guid>http://etiennechassaing.com/teaching/companies/pid-mpc/</guid><description>&lt;p>&lt;img src="featured.png" alt="Principle of MPC">
&lt;em>Principle of Model Predictive Control&lt;/em>&lt;/p>
&lt;p>Training delivered in &lt;strong>February 2025 at &lt;a href="https://nehemis.com">Nehemis&lt;/a>&lt;/strong> for its
&lt;strong>software teams&lt;/strong>. The goal: give developers who are not control engineers what they need
to understand, tune and debug a control loop, then position Model Predictive Control
relative to a PID: when it brings something extra, and at what cost.&lt;/p>
&lt;p>The program below is the one that was delivered; it can be adapted to your system and
the time available.&lt;/p>
&lt;h2 id="part-1-understanding-a-dynamic-system">Part 1: Understanding a Dynamic System&lt;/h2>
&lt;ul>
&lt;li>Open-loop and closed-loop systems: what feedback changes.&lt;/li>
&lt;li>Modeling a simple process, transfer functions, poles and stability.&lt;/li>
&lt;li>Reading a step response: static gain, time constant, overshoot, settling time.&lt;/li>
&lt;li>Experimental identification of a process from real measurements.&lt;/li>
&lt;/ul>
&lt;h2 id="part-2-the-pid-controller-in-practice">Part 2: The PID Controller in Practice&lt;/h2>
&lt;ul>
&lt;li>The role of each term (proportional, integral, derivative) and what each one costs.&lt;/li>
&lt;li>Steady-state error, disturbance rejection, stability margin.&lt;/li>
&lt;li>Tuning methods: empirical approach, Ziegler-Nichols, model-based tuning.&lt;/li>
&lt;li>The implementation pitfalls that make a PID that is correct on paper fail in practice:
integrator windup (&lt;em>anti-windup&lt;/em>), actuator saturation, noise amplified by the
derivative term and how to filter it.&lt;/li>
&lt;li>Discretization: sampling period, recursive form of the controller, effects of
computation time and delays.&lt;/li>
&lt;/ul>
&lt;h2 id="part-3-introduction-to-model-predictive-control">Part 3: Introduction to Model Predictive Control&lt;/h2>
&lt;ul>
&lt;li>Principle: optimize a sequence of control inputs over a receding horizon rather than
reacting to the instantaneous error.&lt;/li>
&lt;li>Formulating the optimization problem: prediction model, cost function, prediction and
control horizons.&lt;/li>
&lt;li>What MPC can do that PID cannot: &lt;strong>explicit constraints&lt;/strong> on inputs and states,
&lt;strong>multivariable&lt;/strong> (MIMO) systems, anticipating a setpoint known in advance.&lt;/li>
&lt;li>The price to pay: a model is required, computational cost, sensitivity to modeling
errors.&lt;/li>
&lt;li>Selection criteria: when a well-tuned PID is still the right answer.&lt;/li>
&lt;/ul>
&lt;h2 id="part-4-applications">Part 4: Applications&lt;/h2>
&lt;ul>
&lt;li>Hands-on work on a system representative of the participants&amp;rsquo; field.&lt;/li>
&lt;li>PID vs. MPC comparison on the same process: performance, constraint satisfaction,
robustness.&lt;/li>
&lt;li>Discussion of internal use cases and possible next steps.&lt;/li>
&lt;/ul>
&lt;p>This training is offered in French or English, on site or remotely, and runs from one to
three days depending on how deep you want to go into MPC.&lt;/p></description></item><item><title>Deep Reinforcement Learning training (5 days)</title><link>http://etiennechassaing.com/teaching/companies/drl/</link><pubDate>Thu, 03 Oct 2024 00:00:00 +0000</pubDate><guid>http://etiennechassaing.com/teaching/companies/drl/</guid><description>&lt;p>
&lt;figure >
&lt;div class="flex justify-center ">
&lt;div class="w-100" >&lt;img alt="Agent-environment loop: the agent&amp;amp;rsquo;s policy produces actions, the environment returns observations and rewards" srcset="
/teaching/companies/drl/featured_hu6aa22660cb161ba72f620666ae4b7356_150020_db0698252c36dfb6b85666ceb8af70b2.webp 400w,
/teaching/companies/drl/featured_hu6aa22660cb161ba72f620666ae4b7356_150020_cd0430aaa62ad2fc17ba109e79964c3e.webp 760w,
/teaching/companies/drl/featured_hu6aa22660cb161ba72f620666ae4b7356_150020_1200x1200_fit_q80_h2_lanczos_3.webp 1200w"
src="http://etiennechassaing.com/teaching/companies/drl/featured_hu6aa22660cb161ba72f620666ae4b7356_150020_db0698252c36dfb6b85666ceb8af70b2.webp"
width="760"
height="428"
loading="lazy" data-zoomable />&lt;/div>
&lt;/div>&lt;/figure>
&lt;em>The agent-environment loop in reinforcement learning&lt;/em>&lt;/p>
&lt;p>A five-day training in &lt;strong>deep reinforcement learning&lt;/strong> that alternates theory and practice
with &lt;strong>Stable-Baselines3&lt;/strong> and &lt;strong>OpenAI Gym&lt;/strong>: lectures in the morning, hands-on work in
the afternoon, and a final project on an environment taken from your own field.&lt;/p>
&lt;p>The program below is a template; it can be adapted to your needs and the time available.&lt;/p>
&lt;h2 id="day-1-reinforcement-learning-basics">Day 1: Reinforcement learning basics&lt;/h2>
&lt;p>&lt;strong>Morning&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>What RL is, and how it differs from supervised and unsupervised learning&lt;/li>
&lt;li>The building blocks: agent, environment, states, actions, rewards&lt;/li>
&lt;li>Real applications: games, robotics, finance&lt;/li>
&lt;li>The vocabulary: policy, state value, action value, reward, return; Markov decision
processes (MDPs)&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Afternoon&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Tabular methods: Monte Carlo, temporal difference (TD), SARSA, Q-learning&lt;/li>
&lt;li>Their strengths, limits and use cases&lt;/li>
&lt;li>&lt;em>Hands-on:&lt;/em> implement a simple Q-learning agent on FrozenLake (OpenAI Gym)&lt;/li>
&lt;/ul>
&lt;h2 id="day-2-openai-gym-environments">Day 2: OpenAI Gym environments&lt;/h2>
&lt;p>&lt;strong>Morning&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>The role of the environment in training an agent&lt;/li>
&lt;li>Standard environments: CartPole, MountainCar…&lt;/li>
&lt;li>Working with an environment: &lt;code>env.reset()&lt;/code>, &lt;code>env.step()&lt;/code>, &lt;code>env.render()&lt;/code>&lt;/li>
&lt;li>Action and observation spaces&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Afternoon&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Building your own environment with the Gym API&lt;/li>
&lt;li>&lt;em>Hands-on:&lt;/em> a simple game with a custom reward logic&lt;/li>
&lt;/ul>
&lt;h2 id="day-3-stable-baselines3">Day 3: Stable-Baselines3&lt;/h2>
&lt;p>&lt;strong>Morning&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Why SB3: reliable implementations of DQN, PPO, A2C…&lt;/li>
&lt;li>Installation and setup&lt;/li>
&lt;li>Plugging a Gym environment into SB3, a first PPO run on CartPole&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Afternoon&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>&lt;em>Hands-on:&lt;/em> tune a PPO agent on CartPole&lt;/li>
&lt;li>Callbacks to steer training: early stopping, metric logging&lt;/li>
&lt;/ul>
&lt;h2 id="day-4-advanced-algorithms">Day 4: Advanced algorithms&lt;/h2>
&lt;p>&lt;strong>Morning: Deep Q-Network (DQN)&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>From tabular methods to neural networks&lt;/li>
&lt;li>Replay buffer, target network&lt;/li>
&lt;li>&lt;em>Hands-on:&lt;/em> DQN with SB3 on LunarLander&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Afternoon: Actor-Critic (A2C, PPO)&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>The actor-critic principle&lt;/li>
&lt;li>What PPO brings over DQN: sample efficiency, stability&lt;/li>
&lt;li>&lt;em>Hands-on:&lt;/em> PPO on a continuous environment, BipedalWalker&lt;/li>
&lt;/ul>
&lt;h2 id="day-5-deployment-and-a-project-from-your-field">Day 5: Deployment and a project from your field&lt;/h2>
&lt;p>&lt;strong>Morning&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>Saving and reloading a trained agent&lt;/li>
&lt;li>Evaluating it on episodes it never saw in training&lt;/li>
&lt;li>Integrating it into a real application: robots, simulations, games&lt;/li>
&lt;/ul>
&lt;p>&lt;strong>Afternoon: final project&lt;/strong>&lt;/p>
&lt;ul>
&lt;li>A complete project on an environment taken from your field: modelling, choice of
algorithm, training, tuning and evaluation&lt;/li>
&lt;li>Wrap-up: the difficulties met and where to go next&lt;/li>
&lt;/ul></description></item></channel></rss>