Video summary
MDP Modelling
Main summary
Key takeaways
Main Ideas and Concepts
-
MDPs (Markov Decision Processes) as the full reinforcement learning formulation
- An MDP is defined by:
- States
- Actions
- Transition probabilities (how the state changes after actions)
- Reward structure (rewards associated with events/outcomes)
- Optionally/implicitly a discount factor (\gamma), which affects the solution (some RL textbooks omit it even though it’s inherent).
- An MDP is defined by:
-
Modeling design is about problem specification → MDP specification
- When constructing an MDP from a real-world scenario, you must decide:
- States
- Actions
- Rewards
- You also need transition probabilities to solve the problem (though RL methods may not require explicit transition dynamics later).
- A major theme: you must choose how detailed the model is—simplifying assumptions trade accuracy for tractability.
- When constructing an MDP from a real-world scenario, you must decide:
-
Example: a recycling robot
- The robot moves around to find and collect empty cans, earning rewards for collection.
- The robot has limited battery, so it must trade off:
- Actively collecting cans
- Avoiding running out of battery (costly to physically recharge)
- Alternatively waiting for cans to appear (idle behavior)
- Or recharging when appropriate
Reinforcement Learning Problem Setup (Robot Example)
1) States (battery-focused, simplified)
- The speaker starts with a simplified state representation emphasizing battery charge.
- Battery charge is discretized into:
- High charge
- Low charge
- A “zero charge” state is avoided by using an equivalent penalty outcome:
- If the robot effectively “runs out,” it transitions from low → (penalty + recovery) behavior rather than explicitly modeling a full empty state.
- Discretization choice is a key modeling decision:
- You could model battery as continuous, or discretize more coarsely (e.g., low/medium/high).
- The exact discretization affects when decisions change.
2) Actions
From the simplified states, the actions discussed are:
- Actively seek cans (collecting actively)
- Wait for cans (becoming like an idle bin)
- Recharge
Additional constraints/assumptions:
- Recharge is only allowed/meaningful in the low-charge state
- If already in high, the model would prohibit/ignore recharge as an action (recharging from fully charged doesn’t make sense in the simplified setup).
3) Rewards (event-based, mapped to state/action)
The speaker defines rewards as occurring upon events:
- Collecting a can → +1
- Running out of battery → −3
- Recharging → 0
- No immediate reward, but indirectly enables future collection.
4) Transition Probabilities (what determines the next state)
To compute expected rewards and solve the MDP, the transition probabilities must be specified for each (state, action) pair.
Key probabilities introduced:
- (\alpha): probability of obtaining a can when the robot is actively seeking
- (\beta): probability of obtaining a can when the robot is waiting
- Typically (\alpha > \beta), though real environments can differ (e.g., lunchtime).
Battery transition depends on:
- Current battery state (high vs low)
- Action taken (seek vs wait)
Examples of transition-probability concepts:
- If in high and seek:
- Probability to remain high
- Probability to move to low
- If in low and seek:
- Probability to remain low
- Probability to move to high (possible under modeling choices) or remain low with high probability
- If in high and wait:
- Probability to remain high
- Probability to move to low
- If in low and wait:
- Probability to remain low
- Probability to move to high
For recharge:
- If in low and action recharge:
- Transition to high with probability 1
- Reward for recharging is 0
The speaker emphasizes stochasticity:
Even if you know the power consumption mechanics, real-world uncertainty (ignored details) makes transitions effectively stochastic.
5) Diagramming the MDP
- The speaker proposes drawing a state transition diagram:
- For each state (high/low), show outgoing transitions for each allowed action
- Include probabilities on edges
- Associate expected rewards (e.g., collecting a can yields expected (+\alpha) when seeking from a given state, etc.)
- A key point:
The process of building the diagram reveals what modeling assumptions you made and what you ignored.
Methodology / Checklist (Implied by the Lecture): Construction of an MDP
-
Pick modeling scope and simplify
- Decide what real-world factors to ignore (robot velocity, exact location, etc.)
- Decide what to approximate (e.g., battery discretization)
-
Define the MDP components
- Choose States
- At minimum, battery level (discretized into high/low)
- Optionally include other variables (e.g., number of cans carried), but the lecture simplifies them away
- Choose Actions
- actively seek cans
- wait for cans
- recharge (only meaningful in low-charge state)
- Choose Reward rules
- +1 for each can collected
- −3 for running out of battery
- 0 for recharging
- Choose Transition probabilities
- For every (state, action), define probabilities of moving to each next state
- Define how action choice affects:
- chance of getting a can
- chance of battery draining (high→low) or recovering (if modeled)
- Choose States
-
Ensure event/reward logic aligns with state/action outcomes
- Example: expected reward depends on probability of can collection for that action.
-
Keep the model tractable
- Simplification choices are intentional “trade-offs” between realism and solvability.
Speakers or Sources Featured
- Instructor / lecturer (unnamed)
- Leads the class and discusses reinforcement learning / MDP modeling, referencing RL textbooks.
- RL textbook(s)
- Referenced as a source for the recycling robot example and discussion of how textbooks present (\gamma) and MDP components (not named specifically).