Value Functions

1 min read

Value functions estimate how good it is to be in a state (or to take an action in a state) under a given policy’

State-value function Vπ(s)V^\pi(s):

Vπ(s)=Eπ[Gt∣St=s]=Eπ[∑k=0∞γkRt+k+1∣St=s]V^\pi(s) = \mathbb{E}_\pi[G_t | S_t = s] = \mathbb{E}_\pi\left[\sum_{k=0}^\infty \gamma^k R_{t+k+1} \bigg| S_t = s\right]

"Expected return starting from state ss and following policy π\pi."

Action-value function Qπ(s,a)Q^\pi(s,a):

Qπ(s,a)=Eπ[Gt∣St=s,At=a]Q^\pi(s,a) = \mathbb{E}_\pi[G_t | S_t = s, A_t = a]

"Expected return starting from state ss, taking action aa, then following π\pi."

Relationship: Vπ(s)=∑aπ(a∣s)Qπ(s,a)V^\pi(s) = \sum_a \pi(a|s) Q^\pi(s,a) — the value of a state is the expected Q-value over actions.

Optimal value functions:

  • V∗(s)=max⁡πVπ(s)V^*(s) = \max_\pi V^\pi(s) — best possible value
  • Q∗(s,a)=max⁡πQπ(s,a)Q^*(s,a) = \max_\pi Q^\pi(s,a) — best possible Q-value
  • Optimal policy: π∗(a∣s)=1\pi^*(a|s) = 1 if a=arg⁡max⁡aQ∗(s,a)a = \arg\max_a Q^*(s,a) — just be greedy w.r.t. Q∗Q^*

Key insight: if you know Q’Q^’ , the optimal policy is trivial — just take the action with highest QQ-value. This is why Q-Learning focuses on learning Q’Q^’.

See also: Bellman Equations , Markov Decision Process , Advantage Function

Linked from