What is the difference between Q-learning and SARSA?

Question

When I was learning this part, I found it very confusing too, so I put together the two pseudo-codes from R.Sutton and A.G.Barto hoping to make the difference clearer.

enter image description here

Blue boxes highlight the part where the two algorithms actually differ. Numbers highlight the more detailed difference to be explained later.

TL;NR:

|             | SARSA | Q-learning |
|:-----------:|:-----:|:----------:|
| Choosing A' |   π   |      π     |
| Updating Q  |   π   |      μ     |

where π is a ε-greedy policy (e.g. ε > 0 with exploration), and μ is a greedy policy (e.g. ε == 0, NO exploration).

Given that Q-learning is using different policies for choosing next action A’ and updating Q. In other words, it is trying to evaluate π while following another policy μ, so it’s an off-policy algorithm.
In contrast, SARSA uses π all the time, hence it is an on-policy algorithm.

More detailed explanation:

The most important difference between the two is how Q is updated after each action. SARSA uses the Q’ following a ε-greedy policy exactly, as A’ is drawn from it. In contrast, Q-learning uses the maximum Q’ over all possible actions for the next step. This makes it look like following a greedy policy with ε=0, i.e. NO exploration in this part.
However, when actually taking an action, Q-learning still uses the action taken from a ε-greedy policy. This is why “Choose A …” is inside the repeat loop.
Following the loop logic in Q-learning, A’ is still from the ε-greedy policy.

Leave a Comment Cancel reply