Exploration by optimism: act greedily with respect to the value estimate plus an uncertainty bonus.
… current value estimate of action
… number of times has been selected so far
… controls the degree of exploration
The square-root term measures the uncertainty in ‘s value estimate. Each time is selected, grows and its uncertainty (and degree of exploration) shrinks; each step is not selected, grows while stays put, so its bonus slowly rises until it gets tried again.
The use of the natural logarithm means that the increases get smaller over time, but are unbounded; all actions will eventually be selected, but actions with lower value estimates, or that have already been selected frequently, will be selected with decreasing frequency over time.