Recurrent neural networks that learn multi-step visual routines with reinforcement learning
Fig 4
RELEARNN for an example stimulus.
A. Correct choice by the network. In the first phase, activity propagates in the regular network that has both feedforward and recurrent connections (squares) until this network reaches a stable state. Here enhanced activity (yellow) spreads over the target curve. In the second phase, a winning unit is selected and activates the corresponding unit of the accessory network (small circles). From there, activity propagates in the accessory network (small orange circles) to tag the connections that influence the Q-value of the chosen action. After a few timesteps, activity in the accessory unit xiacc becomes proportional to the influence of the activity of the corresponding regular unit xi∞ on the chosen output unit Qa. In the third phase, a reward is given if the action was correct, or not in case of an error, and a neuromodulator (green cloud, δ) broadcasts the reward prediction error to the network. Weights are changed according to a four-factor Hebbian learning rule (green connections between the units are increased). B. Incorrect choice by the network. In this case, the enhanced activity spreads over the wrong curve and reward prediction error is negative because of the wrong choice (red cloud). Hence, the weights between units that represent the distractor curve are decreased (red connections).