Emergence of belief-like representations through reinforcement learning
Fig 4
Value RNN activity readout was correlated with beliefs and could be used to decode hidden states.
A. Example observations, states, beliefs, and Value RNN activity from the same Task 2 trials shown in Fig 2. States and beliefs are colored as in Fig 2, with black indicating ITI microstates, and other colors indicating ISI microstates. Note that the states following the second odor observation remain in the ITI (black) because the second trial is an omission trial. Bottom traces depict the linear transformation of the RNN activity that comes closest to matching the beliefs. Total variance explained (R2) is calculated on held-out trials. B. Total variance of beliefs explained (R2), on held-out trials, using different trained and untrained Value RNNs, in both tasks. Same conventions as Fig 3D. C. In purple, the cross-validated log-likelihood of linear decoders trained to estimate true states using RNN activity. Same conventions as Fig 3D. Black circle indicates the log-likelihood when using the beliefs as the decoded state estimate (i.e., no decoder is “trained”).