Skip to main content
Advertisement

< Back to Article

Learning with sparse reward in a gap junction network inspired by the insect mushroom body

Fig 4

Total reward during each episode in Taxi-v3 task for the dynamic routing model, using an infinite step limit per episode.

The reward here is calculated with the inclusion of negative reward per step, although only the positive reward at the final step is used in training the model. Blue line: Episode reward. Yellow line: 100 episode average reward. (A) Episode reward per episode. (B) Episode reward (same data as A) but plotted against the steps making up each episode (which differ in duration) to show how reward changes with time. Inset plots are zoomed in regions (changed y-axis) of outer plots, showing how the reward level stabilises around -5.

Fig 4

doi: https://doi.org/10.1371/journal.pcbi.1012086.g004