Maynard Smith revisited: A multi-agent reinforcement learning approach to the coevolution of signalling behaviour
Fig 6
Q-values Case 3, darker line is average across all, faint lines are average for each state.
Q-values Case 3, darker line is average across all, faint lines are average for each state.