Skip to main content
Advertisement
  • Loading metrics

Flexible navigation with neuromodulated cognitive maps

  • Krubeal Danieli,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Software, Validation, Visualization, Writing – original draft, Writing – review & editing

    Affiliation Center for Integrative Neuroplasticity, FYSCELL, University of Oslo, Oslo, Norway

  • Mikkel Elle Lepperød

    Roles Funding acquisition, Project administration, Resources, Supervision, Validation, Writing – review & editing

    mikkel@simula.no

    Affiliation Simula Research Laboratory, Oslo, Norway

Abstract

Animals develop specialized cognitive maps during navigation, constructing environmental representations that facilitate efficient exploration and goal-directed planning. The hippocampal CA1 region is implicated as the primary neural substrate for cognitive mapping, housing spatially tuned cells that adapt based on behavioral patterns and internal states. Computational approaches to modeling these biological systems have employed various methodologies. Although labeled graphs with local spatial information and deep neural networks have provided computational frameworks for spatial navigation, significant limitations persist in modeling one-shot adaptive mapping. We introduce a biologically inspired place cell architecture that develops cognitive maps during exploration of novel environments. Our model implements a simulated agent for reward-driven navigation that forms spatial representations online. The architecture incorporates behaviorally relevant information through neuromodulatory signals that respond to environmental boundaries and reward locations. Learning combines rapid Hebbian plasticity, lateral competition, and targeted modulation of place cells. Analysis of the model across a variety of environments demonstrates that online map formation and reward-directed navigation can emerge within a single simulated trial, without the multi-epoch training typically required by reinforcement-learning approaches. The simulation results show that the agent successfully explores and navigates to target locations in various environments, adapting when reward positions change. Analysis of neuromodulated place cells reveals dynamic changes in neuronal density and place field size after behaviorally significant events. These findings align with experimental observations of reward effects on hippocampal spatial cells while providing computational support for the efficacy of biologically inspired approaches to cognitive mapping.

Author summary

This work proposes a biologically inspired architecture for neuromodulated cognitive maps that addresses key limitations in current approaches, such as offline training, biologically unrealistic learning, and reliance on external coordinate inputs. The model constructs spatial representations through place cells generated dynamically during navigation, with learning mediated by neuromodulatory signals that integrate boundaries and reward-related information. Simulations show that the model supports navigation, flexible path planning, and robustness to environmental changes. The analysis reveals adaptive strategies including context-dependent regulation of place cell density and place field scaling, aligned with experimental findings from hippocampal CA1. This proof-of-concept highlights how plausible neural dynamics and computationally advantageous strategies can emerge from simple modeling assumptions.

1. Introduction

Survival in complex environments requires efficient navigational strategies. From desert ants to humans, successful wayfinding, defined as navigating toward goals that are not directly visible, depends on emergent internal spatial representations, known as cognitive maps [1,2]. Understanding how such maps are constructed from ongoing experiences and how they can be exploited for flexible goal-directed navigation remains an active area of research in both neuroscience and reinforcement learning.

In the brain, the hippocampus (HP) and the entorhinal cortex (EC) are the central areas involved in spatial representation. They contain specialized neurons encoding spatial and contextual information, including grid, border, speed, and place cells [35]. In particular, the latter are numerous in the CA1 hippocampal sub-region and have attracted interest due to the convergence of inputs from entorhinal grid cells, the CA3 hippocampal sub-region, and the lateral EC [69]. This strategic integration of diverse spatial and contextual signals suggests that CA1 place cells may play a critical role in the formation and maintenance of cognitive maps [10].

Traditional theories of cognitive maps suggest that spatial representations can emerge from multiple navigational strategies. Early frameworks suggested that the hippocampus encodes both spatial location and direction [11], while graph-based models captured structural aspects of spatial organization, although often at the cost of scalability [12].

Another approach is path integration, which involves tracking one’s position by integrating past movements—using actions or velocity vectors—to estimate the trajectory and return to the point of origin.

An important strength of path integration lies in its independence from external cues, relying instead on idiothetic internal-motion signals, which are considered biologically plausible and believed to involve entorhinal grid cells [13,14]. A common computational methodology to investigate path integration involves recurrent neural networks combined with linear readouts. When trained in navigation tasks, these systems spontaneously develop spatially tuned activity patterns reminiscent of grid, place, and border cells [1517].

Another proposal is the Tolman-Eichenbaum machine (TEM), a model that generalizes across spatial and relational tasks while reproducing key biological neural activity patterns [1820]. Learning in TEM occurs through gradient descent over many episodes or epochs, in contrast to adaptive, context-dependent learning observed in biological systems. As with several modeling approaches, its neuronal dynamics remain simplified, typically reduced to artificial neurons defined by weighted sums and fixed non-linear activation functions. An alternative is the successor representation framework (SR), built around the learning of a predictive map of the explored state space, which can be used for active navigation [21,22]. It has been linked to hippocampal cognitive map generation, with the possibility of incorporating reward information and formulation in terms of biologically plausible mechanisms [2326].

Neuromodulation is another relevant element in brain dynamics, affecting physiology, cognition, and behavior [27]. Its functions include filtering meaningful internal and external signals, influencing neuronal dynamics, and learning by altering synaptic states and parameters [2830]. An important neuromodulator is dopamine, long associated with reward information, prediction error encoding [30,31], and novelty detection [29], mechanisms closely related to reinforcement learning principles [32,33]. In addition, projections from the Ventral Tegmental Area (VTA) and Locus Coeruleus (LC) have been shown to target the hippocampal circuits [3436], to be involved in memory formation [37], and affect the adjustment of CA1 place cells [3841], particularly through LEC inputs [42,43].

Other approaches have incorporated neuromodulators with reward-driven Hebbian plasticity [44] or other learning features, such as sparsity and update scaling [45]. However, these architectures often remain limited in their ability to integrate these components into a biologically grounded framework that works without extensive training, gradient-based optimization, or external coordinate inputs while still providing adaptable planning and behavior.

In this work, we addressed these limitations by proposing a bio-inspired model of cognitive map formation learned online from sensory experiences. The model builds on the emergence of place cell representations from grid cells activity and velocity information, that is, internally available idiothetic input, in line with the path integration approach [15,16]. We designed plasticity mechanisms through which neuromodulatory signals could target and modulate place cells, augmenting the resulting map with behaviorally relevant information. Combined with a network-based path planning mechanism, our architecture enabled an artificial agent to construct content-rich topological maps of its environment on-the-fly and to leverage them for targeted navigation. We took direct inspiration from previous work discussing the involvement of neuromodulation in hippocampal spatial processing, for example, in terms of the gating of incoming input streams and the influence of synaptic plasticity [4648].

Our primary aim was to evaluate the potential of the emergent cognitive map to support navigation in unexplored environments, including the ability to record the location of a reward and to generate paths to it from already visited regions. Notably, the model begins without any spatial or qualitative knowledge of its surroundings; such knowledge is acquired incrementally as the agent moves. The lack of separation between map construction and reward collection ensures that the task operates entirely online.

A second aim was to assess the model’s robustness and adaptability to qualitative and spatial changes, such as shifts in reward location or the introduction of new boundaries. Additionally, the contribution of neuromodulation to the regulation of place cell activity was also examined as experience accumulated, investigating the effect on cell density and the size of place fields. Finally, we sought to quantify the effect of neuromodulation on task performance through a series of comparative ablation experiments.

These analyses allowed us to characterize the effects of neuromodulation on place field properties, such as modulation of their size and density, and to discuss potential links to experimental findings [49,50].

The remainder of the paper is organized as follows: Section 2 details the model and experimental setup; Section 3 presents results; Section 4 discusses the broader implications and future directions.

2. Methods

We propose a model of cognitive map formation driven by an agent’s experience within a closed environment.

The architecture operates with minimal external inputs, limited to binary reward and collision signals, as illustrated in Fig 1A. Instead of relying on exteroceptive cues, spatial representations emerge from idiothetic information, that is, the agent’s internal perception of self-motion [51], consistent with path-integration frameworks. In practice, we use the agent’s ground truth velocity vector, that is, its actual displacement within the environment, as the primary navigational signal, reflecting the integration of inertial and proprioceptive signals observed in biological systems [13,52]. Since no visual information is used, the agent effectively navigates in the dark. Importantly, the model does not receive allocentric coordinates as sensory input; any two-dimensional positions used for planning or visualization are derived from the internally constructed place cell representation.

thumbnail
Fig 1. Model layout and spatial representations.

a: layout of the architecture consisting of four main inputs: a reward signal targeting the dopamine modulation component (DA), a collision signal targeting the boundary modulation component (BND), a velocity vector (the actual movement made by the agent) aiming at the grid cells module, and a reward trigger acting as a flag to enable reward-directed navigation. A place cells component represents where a cognitive map actually resides, and receives input from the grid cells for spatial information, modulators for non-spatial data, and it is used for navigation and determining the reward position. Two policy components are responsible for determining the agent’s action: an explorative behaviour and a reward-directed behaviour, the latter being conditioned on the ability to calculate the reward position and the presence of a trigger signal. b: visualization of part of the path-finding algorithm, propagation of an activity wave through the place cell network from top-left to bottom-center, and the calculated path visualization in the bottom-right. c: a module of grid cells defined in a bounded square space of length 1, and an activity representation of their receptive field over a torus. d: formation of a cognitive map during navigation, with place cells (grey) being formed over the agent’s trajectory (black line); boundary-modulated place cells are formed near boundaries (blue). e: spatial distribution of the place cell centers, together with the place fields of two cells. f: representation of the cognitive map as place cells tagged by neuromodulators during collisions (blue), reward events (green), or untagged (grey); the color gradient correlates with the intensity of the neuromodulators’ weights. g: illustration of a path (arrows) over a place cell map connecting an initial position (red) and a goal (green).

https://doi.org/10.1371/journal.pcbi.1013487.g001

2.1. Place cell formation

The primary spatial representation is formed by a set of simplified grid cell modules, each encoding a periodic tiling of 2D space, which directly maps to a toroidal manifold (Fig 1C). Departing from traditional grid cell modeling approaches [53,54], we generate population activity directly by Gaussian tuning on a torus, continuously updated using the agent velocity vector, an approach used in previous work [8].

The grid cell population vector is forwarded to a place cell network with initial zero synaptic weights. When no place cell is sufficiently active for a given input, a silent unit is randomly selected and imprinted with the current grid activity pattern. The rationale is to form a population of place cells as the agent moves around, with the centers placed on the trajectory (Fig 1D). To enforce sparsity in the population vector and clear spatial tuning specificity, a lateral inhibition mechanism compares the cosine similarity between the new weight vector and the existing non-zero weights against a threshold . If it is above threshold, the new place cell is discarded.

The activation of each place cell is calculated using a bounded cosine similarity function, which determines how closely the cell’s tuning curve matches the current grid-cell activity, followed by a generalized sigmoid function. The latter acts as a parametrized filter with being the gain and the bias, responsible for the sensitivity and thresholding of the input, affecting the shape of the resulting place field (Fig 1E).

Additionally, for each neuron, an activity trace m is recorded with time constant , representing the decay of activation over time. A compact summary of the principal model and task parameters is provided in S1 Table. Further details of the implementation, including lateral inhibition and recurrent connectivity, are provided in the Supplementary Information.

2.2. Neuromodulation

Neuromodulators deliver event signals: rewards, denoted DA (for dopamine), and boundary collisions, denoted BND. In particular, the latter is an internal variable that correlates with and tracks the sensory experience of mechanical collisions.

They are driven by binary inputs and defined through a leaky variable with exponential decay.

Each modulator k updates its connection weights with the place cells through a Hebbian rule based on place cell activity and the error:

(1)

The term in brackets can be interpreted as an error term that implements a simple form of predictive coding, and is inspired by temporal-difference learning [33], aligning with evidence that neuromodulatory systems signal prediction errors and update beliefs [19,55,56]. This mechanism was intended to support robustness to environmental changes, such as reward displacement or the introduction of new obstacles.

The weight vectors are restricted to remain non-negative. Reward modulation tags cells near reward locations, whereas boundary modulation builds a representation of the environmental edges. The result is a set of scalar fields over the place cells, and it forms the core of our model of a cognitive map (Fig 1F). Throughout the manuscript, the architecture is understood as containing only place cells: labels such as reward-modulated (or DA-tagged) and boundary-modulated (or BND-tagged) refer to subsets of place cells carrying non-zero modulatory weights, not to distinct neuron types. See the Supplementary Information for full learning rules and parameter settings.

2.3. Modulation of place fields

In addition, we tested whether neuromodulators could directly alter spatial tuning. The place fields were dynamically shifted and resized according to recent salient events, effectively changing the cell density and the area of the receptive fields.

More in detail, following a reward or collision signal, the place field centers were displaced within the grid-cell space according to a vector whose magnitude depended on the neuromodulator and the spatial proximity to the reward or collision location:

(2)

Here, is a Gaussian function with width and a scaling factor. This rule is inspired by BTSP plasticity [49], which shifts CA1 place fields following salient experiences. This action was applied only to recently active cells, that is, with an activity trace greater than a given threshold . The motivation was to prioritize activity within a recent time window, consistent with suggestions of dopaminergic gating in the hippocampus [57,58]. Furthermore, this approach is consistent with mechanisms such as the BCM rule [59] and the aforementioned BTSP, both of which rely on asymmetric kernels that reduce or filter out inactive neurons [39].

Lateral inhibition prevents field overlap during remapping. Moreover, field size is modulated by scaling the gain of recently active neurons, allowing neuromodulators to transiently enhance or suppress the spatial sensitivity of specific cells. This modulation rule involves the gain of each cell being adjusted proportionally to its activity trace , a reference gain constant , and a modulatory scaling variable.

(3)

where is a scaling gain parameter, for which a value of 1 means absence of modulation.

Finally, these constitute the mechanisms through which place fields could be reshaped, and the cells that were not affected maintained stable and unchanged fields.

2.4. Policy and behavior

To evaluate the navigation ability of the model, we implemented a simple policy that switched between exploration and reward-seeking behavior. The transition between these two modes was based on an external trigger, which indicated whether the reward was ready to be discovered, and the state of the internal map, defined as the presence of non-zero dopaminergic connections. In other words, the purpose of the behavior was dictated by the internal availability of a reward representation; see S2 Fig. Planning was defined as the process of determining a target location, computing a path as a sequence of place cells from the one corresponding to the current position to the one nearest the goal, and executing the actions required to traverse such a sequence. The paths were generated using a graph-based pathfinding algorithm in which place cells served as nodes, synaptic connections as edges, and the calculations involved the operation of population vectors, as shown in Fig 1B. To further improve navigation, we introduced a cost function on the graph nodes, assigning lower values to boundary-modulated cells. This bias discouraged, but did not forbid, paths that approach environmental boundaries, as illustrated in Fig 1G.

Exploration consisted of two possible strategies: a random walk, for purely stochastic movements, and periodic goal-directed navigation toward a randomly selected visited location, aimed at preventing stagnation.

In contrast, exploitation, defined as reward-directed navigation, involved identifying the location of the reward within the cognitive map. This location corresponded to the average position of DA-modulated place cells, reflecting mechanisms such as hippocampal replay and value-based navigation [6062].

Lastly, wall-avoidance and goal-planning together constitute a way for the model to functionally distinguish and treat reward and collision signals.

3. Results

The main objective of this work was to assess the model’s ability to construct an adaptive spatial representation that supports effective goal-directed behavior.

To evaluate the model, we generated a set of closed environments with an increasing number of internal walls, thereby varying the difficulty of mapping and navigation. The internal arrangement of walls and corners renders these environments non-convex. They thus qualify as a wayfinding setting [63], where the goal location is hidden from the agent’s immediate line of sight. At each time step, the agent’s selected action updated its position. A binary collision signal was generated (1 if the new coordinates intersected or crossed a wall, and 0 otherwise). Similarly, the reward signal was defined relative to the goal location, delivering a value of 1 with a probability of when the agent entered a circular region of radius around the reward, and 0 otherwise.

A simulation trial was defined by initializing the model without any place cell or learned weights and placing the agent randomly. To facilitate map analysis by increasing the average number of place cells formed, the reward position was initially hidden, becoming available only after a fixed duration. Once available, each time a reward was collected, the agent was randomly teleported to a new location; concurrently, the model’s grid cell modules were re-calibrated to keep the internal map and newly recruited place cells aligned.

We defined the performance metric within a trial as the total number of rewards obtained over a duration Tdur.

This evaluation of map formation and performance within the same trial reflects the online nature of the model, in contrast to approaches based on distinct multi-epoch training and testing stages.

Following these simulation settings and parameter optimization, the remainder of this section is organized around three main questions. First, we investigate whether the model can form a cognitive map online to navigate a representative environment (Figs 2A and 2B). Second, we illustrate how this map is locally refined and updated during specific tasks, including place field modulation, detour behavior, and reward relocation (Figs 2C and 2H). Finally, we quantify these event-dependent changes and establish the functional role of neuromodulation via ablation studies (Figs 3A and 3E).

thumbnail
Fig 2. Cognitive maps and performance results.

a: illustration of a cognitive map composed of boundary-modulated place cells (blue), reward-modulated place cells (green), and a line (red) representing the planned trajectory to reach a target location from the current position. b: representation of the same environment as in panel a but with the reward area shown as a circle (green), together with the past trajectories (red line) and the agent position (black square). - c: a cognitive map in a different environment showing shifts of the place fields for the two types of modulated place cells. Specifically, the starting boundary-modulated place cell centers (magenta circles), final boundary-modulated place cell centers (blue circles), and their displacement vectors with length proportional to the magnitude (blue arrows) are shown. Similarly, they are shown the starting reward-modulated place cell centers (orange circles), final reward-modulated place cell centers (green), and their displacement vectors (green arrows). Finally, there are also non-modulated place cells (grey filled circles). - d: example cognitive map representing place field size modulation. Both reward-modulated place cells (green) and boundary-modulated place cells (blue) are shown as circles whose radius represents the relative width of each cell’s place field, calculated from its gain parameter : larger values correspond to smaller fields. This plot is intended as a qualitative visualization; the corresponding statistical comparison across simulations is reported in Fig 3C. - e: detour experiment, plot of trajectories before (black) and after (red) the insertion of a wall (rectangle) between the starting and goal positions; the wall is also indicated by the boundary-modulated place cells in blue. - f: collision (blue) and reward (green) counts during a simulation trial where a new wall was inserted after 5000 time steps (red dotted line). Curves represent a 1000-step moving average across 64 repetitions, with shaded regions denoting the standard deviation. - g: trajectories across multiple trials in an example environment for the reward-relocation task, with the agent starting at the same position (black square) and the reward location periodically moved among the green circles. - h: same legend and protocol as in plot f but with the reward location changed instead.

https://doi.org/10.1371/journal.pcbi.1013487.g002

thumbnail
Fig 3. Statistics and performance results.

The labels “BND cells” and “DA cells” abbreviate boundary- and reward-modulated place cells, respectively. a: effect of collision count on the number of BND- and DA-modulated place cells (top) and on their average activation gain (bottom). Each sub-plot reports Pearson’s r, R2, and the p-value. - b: same layout as in plot a but using reward count. - c: Dunnett-corrected differences in place field center displacement (top) and average gain magnitude (bottom) for DA-modulated, BND-modulated, and non-modulated control place cells; larger values correspond to narrower place fields. - d: reward-count performance across the chance, no-modulation, full-modulation, DA-only, BND-only, density-only, and gain-only variants. Significant pairwise differences are color-coded as: grey <0.05, black <0.01, and red < 0.001. Results obtained from pairwise t-tests across 128 runs with Bonferroni correction; the numerical report is provided in S4 and S5 Tables - e: same analysis as plot d but for collision count, calculated from the time a reward representation is initially formed (set at the first non-zero DA weights); the collision count of the control model was not considered, since it lacked reward-driven navigation. The numerical results are reported in S6 and S7 Tables.

https://doi.org/10.1371/journal.pcbi.1013487.g003

3.1. Optimization

The model architecture contains 15 hyperparameters, several of which control non-differentiable components of place cell generation, neuromodulation, and planning. Consequently, we optimized the model using the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) [64] instead of gradient-based methods or manual tuning. The fitness function was designed to simultaneously maximize reward collection and minimize collisions across multiple environments of increasing layout complexity. Full parameter distributions and convergence analyses are detailed in S3 Fig.

3.2. Performance in wayfinding

Figs 2A and 2B show a representative simulated trial in which place cells were recruited online during exploration and then used for planned navigation toward the reward. Across the example layouts in Figs 2A and 2D, increasing the number and placement of internal walls could prolong exploration before the cognitive map was sufficiently populated, while also making subsequent route planning more demanding than in simpler layouts.

Fig 2A shows the subsets of place cells tagged by collision (blue) and reward events (green), marking the shape of the boundaries and location of rewards. The composition of these modulated subsets with non-modulated place cells (gray) constitutes what we refer to here as a cognitive map. These elements encode the primary spatial and contextual information used during planned navigation. Fig 2B shows the same environment layout (walls in black), the reward region (green), and several representative trajectories (red).

Importantly, the cognitive map was not precomputed but emerged dynamically within a single trial as place cells were sequentially recruited along the exploration trajectory. In stationary environments, previously formed associations between grid and place cells were preserved, restricting map adaptation to newly explored regions. Salient events nevertheless induced local refinement: collisions and reward encounters modulated the activation gain and shifted the place field centers of recently active, tagged cells. Without these modulatory dynamics, the emergent spatial representation remained stable over time, serving as a reliable instrument for navigation.

To overcome the challenges imposed by non-convex layouts, the model utilized graph-based planning over this emergent representation, relying on local spatial information and boundary-modulated place cells, enabling the agent to generate routes around obstacles. The approximately homogeneous distribution of place field centers further allowed graph distances between place cell nodes to approximate the spatial cost of candidate routes.

Overall, this representative trial illustrates how the model can balance reward-directed navigation and obstacle avoidance in a wayfinding setting. Across simulations, early reward localization remained variable because exploration was stochastic, especially in environments with numerous walls and narrow passages.

3.3. Local modulation of place field structure

We next investigate how neuromodulation alters local place field properties across different environment geometries. To demonstrate the model’s adaptability across distinct structural constraints, the simulations shown in Figs 2C and 2D utilize different environmental layouts than those in panels A and B. Notably, increasing the number and density of internal walls prolonged the initial exploration phase before the cognitive map was sufficiently populated, while simultaneously rendering subsequent route planning more demanding. One important mechanism driving local adaptation is place field center displacement, which is modulated dynamically by experienced salient events (see Eq. 2). Although the evolutionary optimization did not converge to a unique solution in the parameter space, place cell location modulation regularly produced visible changes in field-center position.

Fig 2C illustrates this field remapping at the conclusion of a trial. Specifically, we highlight the displacement vectors between the initial and final positions of boundary-modulated (magenta to blue) and reward-modulated (orange to green) place cell centers, with arrows indicating the direction and magnitude of the shift. For this representative choice of parameters, boundary-modulated place cells were consistently displaced toward the walls. A possible functional interpretation is that this systematic shift elarges representation space near boundaries, thereby promoting the generation of new place cells to cover open internal areas. Field shifts were also prominent within the reward zone, reflecting the high density of salient encounters in that region during the run.

The second neuromodulatory mechanism involves scaling the neural activation gain () via boundary and dopamine-dependent hyperparameters (, ), directly affecting place field size. Fig 2D provides an illustrative example of this scaling, in which the radius of each circle reflects the relative width of the corresponding place field. Because the gain parameter regulates the steepness of the activation function, higher gain values correspond to narrower fields. Thus, the compact fields of the modulated cells in Fig 2D visually capture this gain-dependent field shrinkage, a population-level phenomenon quantified below in Fig 3C.

Fig 2E shows the trajectories before (gray) and after (red) wall placement, illustrating both the disruption of the previous paths and the subsequent formation of new place cell tags near the newly added obstacle. Fig 2F complements this example by reporting the average effect of the layout change on reward and collision counts: rewards decreased and collisions increased markedly after wall insertion. These results were averaged over 64 simulation trials, of which 9 failed to reach the reward, yielding a success rate of approximately 85%, a reflection of the unavoidable variance inherent to noisy exploration.

3.4. Detour task

We next evaluated the flexibility of the model’s planning mechanisms when subjected to unexpected environmental modifications. To this end, we devised a detour task wherein a novel internal wall was introduced mid-trial. We predicted that relying on a cognitive map containing outdated structural information would initially induce planning failures, thereby forcing the agent to resume stochastic exploration until an alternative route was discovered.

As anticipated, the introduction of the wall caused an immediate, sharp increase in collision counts (Figs 2E and 2F). This elevated sensory error triggered the recruitment of a new subset of boundary-modulated place cells, rapidly updating the internal spatial representation. When the existing cognitive map proved insufficient to resolve a detour, the agent reverted to local exploration until a valid path to the reward zone was discovered. Fig 2E illustrates the trajectories before (gray) and after (red) wall placement, capturing both the disruption of previous paths and the subsequent generation of new place cell tags marking the novel obstacle. Fig 2F complements this spatial trace by visualizing the behavioral impact of the layout change over time; across 64 simulation trials, reward rates dropped considerably while collisions spiked immediately following wall insertion. Of these runs, 9 failed to recover the goal within the trial limit, yielding an overall success rate of approximately 85%, a reflection of the unavoidable variance inherent to noisy, online exploration.

3.5. Adaptive goal representation through sensory error

To assess the model’s flexibility to reward displacement as well, we applied a relocation protocol similar to the detour task. Here, the reward location was periodically moved after fixed intervals, requiring the agent to extinguish previous goal associations and discover novel reward zones in a design inspired by computational work [44]. Fig 2G displays representative trajectories as the reward alternated between three distinct locations. The agent consistently transitioned from exploitation and planned navigation at a known site to targeted exploration when a location became unrewarding, as evidenced by the dense overlay of successful navigation paths.

Mechanistically, whenever a planned route resulted in a failed reward prediction, a dopamine-based sensory error signal weakened the synaptic associations between active place cells and the reward representation, driving the rapid extinction of the outdated goal location. Fig 2H tracks the temporal dynamics of this disturbance on performance metrics. Both reward and collision frequencies fluctuated immediately following each relocation event before stabilizing as the agent successfully mapped and exploited the new rewarding zone. These data demonstrate that the architecture effectively recovers from violated sensory expectations.

Several empirical studies support the involvement of dopamine in updating hippocampal spatial representations under shifting reward contingencies. For instance, extinguishing reward expectations reshapes and degrades CA1 spatial maps, an effect replicated by inhibiting ventral tegmental area (VTA) dopaminergic neurons [56]. Likewise, local hippocampal D1/D5 receptor signaling is required for CA1 place cell reorientation following a change in reward-relevant spatial rules [65]. While these findings are broadly consistent with predictive coding and temporal difference learning frameworks [21,33,66], direct empirical evidence for failed reward predictions actively weakening place reward associations precisely as implemented here remains limited. Consequently, we treat our dopamine-based sensory error as a dopamine-inspired, experimentally testable prediction rather than a direct replication of an established hippocampal plasticity rule.

3.6. Event-dependent modulation of place cells

After exploring the individual task examples in Fig 2, we quantified the aggregate impact of neuromodulation on place cell populations across multiple simulations. Figs 3A and 3B summarize the results of 50 independent runs across four environments of increasing structural complexity, with each data point corresponding to a single simulation. These plots map the relationships between behavioral metrics, here collision (Fig 3A) and reward counts (Fig 3B), and network features, namely the final count of modulated cells and their average neural activation gains.

The primary trend pattern is a direct relationship between cell count and environmental events: the population of boundary-modulated (BND-tagged) place cells scaled with collision frequency, while dopamine-modulated (DA-tagged) place cells expanded with reward count. Moreover, we also observed inverse relationships for the cross-pairings: high-reward simulations exhibited fewer BND-tagged cells, while high-collision simulations displayed fewer DA-tagged cells.

Overall, these results are not unexpected. Rather than indicating direct cross-modulation dynamics within the network architecture, these inverse relationships are best interpreted as an indirect consequence of the agent’s behavior. High collision counts typically signify prolonged, stochastic exploration, characterized by repeated boundary hits, or difficulty in planned navigation. These more inefficient runs naturally present fewer opportunities to sample the reward zone, leaving the dopamine-tagged sub-population and its associated gain parameters near baseline levels. Conversely, when the reward is discovered early, it rapidly switches from exploratory wandering to efficient and recurrent reward-directed navigation. This transition promotes reward collection while drastically reducing contacts with the walls, thereby limiting both BND-tagged cell recruitment and BND gain updates. The same behavioral coupling accounts for the gain relationships shown in the lower panels of Fig 3B, where the most prominent trend is a positive correlation between reward count and DA gain. This likely reflects a gradual, cumulative gain modulation induced by reward encounters, contrasting with collision-driven BND updates that tend to saturate rapidly due to the high frequency of wall impacts.

We also quantified the local place field alterations illustrated qualitatively in Figs 2C and 2D. Spatial remapping was statistically significant across both modulated cell populations, as evaluated via a corrected Dunnett’s test across 156 simulations (Fig 3C, top). Reward-modulated place cells exhibited an average field-center displacement of (SEM), whereas boundary-modulated cells underwent a larger displacement of , both measured relative to a non-modulated control group fixed at zero. Furthermore, boundary-modulated cells received the most pronounced modulatory scaling, reaching an average activation gain of (SEM), compared to for reward-modulated cells and for non-modulated controls (Fig 3C, bottom). The tuning properties of the non-modulated control population remained highly stable, displaying only minor, transient fluctuations proximal to salient sensory boundaries although not enough for them to be BND-tagged.

Functionally, this dual-modulation strategy extends spatial remapping by increasing local neuronal representation density. Mechanistically, sharper place fields reduce the lateral inhibition exerted on neighboring units, allowing additional place cells to be recruited within the same local region. Computationally, this density regulation can be interpreted to opportunely increase the spatial resolution of the cognitive map around critical environmental features. By maximizing representation density near boundaries and narrow passages, the model significantly enhances planning accuracy where navigation margins are tight. These population-level adjustments align with recent empirical findings reporting a higher clustering of place fields near points of behavioral relevance or structural interest, such as environmental boundaries and reward sites, reflecting a biological compromise between spatial resolution and total network size [6769].

3.7. Functional effect of modulation on navigation

We next investigated whether experience-driven modulation directly enhances functional aspects of navigation, specifically maximizing reward collection and minimizing collisions. To isolate the individual contributions of these mechanisms, we conducted a series of ablation experiments where the dopamine (DA) and boundary (BND) modulatory components were systematically disabled, separating their effects on place cell density from their effects on neural activation gain. We evaluated each of the six model variants across five environments of increasing structural complexity, running 128 independent simulations per condition. To put the difficulty of reward collection into perspective, we also evaluated a control variant operating under a random-walk policy.

For all learning architectures, reward and collision frequencies were quantified from the exact moment the agent first discovered the reward to isolate navigation efficiency from initial exploration duration. Statistical significance across configurations was then assessed via pairwise t-tests.

The behavioral outcomes illustrated in Figs 3D and 3E show that while all configurations performed above the chance baseline, the magnitude of the performance benefit depended on which modulatory components remained intact. To interpret these results, we focused on specific planned comparisons: the primary test evaluated the full-modulation architecture against the fully non-modulated baseline, while secondary contrasts isolated the independent contributions of reward versus boundary signals (DA-only vs. BND-only) and decoupled the underlying physical mechanisms of adaptation (density-only vs. gain-only).

Our primary comparison revealed that modulation-enabled variants consistently outperformed the control and the fully non-modulated baseline across environments. Despite substantial stochastic variance across individual runs, the non-modulated model harvested significantly fewer rewards.

This global trend suggests that experience-dependent reshaping of the cognitive map provided a behavioral advantage beyond path planning alone.

Among the secondary partial ablations, boundary-related modulation accounted for a noticeably larger share of the navigation benefit than reward-related modulation. This asymmetry suggests that dynamically updating representations of structural walls and narrow corridors is more critical for robust wayfinding than refining representations near known goals, likely because boundary modulation alters the underlying planning graph topology around edges and corners to minimize total path length. Furthermore, the density-modulation variant independently recapitulated most of the performance advantages seen in the full model, whereas gain-only field sharpening yielded weaker and less consistent improvements.

This pattern suggests that the main functional contribution of this neuromodulatory strategy lies not merely in sharpening the tuning curves of an existing population, but in restructuring the spatial density of place cells to better support the planning graph.

Taken together, these ablation studies support the hypothesis that direct, experience-dependent modulation of place field structure can provide a functional advantage for online navigation within complex, non-convex environments.

4. Discussion

Exploration and planning in novel and past environments are essential behaviors of animals that directly affect their success in spatial understanding and achieving goals.

An important element behind these abilities is the formation of a map of their surroundings, enriched with information gained from new experiences, known as a cognitive map. In this work, we presented a rate network model that simulates the CA1 area of the hippocampus [10,70]. It differs from previous approaches as it operates online, uses neuromodulated synaptic plasticity, and does not rely on external coordinates.

We used simplified grid cells and synaptic plasticity to generate information-rich spatial representations rapidly forming in novel environments, dynamics directly following experimental observations and computational modeling of place cell generation [71,72]. The architecture and its mechanics are based on the proposal that CA1 place coding is supported by entorhinal grid cell modules that operate at different spatial frequencies [73,74]. In the spirit of minimizing geometric assumptions in the neural space, we treated the generated place network as a topological graph learned through velocity inputs, reminiscent of path integration [14,75,76]. External sensory information was added locally to each graph node, i.e., a place cell, through the action of neuromodulators acting as a signal filter of salient events such as boundary collisions and reward occurrences. We took inspiration from experimental findings about the involvement of dopaminergic projections in hippocampal representation [65,77]. This idea of enriching a spatial map aligns with the concept of a labeled graph [78,79], a cognitive map proposal that addresses the problems of geometric and metric consistency by harnessing noisy observations and supporting vector-like operations [8082].

Although incorporating a richer set of sensory inputs, such as additional self-motion cues, could further improve performance, our goal in this work was to provide a proof of concept. Specifically, our interest was to document the emergence of advantageous computational strategies even from relatively simple modeling assumptions.

The tasks we applied the agent to consisted of an exploratory and exploitative phase, in which the system was tasked to plan and reach reward positions. For simplicity, the first stage relied on a random walk process, as more elaborate policies were outside the scope of this work. This choice had the side effect that the reward was not always discovered, leading to the formation of incomplete maps and impairing performance. However, this constraint did not qualitatively alter our central findings.

The simulation results support the model’s central assumptions, showing the formation of cognitive maps and their encoding of information collected during experience.

Previous work has used path integration with deep neural networks, but required extensive gradient-based training [13,16,17]. Another important difference is that our resulting neural network was composed solely of place cells, although neuromodulated, and no other types of neurons were present. This distinction is justified by the partially different task protocol and internal architecture, which constrained the tuning dynamics and also did not receive visual information as in [15]. Furthermore, our model relied on predefined grid cell layers, which constituted a strong and sufficient inductive bias, and did not have to be learned from scratch. The selection of the goal location was based on dopaminergic connections and a spatial average of the associated neurons. This approach is consistent with experimental observations and dopamine activity recordings in CA1, suggesting its role in stabilizing long-term spatial memories and supporting reward-guided navigation [60,65,83].

An additional relevant aspect is the consideration of the place cell layer as an explicit graph data structure, on which path-planning was applied, meant to be a neural-based variant of the Dijkstra algorithm [84], a popular choice for graph computations. In fact, our algorithm relied on biologically plausible neural operations, such as activity propagation through the network and the use of synaptic activity traces as masks to identify the best neighboring place cells. This structured representation introduced a strong inductive bias, promoting robustness and flexibility across environments with varying layout complexity. Moreover, this lifted the need to learn an approximation of it through network dynamics and rely on a greater variability in neural receptive fields.

A similar approach using boundary-modulated place cells has also been used for applying the successor representation (SR) framework to large environments [85], with the emergence of place and grid cell tuning. More broadly, on the one hand, several aspects of the SR framework align with our model: the distinction of the state map and the reward/boundary signals, and the use of an update rule inspired by temporal difference learning [2123]. In addition, several studies have investigated biologically plausible implementations [2426], obtaining close and promising matches with experimental results. On the other hand, the construction principles of our architecture required fewer assumptions about the state space size (the environment area can be almost arbitrarily large), while learning the basis of the state representation (place cells) and the successor matrix (associated with the lateral connectivity) is more direct and immediate than in the standard SR formulation. These modeling choices reflect our interest in the online generation of an explicit place cell map with an associated graph representation, while the SR may also produce non-spatial neural profiles. In fact, the type of cognitive map that emerged in our results was solely composed of neurons with a well-defined tuning curve requiring a single grid cell population vector, thus not demanding multiple spatial observations or learning steps, except for adapting the modulation weights to failed predictions. Finally, our architecture was better suited to integrate and study our formulation of neuromodulated plasticity, given direct access to a spatial map.

Adaptability was tested by occasionally moving the reward position, leading to the generation of an internal prediction error that was used to update its representation on the map. The agent was able to weaken previous associations, return to exploration, and encode new reward locations. This behavioral protocol is similar to previous work [86], in which dopaminergic and cholinergic activity was utilized within a Hebbian plasticity rule to strengthen or weaken reward-associated spatial representations. However, as an alternative to exploiting neuromodulators with opposite valence, we followed the direction of temporal-difference learning and predictive coding, a framework linked to hippocampal representations [19,87] and explored by various computational approaches [8890]. This preference arose from our focus on using the data encoded in the cognitive map itself, in which each position has an associated neuromodulation value array, encapsulated in the modulatory connection weights. This array functions as an expected sensory experience to compare with the actual sensory experience to obtain a learning signal. This mechanism is aligned with long-standing theories linking neuromodulation, particularly dopamine, with prediction error signaling [31,38,66,91,92]. It also fits within the larger framework of predictive coding, which interprets cognition as the minimization of surprise [90,93], and matches with temporal-difference learning, where the values of internal states are continuously updated through sensory feedback [19,21,33].

However, a limitation of our current implementation is the strong simplification of neuromodulation to externally defined sensory events represented as a binary two-dimensional input.

An important consequence of this choice is the marked limitation of the information about the environment available to the agent, which is effectively restricted to boolean quantities. This operationalization differs substantially from biology, where signals are encoded in a distributed manner and integrated with multiple pathways, providing neuromodulation circuits with high-dimensional neural representations and richer responses [47,57]. Furthermore, neuromodulation in our model only influences value representation and place field modulation, which captures only a subset of the larger roles neuromodulators play in the brain [34].

Lastly, the relevance of active modulation of the neuronal properties of place cells was supported by simulated ablation experiments. The primary finding was the statistically significant role of active boundary modulation, obtained by comparing it with a model variant with disabled neuromodulation. This supports the importance of having a flexible representation of the environment layout.

Further investigation involved testing the effect of density and gain modulation separately. The results showed that altering the density of the place cells affected the total count of the rewards collected. These results are consistent with experimental observations that demonstrate the remapping of place fields following salient events [68,94]. Notably, behavioral time-scale plasticity (BTSP) has been shown to induce shifts in CA1 place fields in response to salient inputs delivered through pathways from the entorhinal cortex [39,49,72,95]. Additional studies further support this view, reporting spatial reorganization of place cells to better emphasize behaviorally relevant information [96,97], as well as changes in firing rate following contextual stimuli [98,99], or other activity processes [100].

The effect was particularly clear for boundary-modulated place cells, which displayed a trend of shifting place cell centers closer to the walls, yielding a statistically significant improvement in performance. One possible interpretation is that a tighter representation of environmental edges leaves more room for the interior, improving its spatial resolution, and thereby enhancing mapping and supporting more efficient path planning. Consistent with this perspective, previous studies have reported an increase in the clustering of place field edges near surface boundaries [101], and have shown that the presence of visual cue edges stabilizes spatial tuning [102]. It should be noted that in our model, the boundary information was limited solely to collision signals rather than feature-rich sensory input. Nevertheless, the emergence of strategies that parallel biological observations suggests the existence of shared computational advantages, supporting previous work documenting the benefit of increased cell clustering near boundaries for spatial accuracy and context discrimination [69,103].

Regarding reward-modulated place cells, despite the reliable experimental evidence to support their remapping, the effect was weak and did not substantially improve behavior, a pattern that likely followed from the baseline constraints of our simulation settings.

Concerning the modulation of place fields, there is significant experimental evidence of their alteration during reward events [104106], some reporting shrinkage near reward objects [107], and more markedly near boundaries [108]. In this direction, two possible experimental predictions can be made, based on the constraint of using self-motion cues for the generation of cognitive maps.

First, in the model, place field sizes near boundaries are initially approximately homogeneous, and are reshaped only after meaningful events, such as collisions. This can be tested by correlating individual place-field structural metrics with the cumulative count of relevant events occurring near their fields. Secondly, a similar test could be performed for rewarding events, potentially going further by blocking dopamine release in order to assess its causal involvement.

In our setup, the place-field modulation was implemented by scaling the neural activation gain. This led to a general reduction in field size across all modulated cells.

In the case of boundary-modulated place cells, this shrinking effect may account for the improved performance in two ways: first, by increasing the spatial resolution of the environment layout [109]; and second, by freeing up space for other place cells to form nearby. The latter may result from reduced lateral inhibition and the inward shift of the field center toward the boundaries.

However, gain modulation alone did not produce statistically significant effects, suggesting that, in this protocol, it may need to be coupled with density modulation to produce measurable improvements.

For reward place cells, the impact of gain modulation was less pronounced. This may be due to the limits of our experimental protocol or reward signal, with the latter being defined purely in spatial terms and lacking any meaningful non-spatial feature.

Previous studies have proposed that the Locus Coeruleus (LC) is involved in place cell reorganization [41] and overrepresentation near reward locations through the co-release of dopamine and noradrenaline [40]. A possible experimental prediction that emerges from our results is that preventing place cell density modulation may indirectly decrease the quality of reward-directed navigation, as suggested by the remapping of boundary-modulated cells shown in Fig 2C and the statistics reported in Fig 3D.

More specifically, this can be tested in animal models by pharmacological blockade of LC afferent connections to CA1, once a reward has been experienced. Then, performance can be evaluated by considering: the count of rewards fetched; the difference between the animal’s trajectory between the initial and reward location and the theoretical spatial geodesic (shortest path); and the average distance from the walls. However, a relevant confounding variable in this setup is the difficulty of isolating CA1 inputs, alongside the involvement of multiple and complex additional signals in the generation of the cognitive map.

In conclusion, this work showed a possible architecture for coupling emergent spatial representations with neuromodulated plasticity to achieve an experience-driven cognitive map. The reliance on some spatial and algorithmic inductive biases, grid cells, and a planning algorithm supports the idea of a labeled graph for goal navigation. Future work can investigate the application to other spatial domains, such as motor control and three-dimensional navigation. In addition, a richer input feature can be added, such as visual information [110], as well as new neuromodulators that encode different sensory dimensions or internally generated signals.

Supporting information

S1 Fig. Place fields obtained from grid-cell activity.

a: grid cell modules with different granularity shown along a continuous trajectory in an open space. For visualization purposes, each module is represented as composed of four sub-modules of 9 grid cells each, whose periodic tuning generates activity that repeats in space. b: place cells whose spatial tuning has been obtained from the concatenation of the grid cells population vector.

https://doi.org/10.1371/journal.pcbi.1013487.s002

(TIF)

S2 Fig. Diagram of the behaviour selection process.

https://doi.org/10.1371/journal.pcbi.1013487.s003

(TIF)

S3 Fig. Distribution of evolved hyper-parameters.

- a-o: results from the last generation, from a run with population size of 96 individuals. Each point represents one individual from the final population, with parameter value on the x-axis and average reward count on the y-axis.

https://doi.org/10.1371/journal.pcbi.1013487.s004

(TIF)

S4 Fig. Sample of generated environments.

https://doi.org/10.1371/journal.pcbi.1013487.s005

(TIF)

S5 Fig. Detour-task comparison between the full and modulation-ablated models.

https://doi.org/10.1371/journal.pcbi.1013487.s006

(TIF)

S1 Table. Summary of principal model and task parameters.

https://doi.org/10.1371/journal.pcbi.1013487.s007

(XLSX)

S2 Table. Comparison of neural network models for spatial navigation and representation.

https://doi.org/10.1371/journal.pcbi.1013487.s008

(XLSX)

S3 Table. Post-wall performance in the detour task for the full and modulation-ablated models.

https://doi.org/10.1371/journal.pcbi.1013487.s009

(XLSX)

S4 Table. Mean reward counts across conditions and environments.

Values shown as mean ± SEM.

https://doi.org/10.1371/journal.pcbi.1013487.s010

(XLSX)

S5 Table. Significant pairwise comparisons (Bonferroni corrected).

https://doi.org/10.1371/journal.pcbi.1013487.s011

(XLSX)

S6 Table. Mean collision counts across conditions and environments.

Values shown as mean ± SEM.

https://doi.org/10.1371/journal.pcbi.1013487.s012

(XLSX)

S7 Table. Significant pairwise comparisons (Bonferroni corrected).

https://doi.org/10.1371/journal.pcbi.1013487.s013

(XLSX)

References

  1. 1. Golledge R, Jacobson RD, Kitchin R, Blades M. Cognitive Maps, Spatial Abilities, and Human Wayfinding. Geogr Rev Jpn, Ser B. 2000;73(2):93–104.
  2. 2. Epstein RA, Vass LK. Neural systems for landmark-based wayfinding in humans. Philos Trans R Soc Lond B Biol Sci. 2013;369(1635):20120533. pmid:24366141
  3. 3. Sargolini F, Fyhn M, Hafting T, McNaughton BL, Witter MP, Moser M-B, et al. Conjunctive Representation of Position, Direction, and Velocity in Entorhinal Cortex. Science. 2006;312(5774):758–62.
  4. 4. Kropff E, Carmichael JE, Moser M-B, Moser EI. Speed cells in the medial entorhinal cortex. Nature. 2015;523(7561):419–24. pmid:26176924
  5. 5. Solstad T, Moser EI, Einevoll GT. From grid cells to place cells: a mathematical model. Hippocampus. 2006;16(12):1026–31. pmid:17094145
  6. 6. Bush D, Barry C, Burgess N. What do grid cells contribute to place cell firing? Trends in Neurosciences. 2014;37(3):136–45.
  7. 7. Neher T, Azizi AH, Cheng S. From grid cells to place cells with realistic field sizes. PLoS One. 2017;12(7):e0181618. pmid:28750005
  8. 8. Li T, Arleo A, Sheynikhovich D. Modeling Place Cells and Grid Cells in Multi-Compartment Environments: Hippocampal-Entorhinal Loop as a Multisensory Integration Circuit. Neural Networks. 2019;121.
  9. 9. Bilash OM, Chavlis S, Johnson CD, Poirazi P, Basu J. Lateral entorhinal cortex inputs modulate hippocampal dendritic excitability by recruiting a local disinhibitory microcircuit. Cell Rep. 2023;42(1):111962. pmid:36640337
  10. 10. Donato F, Xu Schwartzlose A, Viana Mendes RA. How Do You Build a Cognitive Map? The Development of Circuits and Computations for the Representation of Space in the Brain. Annu Rev Neurosci. 2023;46:281–99. pmid:37428607
  11. 11. Poucet B. Spatial cognitive maps in animals: new hypotheses on their structure and neural mechanisms. Psychol Rev. 1993;100(2):163–82. pmid:8483980
  12. 12. Werner S, Krieg-Brückner B, Herrmann T. Modelling Navigational Knowledge by Route Graphs. Freksa C, Habel C, Brauer W, Wender KF. Spatial Cognition II: Integrating Abstract Theories, Empirical Studies, Formal Methods, and Practical Applications. 295–316. Springer, Berlin, Heidelberg, 2000.
  13. 13. Whishaw IQ, Brooks BL. Calibrating space: exploration is important for allothetic and idiothetic navigation. Hippocampus. 1999;9(6):659–67. pmid:10641759
  14. 14. McNaughton BL, Battaglia FP, Jensen O, Moser EI, Moser M-B. Path integration and the neural basis of the “cognitive map”. Nat Rev Neurosci. 2006;7(8):663–78. pmid:16858394
  15. 15. Banino A, Barry C, Uria B, Blundell C, Lillicrap T, Mirowski P, et al. Vector-based navigation using grid-like representations in artificial agents. Nature. 2018;557(7705):429–33. pmid:29743670
  16. 16. Sorscher B, Mel GC, Ocko SA, Giocomo LM, Ganguli S. A unified theory for the computational and mechanistic origins of grid cells. Neuron. 2023;111(1):121-137.e13. pmid:36306779
  17. 17. Christopher J. Cueva and Xue-Xin Wei. Emergence of grid-like representations by training recurrent neural networks to perform spatial localization. 2018.
  18. 18. Stoewer P, Schilling A, Maier A, Krauss P. Neural network based formation of cognitive maps of semantic spaces and the putative emergence of abstract concepts. Sci Rep. 2023;13(1):3644. pmid:36871003
  19. 19. de Cothi W, Nyberg N, Griesbauer E-M, Ghanamé C, Zisch F, Lefort JM, et al. Predictive maps in rats and humans for spatial navigation. Curr Biol. 2022;32(17):3676-3689.e5. pmid:35863351
  20. 20. Whittington JCR, Muller TH, Mark S, Chen G, Barry C, Burgess N, et al. The Tolman-Eichenbaum Machine: Unifying Space and Relational Memory through Generalization in the Hippocampal Formation. Cell. 2020;183(5):1249–63.
  21. 21. Dayan P. Improving Generalization for Temporal Difference Learning: The Successor Representation. Neural Computation. 1993;5(4):613–24.
  22. 22. Blier L, Tallec C, Ollivier Y. Learning Successor States and Goal-Dependent Values: A Mathematical Viewpoint. 2021.
  23. 23. Matthew PHG, Schoenbaum G, Gershman SJ. Rethinking dopamine as generalized prediction error. 2018.
  24. 24. Lee H. Toward the biological model of the hippocampus as the successor representation agent. Biosystems. 2022;213:104612. pmid:35093444
  25. 25. Bono J, Zannone S, Pedrosa V, Clopath C. Learning predictive cognitive maps with spiking neurons during behavior and replays. eLife. 2023;12:e80671.
  26. 26. Fang C, Dmitriy A, Abbott LF, Mackevicius EL. Neural learning rules for generating flexible predictions and computing the successor representation. eLife. 2023;12:e80680.
  27. 27. Nadim F, Bucher D. Neuromodulation of neurons and synapses. Curr Opin Neurobiol. 2014;29:48–56. pmid:24907657
  28. 28. Lisman JE, Grace AA. The hippocampal-VTA loop: controlling the entry of information into long-term memory. Neuron. 2005;46(5):703–13. pmid:15924857
  29. 29. Duszkiewicz AJ, McNamara CG, Takeuchi T, Genzel L. Novelty and Dopaminergic Modulation of Memory Persistence: A Tale of Two Systems. Trends in Neurosciences. 2019;42(2):102–14.
  30. 30. Schultz W, Dayan P, Montague PR. A Neural Substrate of Prediction and Reward. Science. 1997;275(5306):1593–9.
  31. 31. Sheynikhovich D, Otani S, Bai J, Arleo A. Long-term memory, synaptic plasticity and dopamine in rodent medial prefrontal cortex: Role in executive functions. Front Behav Neurosci. 2023;16:1068271. pmid:36710953
  32. 32. Avery MC, Krichmar JL. Neuromodulatory Systems and Their Interactions: A Review of Models, Theories, and Experiments. Front Neural Circuits. 2017;11:108. pmid:29311844
  33. 33. Richard S. Sutton and Andrew G. Barto. The Reinforcement Learning Problem. In Reinforcement Learning: An Introduction. 51–85. MIT Press, 1998.
  34. 34. Fuchsberger T, Paulsen O. Modulation of hippocampal plasticity in learning and memory. Curr Opin Neurobiol. 2022;75:102558. pmid:35660989
  35. 35. Edelmann E, Lessmann V. Dopaminergic innervation and modulation of hippocampal networks. Cell Tissue Res. 2018;373(3):711–27. pmid:29470647
  36. 36. Chowdhury A, Luchetti A, Fernandes G, Filho DA, Kastellakis G, Tzilivaki A, et al. A locus coeruleus-dorsal CA1 dopaminergic circuit modulates memory linking. Neuron. 2022;110(20):3374-3388.e8. pmid:36041433
  37. 37. Takeuchi T, Duszkiewicz AJ, Sonneborn A, Spooner PA, Yamasaki M, Watanabe M, et al. Locus coeruleus and dopaminergic consolidation of everyday memory. Nature. 2016;537(7620):357–62. pmid:27602521
  38. 38. Kempadoo KA, Mosharov EV, Choi SJ, Sulzer D, Kandel ER. Dopamine release from the locus coeruleus to the dorsal hippocampus promotes spatial learning and memory. Proc Natl Acad Sci U S A. 2016;113(51):14835–40. pmid:27930324
  39. 39. Bittner KC, Milstein AD, Grienberger C, Romani S, Magee JC. Behavioral time scale synaptic plasticity underlies CA1 place fields. Science. 2017;357(6355):1033–6. pmid:28883072
  40. 40. Kaufman AM, Geiller T, Losonczy A. A Role for the Locus Coeruleus in Hippocampal CA1 Place Cell Reorganization during Spatial Reward Learning. Neuron. 2020;105(6):1018-1026.e4. pmid:31980319
  41. 41. Retailleau A, Boraud T. The Michelin red guide of the brain: role of dopamine in goal-oriented navigation. Front Syst Neurosci. 2014;8:32. pmid:24672436
  42. 42. Igarashi KM, Ito HT, Moser EI, Moser M-B. Functional diversity along the transverse axis of hippocampal area CA1. FEBS Lett. 2014;588(15):2470–6. pmid:24911200
  43. 43. Ito HT, Schuman EM. Functional division of hippocampal area CA1 via modulatory gating of entorhinal cortical inputs. Hippocampus. 2012;22(2):372–87. pmid:21240920
  44. 44. Brzosko Z, Mierau SB, Paulsen O. Neuromodulation of Spike-Timing-Dependent Plasticity: Past, Present, and Future. Neuron. 2019;103(4):563–81. pmid:31437453
  45. 45. Mei J, Meshkinnejad R, Mohsenzadeh Y. Effects of neuromodulation-inspired mechanisms on the performance of deep neural networks in a spatial learning task. iScience. 2023;26(2):106026. pmid:36818295
  46. 46. Tobler PN, Fiorillo CD, Schultz W. Adaptive coding of reward value by dopamine neurons. Science. 2005;307(5715):1642–5.
  47. 47. Cools R. Chemistry of the Adaptive Mind: Lessons from Dopamine. Neuron. 2019;104(1):113–31. pmid:31600509
  48. 48. Sosa M, Plitt MH, Giocomo LM. Hippocampal sequences span experience relative to rewards.2024;2023.
  49. 49. Milstein AD, Li Y, Bittner KC, Grienberger C, Soltesz I, Magee JC, et al. Bidirectional synaptic plasticity rapidly modifies hippocampal representations. Elife. 2021;10:e73046. pmid:34882093
  50. 50. Fenton AA. Remapping revisited: how the hippocampus represents different spaces. Nat Rev Neurosci. 2024;25(6):428–48. pmid:38714834
  51. 51. Zhou L, Gu Y. Cortical mechanisms of multisensory linear self-motion perception. Neurosci Bull. 2023;39(1):125–37. pmid:35821337
  52. 52. Jerjian SJ, Harsch DR, Fetsch CR. Self-motion perception and sequential decision-making: where are we heading?. Philos Trans R Soc Lond B Biol Sci. 2023;378(1886):20220333. pmid:37545301
  53. 53. Dabaghian Y. Grid cells, border cells, and discrete complex analysis. Front Comput Neurosci. 2023;17:1242300. pmid:37881247
  54. 54. Schøyen VS, Beshkov K, Pettersen MB, Hermansen E, Holzhausen K, Malthe-Sørenssen A, et al. Hexagons all the way down: Grid cells as a conformal isometric map of space. PLOS Computational Biology. 2025;21(2):e1012804.
  55. 55. Montague PR, Dayan P, Sejnowski TJ. A framework for mesencephalic dopamine systems based on predictive Hebbian learning. J Neurosci. 1996;16(5):1936–47. pmid:8774460
  56. 56. Krishnan S, Heer C, Cherian C, Sheffield MEJ. Reward expectation extinction restructures and degrades CA1 spatial maps through loss of a dopaminergic reward proximity signal. Nat Commun. 2022;13(1):6662. pmid:36333323
  57. 57. Tsetsenis T, Broussard JI, Dani JA. Dopaminergic regulation of hippocampal plasticity, learning, and memory. Front Behav Neurosci. 2023;16:1092420. pmid:36778837
  58. 58. Sosa M, Plitt MH, Giocomo LM. A flexible hippocampal population code for experience relative to reward. Nat Neurosci. 2025;28(7):1497–509. pmid:40500314
  59. 59. Lubica Benuskova and Wickliffe C. Abraham. STDP rule endowed with the BCM sliding threshold accounts for hippocampal heterosynaptic plasticity. Journal of Computational Neuroscience, 22(2):129–33, April 2007.
  60. 60. McNamara CG, Tejero-Cantero Á, Trouche S, Campo-Urriza N, Dupret D. Dopaminergic neurons promote hippocampal reactivation and spatial memory persistence. Nat Neurosci. 2014;17(12):1658–60. pmid:25326690
  61. 61. Michon F, Krul E, Sun J-J, Kloosterman F. Single-trial dynamics of hippocampal spatial representations are modulated by reward value. Curr Biol. 2021;31(20):4423-4435.e5. pmid:34416178
  62. 62. Philip Shamash and Tiago Branco. Mice identify subgoal locations through an action-driven mapping process. 2021.
  63. 63. Meilinger T, Strickrodt M, Bülthoff HH. Qualitative differences in memory for vista and environmental spaces are caused by opaque borders, not movement or successive presentation. Cognition. 2016;155:77–95. pmid:27367592
  64. 64. Igel C, Hansen N, Roth S. Covariance matrix adaptation for multi-objective optimization. Evol Comput. 2007;15(1):1–28. pmid:17388777
  65. 65. Retailleau A, Morris G. Spatial Rule Learning and Corresponding CA1 Place Cell Reorientation Depend on Local Dopamine Release. Curr Biol. 2018;28(6):836-846.e4. pmid:29502949
  66. 66. Schultz W. Dopamine reward prediction error coding. Dialogues Clin Neurosci. 2016;18(1):23–32. pmid:27069377
  67. 67. Ormond J, Serka SA, Johansen JP. Enhanced Reactivation of Remapping Place Cells during Aversive Learning. J Neurosci. 2023;43(12):2153–67. pmid:36596695
  68. 68. Blair GJ, Guo C, Wang S, Fanselow MS, Golshani P, Aharoni D, et al. Hippocampal place cell remapping occurs with memory storage of aversive experiences. eLife. 2023;12:e80661.
  69. 69. Rooke S, Wang Z, Di Tullio RW, Balasubramanian V. Trading Place for Space: Increasing Location Resolution Reduces Contextual Capacity in Hippocampal Codes. 2025;2024.10.29.620785.
  70. 70. Ziv Y, Burns LD, Cocker ED, Hamel EO, Ghosh KK, Kitch LJ, et al. Long-term dynamics of CA1 hippocampal place codes. Nat Neurosci. 2013;16(3):264–6. pmid:23396101
  71. 71. Priestley JB, Bowler JC, Rolotti SV, Fusi S, Losonczy A. Signatures of rapid plasticity in hippocampal CA1 representations during novel experiences. Neuron. 2022;110(12):1978-1992.e6. pmid:35447088
  72. 72. Vaidya SP, Li G, Chitwood RA, Li Y, Magee JC. Formation of an expanding memory representation in the hippocampus. Nature Neuroscience. 2025;28(7):1510–8.
  73. 73. Hok V, Lenck-Santini P-P, Roux S, Save E, Muller RU, Poucet B. Goal-related activity in hippocampal place cells. J Neurosci. 2007;27(3):472–82. pmid:17234580
  74. 74. Kubie JL, Fox SE. Do the spatial frequencies of grid cells mold the firing fields of place cells?. Proc Natl Acad Sci U S A. 2015;112(13):3860–1. pmid:25829538
  75. 75. Gallistel CR, Cramer AE. Computations on metric maps in mammals: getting oriented and choosing a multi-destination route. J Exp Biol. 1996;199(Pt 1):211–7. pmid:8576692
  76. 76. Sabine Gillner and Hanspeter A. Mallot. Navigation and Acquisition of Spatial Knowledge in a Virtual Maze. Journal of Cognitive Neuroscience, 10(4):445–63, July 1998.
  77. 77. Tran AH, Uwano T, Kimura T, Hori E, Katsuki M, Nishijo H, et al. Dopamine D1 receptor modulates hippocampal representation plasticity to spatial novelty. J Neurosci. 2008;28(50):13390–400. pmid:19074012
  78. 78. Ishikawa T, Montello DR. Spatial knowledge acquisition from direct experience in the environment: individual differences in the development of metric knowledge and the integration of separately learned places. Cogn Psychol. 2006;52(2):93–129. pmid:16375882
  79. 79. Warren WH. Non-Euclidean navigation. J Exp Biol. 2019;222(Pt Suppl 1):jeb187971. pmid:30728233
  80. 80. Meilinger T. The Network of Reference Frames Theory: A Synthesis of Graphs and Cognitive Maps. In Christian Freksa, Nora S. Newcombe, Peter Gärdenfors, and Stefan Wölfl, editors, Spatial Cognition VI. Learning, Reasoning, and Talking about Space. 344–60, Berlin, Heidelberg, 2008. Springer.
  81. 81. Wang JX, Kurth-Nelson Z, Tirumala D, Soyer H, Leibo JZ, Munos R, et al. Learning to reinforcement learn. 2017.
  82. 82. Schinazi VR, Nardi D, Newcombe NS, Shipley TF, Epstein RA. Hippocampal size predicts rapid learning of a cognitive map in humans. Hippocampus. 2013;23(6):515–28. pmid:23505031
  83. 83. Matthew R. Spatial localization of hippocampal replay requires dopamine signaling. eLife. 2025.
  84. 84. Javaid A. Understanding Dijkstra’s Algorithm. 2013.
  85. 85. de Cothi W, Barry C. Neurobiological successor features for spatial navigation. Hippocampus. 2020;30(12):1347–55. pmid:32584491
  86. 86. Brzosko Z, Zannone S, Schultz W, Clopath C, Paulsen O. Sequential neuromodulation of Hebbian plasticity offers mechanism for effective reward-based navigation. Elife. 2017;6:e27756. pmid:28691903
  87. 87. Aitken F, Kok P. Hippocampal representations switch from errors to predictions during acquisition of predictive associations. Nat Commun. 2022;13(1):3294. pmid:35676285
  88. 88. Halvagal MS, Zenke F. The combination of Hebbian and predictive plasticity learns invariant object representations in deep sensory networks. Nat Neurosci. 2023;26(11):1906–15. pmid:37828226
  89. 89. Ororbia A. Spiking neural predictive coding for continually learning from data streams. Neurocomputing. 2023;544:126292.
  90. 90. Stachenfeld KL, Botvinick MM, Gershman SJ. The hippocampus as a predictive map. Nature Neuroscience. 2017;20(11):1643–53.
  91. 91. Fiorillo CD, Tobler PN, Schultz W. Discrete coding of reward probability and uncertainty by dopamine neurons. Science. 2003;299(5614):1898–902. pmid:12649484
  92. 92. Sosa M, Giocomo LM. Navigating for reward. Nat Rev Neurosci. 2021;22(8):472–87. pmid:34230644
  93. 93. Ali A, Ahmad N, de Groot E, Marcel AJ, Kietzmann TC. Predictive coding is a consequence of energy efficiency in recurrent neural networks. 2021.
  94. 94. Nair IR, Bhasin G, Roy D. Hippocampus maintains a coherent map under reward feature-landmark cue conflict. Front Neural Circuits. 2022;16:878046. pmid:35558552
  95. 95. Grienberger C, Magee JC. Entorhinal cortex directs learning-related changes in CA1 representations. Nature. 2022;611(7936):554–62. pmid:36323779
  96. 96. Tryon VL, Penner MR, Heide SW, King HO, Larkin J, Mizumori SJY. Hippocampal neural activity reflects the economy of choices during goal-directed navigation. Hippocampus. 2017;27(7):743–58. pmid:28241404
  97. 97. Hannah S Wirtshafter and Matthew A Wilson. Differences in reward biased spatial representations in the lateral septum and hippocampus. eLife. 2020;9:e55252.
  98. 98. Anderson MI, Jeffery KJ. Heterogeneous modulation of place cell firing by changes in context. J Neurosci. 2003;23(26):8827–35. pmid:14523083
  99. 99. Lee I, Griffin AL, Zilli EA, Eichenbaum H, Hasselmo ME. Gradual Translocation of Spatial Correlates of Neuronal Firing in the Hippocampus toward Prospective Reward Locations. Neuron. 2006;51(5):639–50.
  100. 100. Schoenenberger P, O’Neill J, Csicsvari J. Activity-dependent plasticity of hippocampal place maps. Nat Commun. 2016;7:11824. pmid:27282121
  101. 101. Wang CH, Monaco JD, Knierim JJ. Hippocampal Place Cells Encode Local Surface-Texture Boundaries. Current biology: CB. 2020;30(8):1397–409.
  102. 102. Yang X, Cacucci F, Burgess N, Wills TJ, Chen G. Visual boundary cues suffice to anchor place and grid cells in virtual reality. Curr Biol. 2024;34(10):2256-2264.e3. pmid:38701787
  103. 103. Wiener SI, Paul CA, Eichenbaum H. Spatial and behavioral correlates of hippocampal neuronal activity. J Neurosci. 1989;9(8):2737–63. pmid:2769364
  104. 104. Fyhn M, Molden S, Hollup S, Moser M-B, Moser EI. Hippocampal Neurons Responding to First-Time Dislocation of a Target Object. Neuron. 2002;35(3):555–66.
  105. 105. Lenck-Santini P-P, Rivard B, Muller RU, Poucet B. Study of CA1 place cell activity and exploratory behavior following spatial and nonspatial changes in the environment. Hippocampus. 2005;15(3):356–69. pmid:15602750
  106. 106. Dupret D, O’Neill J, Pleydell-Bouverie B, Csicsvari J. The reorganization and reactivation of hippocampal maps predict spatial memory performance. Nat Neurosci. 2010;13(8):995–1002. pmid:20639874
  107. 107. Burke SN, Maurer AP, Nematollahi S, Uprety AR, Wallace JL, Barnes CA. The Influence of Objects on Place Field Expression and Size in Distal Hippocampal CA1. Hippocampus. 2011;21(7):783–801.
  108. 108. Tanni S, de Cothi W, Barry C. State transitions in the statistically stable place cell population correspond to rate of perceptual change. Curr Biol. 2022;32(16):3505-3514.e7. pmid:35835121
  109. 109. Scleidorovich P, Fellous J-M, Weitzenfeld A. Adapting hippocampus multi-scale place field distributions in cluttered environments optimizes spatial navigation and learning. Front Comput Neurosci. 2022;16:1039822. pmid:36578316
  110. 110. Wen JH, Sorscher B, Aery Jones EA, Ganguli S, Giocomo LM. One-shot entorhinal maps enable flexible navigation in novel environments. Nature. 2024;635(8040):943–50. pmid:39385034