Figures
Abstract
Vehicle trajectory prediction (VTP) is the primary part of the perception-planning-control pipeline in autonomous driving. The inter-vehicle interactions strongly shape the trajectory of vehicles in complex traffic scenarios. Neural networks that use recurrent or convolutional architectures often exhibit performance degradation in longer prediction horizons. Many existing approaches confine interaction modeling to a predefined spatial neighborhood and a narrow temporal window, while implicitly assuming that all neighboring vehicles have equal influence on the future trajectory of target vehicle. This results in error accumulation and the inability to capture long-range spatiotemporal dependencies. To overcome these limitations, a non-local multi-head spatiotemporal attention based long short-term memory model (NL-MHA-LSTM) is introduced which employs an attention mechanism to assign context weights to relevant neighbor vehicles. It extends beyond pairwise effects to model long-range dependencies. The model emphasizes the most influential vehicles and formulate position-aware interaction representations. A comprehensive set of experiments is conducted on the pubicly available HighD dataset. The results demonstrate that the proposed model outperforms all state-of-the-art methods, achieving a 53.4% reduction in RMSE at the 5s prediction horizon relative to the next-best model. In addition, a detailed ablation study is conducted to systematically evaluate the impact of different attention mechanisms and varying numbers of attention heads on prediction accuracy across multiple horizons.
Citation: Rashid S, Khan MA, Akram U, Mumtaz R, Lee IE, Syed TA (2026) A non-local multi-head spatiotemporal attention LSTM for vehicle trajectory prediction. PLoS One 21(9): e0357735. https://doi.org/10.1371/journal.pone.0357735
Editor: Yongxiang Zhang, Southwest Jiaotong University, CHINA
Received: February 24, 2026; Accepted: August 20, 2026; Published: September 21, 2026
Copyright: © 2026 Rashid et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The HighD Dataset used in this study is available at https://levelxdata.com/highd-dataset/.
Funding: This research work is supported by the Ministry of Higher Education (MOHE) under the 2023 Translational Research Program for the Energy Sustainability Focus Area (Project ID: MMUE/240001), the 2024 ASEAN IVO (Project ID: 2024-02), Multimedia University and Islamic University of Madinah. The funders contributed to the study design, experimentation, and the Article Processing Charge (APC) for this manuscript.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Accurate modeling of the surrounding vehicle ensures safe navigation and collision avoidance in autonomous driving. In real-world scenarios, when vehicle-to-vehicle (V2V) communication is not available, autonomous systems infer the intentions of surrounding vehicles using their on-board sensory data. A detailed analysis of vehicle interactions and historical trajectories is required to improve situational awareness [1]. Intelligent Transportation Systems (ITS) integrates data analytics, communication, and control in vehicular networks. Using real-time data from sensors and connected platforms, ITS improves mobility while addressing key challenges such as traffic congestion and accidents [2]. Autonomous vehicles follow a tightly integrated pipeline; perception, prediction, planning, and control [3]. Multi-modal sensors provide information about the environment, which helps to detect lanes, vehicles, and other road agents in real time. The limitations of individual sensors are overcome by a combination of sensor modalities for low-light and adverse weather conditions [4,5].
Traditional trajectory prediction methods are lightweight, but often struggle to model multi-agent interactions of dynamic traffic scenarios [6]. The shift toward data-driven methods has led to the adoption of deep neural networks, which often consider trajectory prediction as a sequence prediction task. In this framework, convolutional neural networks (CNN) and recurrent neural networks (RNNs) [7] have been extensively applied either individually or in combination with their advanced variants, namely gated recurrent units (GRUs) [8] and LSTM [9] networks. These architectures do not use any rules to capture inter-vehicle interactions and lane-changing behaviors.
Classical pooling schemes are computationally efficient, but often oversimplify the complex interactions, which decrease the overall prediction accuracy. Whereas, attention-based methods concentrate on the important dependencies and events in the sequence [10]. The recent literature has shown the high potentials of attention mechanisms for trajectory prediction [11].
In this paper, we propose a deep learning architecture based on LSTM with a non-local multi-head spatiotemporal attention mechanism (NL-MHA-LSTM) for vehicle trajectory prediction. Our model selectively aggregates the states of the most relevant surrounding vehicles regardless of their distance from the target vehicle. It also captures long-range critical dependencies in complex traffic scenarios. The model avoids any limitation of fixed spatial neighborhood while maintaining computational efficiency. A series of extensive experiments validate that proposed model works better than state-of-the-art methods on long prediction horizons in complex driving scenarios.
Vehicle trajectory prediction has become very popular in the last decade due to its significance in autonomous driving. A number of research works addressed this challenge, coming up with different modeling approaches and architecture designs. In this section, we narrow our research focus that specifically explores attention-enhanced LSTM architectures and their variants, highlighting how these approaches advance the state of the art in capturing complex spatiotemporal dependencies.
1. Related research
Social interactions have been considered for trajectory prediction in recent studies. Recurrent architectures, particularly LSTM networks, remain a cornerstone in capturing temporal dependencies [12]. LSTM-based approaches make use of the historical trajectory patterns to predict trajectories in extended horizons with a lower computational cost [13]. LSTM frameworks often integrate lane orientation and driving intentions to improve predictive performance [14]. Other studies improved LSTMs through multi-layer or hierarchical designs.
The ST-LSTM is proposed by Dai et al. [15]. It is a multi-layer model with two layers; the first layer learns the trajectory of an individual vehicle while the second one learns the interaction among the agents. Yang et al. [16] used similarity metrics on historical trajectories to improve prediction accuracy. The real time prediction of vehicle trajectory is also addressed using efficient LSTM-based architectures in dynamic traffic scenarios [17].
More advanced frameworks include the use of dedicated components or the combination of LSTMs with other models. Karimzadeh et al. [18] combined transfer and reinforcement learning to model urban traffic dynamics. Tang et al. [19] introduced stacked spatio-temporal layers along with an LSTM trajectory generator for modeling the motion of multiple agents. Chen et al. [20] designed a dynamic feature extraction component for capturing detailed interactions between vehicles. Min [21] proposed an architecture with hierarchical LSTMs to model driving intents and lane changes together with future trajectories.
Attention mechanisms address this challenge by adaptively focusing on the most informative components of the input, thereby enhancing the model’s capacity to identify critical dependencies and positional patterns within sequential data [10]. Their effectiveness in VTP has been demonstrated in recent works [11,22]. Attention mechanisms have been widely used in combination with LSTM and its variants. Messaoud et al. [23] designed an architecture based on LSTM which measures the effect of surrounding vehicles through an attention mechanism. Zhiming et al. [24] captured long-range dependencies to address the bottleneck of fixed-length context vectors in LSTM models. Jin et al. [25] designed a temporal attention mechanism within the LSTM decoder to identify events from historical trajectories. Building on this, Wang et al. [26] applied multi-head temporal attention to model vehicle behavior across different time scales, extracting rich interaction features through social tensors. Gao et al. [27] employed a similar approach of self-attention to capture global dependencies through scaled dot-product attention, generating context-aware trajectory representations. Qiao et al. [28] also applied self-attention to hidden LSTM states deriving attention-based adjacency matrices to reveal strong inter-vehicle dependencies.
In addition to LSTM-based models, attention mechanisms have been widely used with other deep learning architectures for trajectory prediction and related tasks. Cai et al. [29] utilized the graph attention network with convolutional social pooling in order to enhance trajectory prediction using contextual features. Similarly, Yuhuan et al. [30] introduced a temporal attention mechanism that adaptively weights scene images captured at different time intervals, inspired by channel attention in convolutional networks. This design computes pairwise attention scores between frames, allowing the model to focus on the most informative time steps. Tian et al. [9] incorporated attention with an encoder-decoder architecture based on interaction-based modeling for better trajectory predictions. Several other studies have incorporated transformer-based layers to encode agent-to-agent interactions in dynamic traffic scenarios [31–36]. Mo et al. [37] introduced edge-enhanced attention to predict trajectories for heterogeneous agents in complex traffic scenarios. The local graph and the temporal attention are combined to identify road structures [38]. An extension to this work introduces multi-head social attention layer. This layer enables models to selectively attend to spatially relevant neighbors while preserving long-range interaction modeling [39]. Spatiotemporal dual attention is employed to refine multimodal predictions by capturing dependencies in both historical and predicted future trajectories [40]. Cheng et al. [41] combined temporal attention for target vehicle history with spatial attention for surrounding agents. A hybrid framework is proposed that uses a graph attention network encoder for spatial interaction modeling and a transformer encoder for temporal dependencies [42]. Luo et al. [43] enhanced spatial attention by assigning adaptive weights to neighbors within a graph attention network (GAT). The diverse interaction patterns are captured using multi-head aggregation. Li et al. [44] integrated attention directly into the adjacency matrix allowing time-varying weights for connected nodes. A spatiotemporal attention model is developed that reflects human driving behavior, adaptively weighting both historical trajectories and influential spatial neighbors [45]. Similarly, Wang et al. [35] proposed an attention-based fusion framework that integrates multi-head self-attention to capture long-range temporal dependencies along with cross-attention mechanisms that capture interactions between traffic agents and the surrounding road infrastructure. A novel Lane-Specific Spatial-Temporal Attention Network (LSTAN) is introduced to handle the lane-level traffic data in predicting trajectories of vehicles [46]. More specifically, an encoder based on the long short-term memory networks is employed for learning the temporal information from both the target vehicles and the nearby vehicles. Lane attention module (LAM) and temporal self-attention module (TSAM) are designed for collecting the spatial and temporal information respectively. A trajectory prediction approach is proposed for highway weaving areas that considers the transfer properties of vehicle interactions as an integral part of the modeling process [47]. In particular, a hierarchical dynamic graph is used to consider transfer properties in two degrees of coupling; feature level and decision level. Thus, the problem of vehicle interactions modeling in mixed traffic is solved. Baicang et al. [48] proposed a novel multi-modal, end-to-end model for autonomous driving to predict vehicle speed and steering angle with high precision. The architecture utilizes spatial encoding, temporal dependencies and attention mechanism to identify critical features. Jin et al. [25] introduced STA-LSTM, a multi-modal trajectory prediction model to handle complex vehicle interactions in connected vehicle environments. The framework utilizes spatial grid occupancy to model interactions and employs convolutional pooling to extract dynamic interactive features. It effectively recognizes driving intentions and provides more reliable long-term trajectory predictions. Fig 1 provides insight into emerging research areas in vehicle trajectory prediction based on our literature survey.
Despite the fact that attention modules have proved to be effective at capturing the long-range temporal dependencies between relevant neighbors, lack of adequate analysis with respect to positional relevance and interactions dynamics may result in insufficient analysis of the traffic scene. Inaccurate choice of the attention heads may make the model miss important interactions, leading to low prediction accuracy. Our latest research presents an optimized LSTM architecture that utilizes spatio-temporal interactions to capture long-range dependencies without limiting interactions to any predefined spatial or temporal neighborhood [49].
In extension to our previous work, a non-Local multi-head spatiotemporal attention LSTM (NL-MHA-LSTM) is proposed for vehicle trajectory prediction. The proposed model captures a comprehensive range of interactions by selectively attending to the most relevant neighboring vehicles, without being constrained by all neighboring agents or predefined distance thresholds. Long-range dependencies can be efficiently modeled by it without adding any more trainable parameters into it.
This paper offers several distinctive contributions, including the following:
- This work is an extension of our previous work [49]. A Non-Local Multi-Head Spatiotemporal Attention LSTM (NL-MHA-LSTM) is proposed for vehicle trajectory prediction. A multi-layer LSTM network is used for encoding trajectory history and relations among vehicles. The relevant neighbors are assigned attention weights using an attention mechanism. The LSTM decoder predicts trajectories by capturing long range dependencies.
- An ablation study is performed to compare the performance of different attention mechanisms, attention heads and efficiency metrics.
- The NL-MHA-LSTM model is evaluated on widely used HighD dataset, a well-established benchmark for highway vehicle trajectory prediction. Experimental results demonstrate that the proposed model achieves competitive performance relative to existing state-of-the-art approaches across long-range prediction horizons.
The structure of the paper is organized as follows. Section I provides an outline of the challenges and importance of vehicle trajectory prediction. Section II presents a literature review. Section III describes the architectural design of the NL-MHA-LSTM model. Section IV provides the experimental set up which includes implementation details, performance evaluation, models used for comparison and ablation study. Section V discusses experimental results, performance analysis, and comparative analysis of the proposed model. Finally, section VI concludes the paper and suggests directions for further research.
2. Overview of the proposed NL-MHA-LSTM model
This section presents the architecture of the proposed NL-MHA-LSTM model. Fig 2 presents the proposed NL-MHA-LSTM. First, the global scene is decomposed into multiple time steps. These concatenated vectors are fed to the LSTM Encoder to learn temporal dynamics and interactions. A multihead attention block computes head-specific projections with similarity scores and a softmax over neighbors produces interaction weights. The contexts from all heads are then concatenated and passed through a linear layer to form a social embedding. The decoder then generates future positions for each time step. The design prioritizes behaviorally relevant neighbors without fixed distance thresholds and captures long-range dependencies efficiently.
The aim is to use the history of target vehicle and its interaction with neighbors to predict the future trajectory. Let represent the collection of all states of vehicles, where
is target vehicle and
is the neighbor vehicle for
. In highway scenes modeled after HighD geometry (six lanes, three per direction), we adopt
as a practical upper bound corresponding to the usual interaction set: preceding, following, left-alongside, right-alongside, left-preceding, left-following, right-preceding, and right-following. This cap can be relaxed via adaptive neighbor selection that learns relevance scores per agent, or replace it with a fixed encoder for a graph-based architecture to scale to variable agents. The state vector
could contain the planar position (x, y) along with many other features like velocity and acceleration. The superscript refers to vehicle id and the subscript shows the time step. The future positions of the target vehicle from T + 1 to T + P can be expressed as Eq 1.
Throughout this work, P refers to the prediction horizon, which spans from 1 to 5 seconds. All coordinates are in target-centric frame with origin at the target’s position at time T.
2.1 Non-local spatio-temporal feature extraction
The model input is defined as for all K tracks, where
reflecting the 60 tracks in the dataset. For this input, features of target vehicle and its neighbors are combined into a structured feature matrix. Each track has a different number of unique vehicles, we process every track and build a 60-dimensional feature vector for each vehicle at all timestamps during which it appears as a neighbor to another vehicle.
The term “non-local” means that all surrounding vehicles states are incorporated into the feature matrix globally, without imposing any spatial distance threshold or temporal window restriction throughout the observation window. Unlike local methods that restrict neighborhood to a fixed spatial boundary (10-15m) or temporal window (1-3s), our non-local formulation incorporates the states of all surrounding vehicles across the entire observation window without any proximity constraint. This is significant in lane-changing scenarios, where a distant vehicle in an adjacent lane may have a greater influence on the target vehicle’s future trajectory than a spatially closer one. This design guarantees that attention mechanism can access distant spatiotemporal interactions, which captures complete context to assign influence weights to surrounding vehicles, a richer foundation as compared to existing approaches. For the target vehicle, twelve features are extracted as detailed in Table 1, while six features are extracted for neighbor vehicle covering position (x, y), velocity , and acceleration
. These are brought together into a unified non-local spatiotemporal feature matrix that serves as the input to the attention module. An insightful description of this matrix is provided in our previous work [49].
2.2 LSTM architecture
The proposed NL-MHA-LSTM adopts a two-layer LSTM encoder-decoder architecture augmented with a multi-head attention mechanism that quantifies the influence of surrounding vehicles on the target vehicle’s future trajectory. The architecture consists of three core components: an encoder that processes the observed trajectory history, an attention module that learns to selectively weight inter-vehicle interactions, and a decoder that generates the predicted future trajectory based on the encoded context and attention-weighted representations. The mathematical equations that describe their functioning are given below.
2.2.1 Encoder.
The encoder takes input sequences , where B is the batch size, T is the sequence length, and dinput represents feature dimension. It iteratively updates its states as Eq (2) and outputs the terminal hidden and cell states
that summarize the entire input sequence as Eq (3).
2.2.2 Multi-head attention module.
After the LSTM encodes every vehicle’s motion history into a single vector, we feed the hidden terminal states into the attention module. In self-attention mechanism, each token or timestep forms a context-aware representation by attending to all others including itself. It considers a single attention head. This operation captures long-range dependencies in a translation invariant way and has become a standard mechanism for sequence modeling. To increase expressiveness and stabilize training, multi-head attention runs H independent attention heads in parallel, each with its own projections. The vehicle interactions are computed as follows:
are the learnable projection matrices in each attention head l. Each head thus captures a dynamics-rich representation of how every neighbour vehicle has been moving relative to the target vehicle. The l-th head for the target i aggregates neighbors via a softmax-weighted sum as Eq (7). Here,
,
encode the influence of neighbor j on target i based on their relative dynamics, and
is the value vector.
This yields a permutation-invariant social context that emphasizes the most influential neighbors. We explore three formulations for computing the attention weights . In the standard neighborhood-based softmax formulation, the interaction weights are defined as Eq (8), where
denotes the similarity score between the target i and the neighbor j. Hence,
measures how much the neighbor j influences the target of the attention head l. The similarity score
is then computed using one of several learnable functions (e.g., dot-product, additive, or distance-based attention) as expressed in Eq (9), where
denotes vector concatenation and
is the dimension of each key/query vector). The attention weights are computed exclusively from the encoding position and related dynamics, allowing the model to identify behaviorally relevant neighbors without revisiting the raw coordinates or relying on hand-crafted features.
Moreover, while a single attention map can re-weight neighbors, a multi-head design probes complementary subspaces of trajectory features, e.g., one head emphasizing longitudinal gaps, another lateral incursions, and a third distant high-speed agents thereby enriching the interaction representation. Finally, the M head contexts are concatenated and linearly projected to form a single social context vector as expressed in Eq (10). Different heads discover complementary relations, e.g., short- vs. long-range motion features, lane-level vs. interaction-level context, which is crucial in trajectory prediction.
2.2.3 Decoder.
The decoder produces the future trajectory conditioned on the encoder context
. At t = 0, the decoder is initialized with the encoder’s terminal states and the last encoder input:
For , the LSTM updates are expressed in Eq (11) and Eq (12). The output head maps the hidden state to a 2D position.
In an autoregressive setting, (or a teacher-forced target during training) is fed back as the next input
, yielding a sequential rollout over the horizon.
3. Experimental evaluation
The proposed approach is evaluated using the publicly available highD dataset [50]. The details about the data set, data processing pipeline, training process, and the experimental setup can be found in our previous paper [49]. The data set is split such that 75% are used for training and 25% for testing. Each trajectory is segmented into an 8-second window, with the first 3 seconds serving as the observed input sequence and the remaining 5 seconds treated as the prediction horizon. Both the sequences are downsampled to 5 Hz. Two layers of LSTM are used during training with batch size of 32 and a learning rate of 10−4. The hidden state of LSTM encoder and decoder has a dimension of 128, and an output dimension of 2. RMSprop optimizer is used in combination with PyTorch framework and training is performed on an NVIDIA GeForce GTX 1080 GPU.
3.1 Evaluation metrics
The proposed model is evaluated using following metrics.
3.1.1 Root Mean Square Error (RMSE).
RMSE measures the error between ground truth and predicted trajectories. In equation 13, and
denote the predicted x,y coordinates of the ith vehicle at time t, while
and
represent the corresponding actual coordinates.
denotes the RMSE at time t. Furthermore,
represents RMSE at time t.
3.1.2 Longitudinal error (
).
Longitudinal error captures how far the predicted position is from the actual position in the direction of the vehicle motion. Given a trajectory along the x-axis, the longitudinal error is determined by the difference in predicted and actual x-positions.
3.1.3 Lateral error (
).
Lateral error captures how far the predicted position is from the actual position in the direction perpendicular to the vehicle’s path. In Eq (15), and
the predicted and actual y-coordinates, respectively.
3.2 Models compared
The performance of proposed model is compared with following state-of-the-art models.
- AS-LSTM [28]: This model encodes vehicle interactions using a social-pooling layer.It also utilizes the self-attention mechanism to improve the weight distribution between inputs and outputs, which captures better positional effects.
- Planning-informed prediction LSTM (PiP-LSTM) [51]: It applies planning information as an informed condition to produce multimodal trajectories.
- NLS-LSTM [52]: A multi-modal prediction system based on an encoder-decoder architecture is utilized to provide a better insight into the variations in trajectories.
- Multi-Head Attention Social Pooling (MHA-LSTM) [23]: The interaction modeling relies on a multi-head attention mechanism with four heads, though it works purely from vehicle position data, leaving other kinematic features.
- Multi-Head Attention Social Pooling (MHA-LSTM(+f)) [23]: This model is an extension of the MHA-LSTM, where additional inputs like velocity, acceleration, and classes have been added along with three attention heads.
- Sparse Graph Convolutional Network (SGCN) [53]: It models interactions by building a sparse directed spatial graph and learning information flow along its edges.
- Scene LSTM Encoder-Decoder (S-LSTM) [54]: model is based on the social encoder-decoder model, which includes full connection pooling.
- CS-LSTM [55]: It uses the temporal backbone of the target vehicle to predict its future trajectory by utilizing an LSTM. It ignores the effect of other vehicles around it in predicting trajectory.
- LS-LSTM [56]: This model integrates an attention system within the traditional LSTM framework. It encodes both the target vehicle’s kinematic state and the behavioral patterns of surrounding vehicles across lanes, while dynamically weighing the relevance of each feature at every time step.
- Dual Transformer [27]: This approach employs a dual-transformer architecture to simultaneously infer lane-change intent and predict vehicle trajectories. A multihead attention mechanism is utilized to characterize the interactions among target vehicle and surrounding vehicles.
- Hybrid-LSTM [45]: This hybrid framework models how surrounding vehicles react to the intended trajectory of the ego vehicle using spatiotemporal attention. The attention mechanism adaptively weights past trajectories and inter-vehicle interactions, enabling the model to understand driving intent and predict motion more accurately.
- Graph Recurrent Attentive Neural Process (GRANP) model [43]: This model leverages neural processes to directly quantify and visualize the uncertainty of the prediction. Its backbone couples a GAT with an LSTM and three stacked 1-D convolutional layers to model spatio-temporal dependencies efficiently.
- GAT-TR-LSTM-LSTM [42]: This hybrid model fuses a GAT to model interactions on the road-agent graph and a transformer encoder to extract global temporal dependencies. An LSTM encoder captures local temporal patterns and the LSTM decoder generates future trajectories.
3.3 Ablation study
We performed an ablation study to identify key factors driving attention performance, examining the module along three complementary problems. The first problem is based on the number of attention heads. For this, models were trained with two to eight (2–8) attention heads to assess how multi-head diversity influences prediction accuracy. The second study is based on the type of attention method. For this, three attention methods are compared including attention, scaled dot product, and concatenation-based attention. The third ablation study reports the performance parameters including inference time, number of parameters, number of FLOPs for different attention methods, and number of attention heads.
The findings of the first study are reported in next section. Using insufficient attention heads can not capture long range spatial and temporal dependencies appropriately. Prediction accuracy and the choice of attention head count must also be balanced against computational cost as more heads increase execution time and resource usage.
4. Results and discussion
For the first ablation study of the number of attention heads, we performed a series of experiments. For this purpose, our model is evaluated for prediction horizons 1-5s and number of attention heads varying from 2 to 8. We have provided RMSE, longitudinal and lateral errors in Table 2. It shows that increasing the number of attention heads from (h = 2) to (h = 3) yields consistent gains at all horizons with RMSE at 5s dropping from 0.87 to 0.68, corresponding reductions in longitudinal 1.07 to 0.82 and lateral errors 0.34 to 0.29. Using h = 4 is broadly comparable to h = 3, but performance degrades for at 5s. RMSE increases to 0.67 with h = 6 and 0.69 with h = 8, while the longitudinal error increases to 0.81 and 0.85, respectively. These trends indicate a capacity fragmentation trade-off with fixed model width, larger ’h’ shrinks the per-head dimensionality
, diffuses attention and weakens long-horizon temporal coherence, while h = 5 provides the best balance of interaction modeling and stability over 1-5s. Considering the results, h = 5 is chosen for all subsequent experiments, as it offers the optimal trade-off in terms of interaction modeling and predictive robustness in both short- and long-term prediction perspectives.
The second ablation study is carried out on the type of attention method used. For this purpose, our model is evaluated for -attention, scaled dot-product, and concatenation attention over the 1-5s prediction horizon. Table 3 shows that the scaled dot product achieves the lowest RMSE (0.09 to 0.29) for horizons 1-3s. However, at the 4s and 5s horizon, concatenation yields the lowest RMSEs of 0.39 and 0.35, respectively. alpha attention remains at highest RMSE at all horizons except 1s. The reason is that it does not consider encoding vectors of the target and relies on encoding vectors of the surrounding vehicles only. These results indicate that the concatenation yields better logit calibration and more stable long-horizon behavior, and the scaled-dot product performs well in short horizons.
For qualitative analysis of attention heads, Fig 3 shows four heads attention maps for lane keeping/accelerate, de-accelerate, right and left lane-change driving scenarios. The four attention heads are represented as h0....h4. The target vehicle is represented with green, and its neighbors are shown in blue color. The color of neighbors indicates their attention weight; it is darker when they have more weight. We can remark that the vehicles in front of target vehicle are assigned a higher attention weight compared to the vehicles following it. In each scenario, there is one head that assigns equal weight to all surrounding vehicles. In lane keeping (h3), de-accelerate (h1), right lane change (h2) and left lane change (h3). Similarly, there is one attention head that assigns the most distinctive weights to all neighbors in each scenario. In lane keeping (h2), de-accelerate (h2), right lane change (h1) and left lane change (h2). We can also notice that in the left lane change scenario (h2), the vehicle at the longest distance is assigned the most weight and the closest neighbors are assigned less weight. The closest neighbors do not have the greatest influence in predicting the trajectory of the vehicle.
The third ablation study reports the inference time, the number of parameters, the number of FLOPs for different attention methods, and the number of attention heads. Fig 4 shows the performance of our model with different attention methods. The bubble sizes represent the parameters of the proposed model. The scaled dot attention (green bubble) exhibits lower FLOPs and latency because its scores include batched dot products and a softmax. Alpha attention (blue bubble) involves a nonlinearity and a learned vector which increases parameter count and latency moderately. Concat attention (orange bubble) doubles the score input dimension due to query and key vector stack, resulting in higher FLOPs, latency, and a larger parameter bubble. Overall, scaled dot-product can be a good choice if compute or latency-bounded as it appears fast and cheap. We can choose additive or concatenation attention only when they provide higher accuracy to justify the extra cost. Fig 5 presents performance of our model using different number of attention heads. During these experiments, the scale-dot attention method is used. Practically, a moderate head count of four to five is a favorable latency and FLOPs trade-off. Using more than five attention heads should be considered only when they deliver measurable accuracy that justify the added runtime and memory budget.
Table 4 reports the statistical validation of proposed NL-MHA-LSTM model across five independent runs with different random seeds. RMSE increases from 0.093m at 1s to 0.675m at 5s, reflecting growing prediction uncertainty over longer horizons. The narrow confidence intervals across all metrics confirm that the model converges consistently and its performance is reproducible rather than dependent on a single favorable initialization. Lateral errors remain consistently lower than longitudinal errors across all horizons, which is expected given the structured lane-keeping behavior characteristic of highway driving.
4.1 Performance evaluation of NL-MHA-LSTM against existing attention-based models
Table 5 benchmarks NL-MHA-LSTM against several attention-based models on the HighD dataset, including AS-LSTM [28], PiP-LSTM [51], NLS-LSTM [52], MHA-LSTM [23], MHA-LSTM(+f) [23], SGCN [53], S-LSTM [54], CS-LSTM [55], LS-LSTM [56], Dual Transformer [27], Hybrid-LSTM [45], GRANP [43], and GAT-TR-LSTM-LSTM [42]. AS-LSTM combines social pooling with a self-attention mechanism built on scaled dot-product attention, enabling it to learn cooperative driving behaviors and inter-vehicle interactions effectively. This design places it among the top performers at shorter horizons (1-4s). The performance of AS-LSTM becomes extremely poor for longer prediction horizons (5s: 0.93), however, due to the limitation of its self-attention mechanism for preserving temporal dependencies over extended prediction horizons.
The MHA-LSTM(+f) model incorporates additional kinematic and class-type features with multi-head attention mechanism, including velocity, acceleration, and vehicle class, using scaled dot-product attention. This enriched feature set translates into strong short-range predictive performance, achieving an RMSE of 0.06m at 1s, placing it second overall at that horizon. Nevertheless, even these models struggle with long-range predictions due to the problems associated with memory capacity and prediction error accumulation. Methods that rely mainly on the target vehicle’s history S-LSTM, CS-LSTM, PiP-LSTM and Hybrid-LSTM exhibit markedly higher RMSE as 3.41, 3.27, 2.36 and 2.38 at 5s, underscoring the limitations of models without explicit interaction modeling. In contrast, our NL-MHA-LSTM achieves 0.41 at 5s, representing a 53.4% reduction in comparison to the next-best with 0.88 RMSE (LS-LSTM). The graph and transformer variants such as SGCN and Dual Transformer are competitive at short horizons, these approaches degrade more steeply as the prediction window extends. Among these models, NL-MHA-LSTM yields the most stable long-range prediction performance (5s: 0.41). These results are reported using the best-performing checkpoint during training process.
Fig 6 illustrates how the RMSE of each model evolves across prediction horizons ranging from 1-5s. Our model shows better performance for longer horizons providing the best results on all horizons. It proves the significance of developing spatiotemporal modeling techniques and using proper features for making accurate predictions on longer horizons. The experimental results emphasize the weaknesses of the current approaches in dealing with long-term predictions. Our proposed model is stable on all horizons and demonstrates the best performance on longer horizons. The reason is improved feature representation, optimized training and attention mechanism.
4.2 Scenario-specific comparison
The evaluation compares proposed NL-MHA-LSTM against a set of established attention-based models for vehicle trajectory prediction. NST-LSTM [49] is not included in above comparison (Table 5) because it uses a fixed uniform pooling rather than an attention mechanism. A dedicated scenario-specific analysis is performed for comparison between NST-LSTM and NL-MHA-LSTM.
To better understand where the performance gains actually come from in both models, Table 6 outlines lane-following and lane change scenarios extracted from the HighD dataset. Out of 110,516 trajectories, 24,617 fall into the lane-following category; where interactions between vehicles are sparse but they carry significant weight. The remaining 85,899 trajectories involve lane-change scenarios, where interactions are dense. Table 7 shows the RMSE for both scenarios. In lane-following scenarios, NL-MHA-LSTM shows 1.50m RMSE and NST-LSTM reduces this by approximately 50%, achieving 0.75m. In lane change scenarios, the trend reverses: NST-LSTM achieves an RMSE of 0.52m compared to 0.70m for NL-MHA-LSTM. NL-MHA-LSTM demonstrates better performance in lane-following scenarios, where the ability to selectively assign attention weights to the most relevant neighboring vehicles is critical, while NST-LSTM performs better in lane-change scenarios where uniform pooling is sufficient. The overall RMSE favors NST-LSTM simply because HighD is dominated by dense lane-change scenarios.
Table 8 compares the computational complexity of NST-LSTM and the three NL-MHA-LSTM attention variants. NST-LSTM is the most lightweight, with an inference time of 8.015ms and 493.83K parameters, due to the absence of an attention module. The NL-MHA-LSTM variants are already discussed in section 4. The higher computational cost of NL-MHA-LSTM over NST-LSTM is a direct consequence of the attention mechanism, which enables learned neighbor weighting and improved model explainability.
4.3 Limitations
The proposed NL-MHA-LSTM has several limitations. The model is trained and evaluated exclusively on the highway dataset, highD. This limits the scope of applicability of NL-MHA-LSTM in more complicated or unstructured scenarios like urban intersections, roundabouts and high occlusion traffic. Extending evaluation to other datasets is not straight-forward, most alternative datasets require substantial custom feature engineering. The model predicts the future position of the target vehicle only, without predicting velocity or acceleration, which limits its direct utility for motion-planning modules that require full kinematic state estimates rather than position alone. In addition, the decoder produces a single deterministic trajectory rather than a distribution over plausible future. The real driving behavior is inherently multi-modal; a vehicle approaching a gap in traffic may accelerate through it, cautiously merge, or yield entirely depending on contextual patterns that are difficult to capture from position alone.
5. Conclusion and future work
This paper introduces an attention-based multi-head spatiotemporal attention model that captures the interaction among vehicles. The proposed model captures long-range spatiotemporal interactions between the vehicles without any constraints that limit conventional neighborhood-based approaches. The high-order interactions are processed by deep LSTM encoder. We subsequently use the attention mechanism to weight encoded interaction features by relative position and motion. It emphasizes key vehicles and delivers strong position-aware features. The experimental evaluations show that the approach suggested performs much better than existing techniques concerning RMSE, longitudinal error, and lateral error. Besides, the ablation study conducted for attention types, head numbers, and performance parameters has shown that parameter selection is crucial for an attention mechanism. The evaluation findings suggest that combining non-local spatiotemporal interaction modeling with a multi-head attention mechanism and a deep LSTM architecture collectively drives the prediction accuracy gains observed in highway driving scenarios.
Future work will evaluate this framework on unstructured traffic environments, such as intersections and roundabouts with the goal of improving model’s generalization and prediction capability. For generalization, the proposed model will be tested on other publicly available datasets like INTERACTION [57], nuScenes [58], and V2X-Sim [59,60] including more diverse traffic scenarios. For a richer feature space, kinematics (speed, acceleration) and semantics (vehicle type) will be integrated in the model. In addition, future work will include generating a distribution of possible future trajectories instead of one deterministic output to reflect the inherent multimodality of real-world driving.
References
- 1.
Messaoud K, Yahiaoui I, Verroust-Blondet A, Nashashibi F. Relational recurrent neural networks for vehicle trajectory prediction. 2019 IEEE Intelligent Transportation Systems Conference (ITSC); 2019. p. 1813–8. https://doi.org/10.1109/ITSC.2019.8916887
- 2.
Paul A, Chilamkurti N, Daniel A, Rho S. Intelligent transportation systems; 2017. p. 21–41. https://doi.org/10.1016/B978-0-12-809266-8.00002-8
- 3. Janai J, Güney F, Behl A, Geiger A. Computer vision for autonomous vehicles: problems, datasets and state of the art. Found Trends Comput Graph Vis. 2020;12(1–3):1–308.
- 4.
Bai X, Hu Z, Zhu X, Huang Q, Chen Y, Fu H, et al. TransFusion: robust LiDAR-camera fusion for 3D object detection with transformers. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022. p. 1080–9. https://doi.org/10.1109/CVPR52688.2022.00116
- 5.
Liu Z, Tang H, Amini A, Yang X, Mao H, Rus DL, et al. BEVFusion: multi-task multi-sensor fusion with unified bird’s-eye view representation. 2023 IEEE International Conference on Robotics and Automation (ICRA); 2023. p. 2774–81. https://doi.org/10.1109/ICRA48891.2023.10160968
- 6.
Liu Y, Fan Y. A review of vehicle trajectory prediction methods in driverless scenarios. 2024 9th International Conference on Intelligent Informatics and Biomedical Sciences (ICIIBMS), vol. 9; 2024. p. 293–9. https://doi.org/10.1109/ICIIBMS62405.2024.10792666
- 7.
Servan-Schreiber D, Cleeremans A, McClelland J. Learning sequential structure in simple recurrent networks. In: Touretzky D, editor. Advances in neural information processing systems. vol. 1. Morgan-Kaufmann; 1988. Available from: https://proceedings.neurips.cc/paper_files/paper/1988/file/9dcb88e0137649590b755372b040afad-Paper.pdf
- 8. Sheng Z, Xu Y, Xue S, Li D. Graph-based spatial-temporal convolutional network for vehicle trajectory prediction in autonomous driving. IEEE Trans Intell Transp Syst. 2022;23(10):17654–65.
- 9. Hochreiter S, Schmidhuber J. Long short-term memory. Neural Comput. 1997;9(8):1735–80.
- 10.
Yang T, Nan Z, Zhang H, Chen S, Zheng N. Traffic agent trajectory prediction using social convolution and attention mechanism. 2020 IEEE Intelligent Vehicles Symposium (IV); 2020. p. 278–83. https://doi.org/10.1109/IV47402.2020.9304645
- 11. Fernando T, Denman S, Sridharan S, Fookes C. Soft + Hardwired attention: an LSTM framework for human trajectory prediction and abnormal event detection. Neural Netw. 2018;108:466–78. pmid:30317132
- 12. Soni D, Kumar N. Machine learning techniques in emerging cloud computing integrated paradigms: a survey and taxonomy. J Netw Comput Appl. 2022;205:103419.
- 13.
Ip A, Irio L, Oliveira R. Vehicle trajectory prediction based on LSTM recurrent neural networks. 2021 IEEE 93rd Vehicular Technology Conference (VTC2021-Spring); 2021. p. 1–5. https://doi.org/10.1109/VTC2021-Spring51267.2021.9449038
- 14. Tian W, Wang S, Wang Z, Wu M, Zhou S, Bi X. Multi-modal vehicle trajectory prediction by collaborative learning of lane orientation, vehicle interaction, and intention. Sensors (Basel). 2022;22(11):4295. pmid:35684916
- 15. Dai S, Li L, Li Z. Modeling vehicle interactions via modified LSTM models for trajectory prediction. IEEE Access. 2019;7:38287–96.
- 16.
Yang Z, Liu D, Ma L. Vehicle trajectory prediction based on LSTM network. 2022 International Conference on Artificial Intelligence and Computer Information Technology (AICIT); 2022. p. 1–4. https://doi.org/10.1109/AICIT55386.2022.9930177
- 17. Katariya V, Baharani M, Morris N, Shoghli O, Tabkhi H. DeepTrack: lightweight deep learning for vehicle trajectory prediction in highways. IEEE Trans Intell Transp Syst. 2022;23(10):18927–36.
- 18.
Karimzadeh M, Aebi R, Souza AMd, Zhao Z, Braun T, Sargento S, et al. Reinforcement learning-designed LSTM for trajectory and traffic flow prediction. 2021 IEEE Wireless Communications and Networking Conference (WCNC); 2021. p. 1–6. https://doi.org/10.1109/WCNC49053.2021.9417511
- 19. Tang L, Yan F, Zou B, Li W, Lv C, Wang K. Trajectory prediction for autonomous driving based on multiscale spatial‐temporal graph. IET Intell Transp Syst. 2022;17.
- 20. Chen J, Zhou S, Wang W, Hu Y, Li J, He B, et al. Vehicle dynamics and interaction for trajectory prediction and traffic control. ACM Trans Auton Adapt Syst. 2025;20(2):1–19.
- 21. Min H, Xiong X, Wang P, Zhang Z. A hierarchical LSTM-based vehicle trajectory prediction method considering interaction information. Automotive Innov. 2024;7(1).
- 22.
Vemula A, Muelling K, Oh J. Social attention: modeling attention in human crowds. 2017. https://doi.org/10.48550/arXiv.1710.04689
- 23. Messaoud K, Yahiaoui I, Verroust-Blondet A, Nashashibi F. Attention based vehicle trajectory prediction. IEEE Trans Intell Veh. 2021;6(1):175–85.
- 24. Gui Z, Wang X, Li W. Dynamic perception-based vehicle trajectory prediction using a memory-enhanced spatio-temporal graph network. ISPRS IJGI. 2024;13(6):172.
- 25. Jin L, Liu X, Wang Y, Han Z, Guo B, Luo G, et al. Multi-modality trajectory prediction with the dynamic spatial interaction among vehicles under connected vehicle environment. Sci Rep. 2024;14(1):2873. pmid:38311625
- 26. Wang T, Fu Y, Cheng X, Li L, He Z, Xiao Y. Vehicle trajectory prediction algorithm based on hybrid prediction model with multiple influencing factors. Sensors (Basel). 2025;25(4):1024. pmid:40006253
- 27. Gao K, Li X, Chen B, Hu L, Liu J, Du R, et al. Dual transformer based prediction for lane change intentions and trajectories in mixed traffic environment. IEEE Trans Intell Transp Syst. 2023;24(6):6203–16.
- 28. Qiao S, Gao F, Wu J, Zhao R. An enhanced vehicle trajectory prediction model leveraging LSTM and social-attention mechanisms. IEEE Access. 2024;12:1718–26.
- 29. Cai Y, Wang Z, Wang H, Chen L, Li Y, Sotelo MA, et al. Environment-attention network for vehicle trajectory prediction. IEEE Trans Veh Technol. 2021;70(11):11216–27.
- 30. Lu Y, Wang W, Hu X, Xu P, Zhou S, Cai M. Vehicle trajectory prediction in connected environments via heterogeneous context-aware graph convolutional networks. IEEE Trans Intell Transp Syst. 2023;24(8):8452–64.
- 31.
Huang Z, Mo X, Lv C. Multi-modal motion prediction with transformer-based neural network for autonomous driving. 2022 International Conference on Robotics and Automation (ICRA); 2022. p. 2605–11. https://doi.org/10.1109/ICRA46639.2022.9812060
- 32.
Sun H, Sun F. Social-transformer: pedestrian trajectory prediction in autonomous driving scenes; 2022. p. 177–90. https://doi.org/10.1007/978-981-16-9247-5_13
- 33.
Xie C, Li Y, Liang R, Dong L, Li X. Synchronous Bi-directional pedestrian trajectory prediction with error compensation. In: Wang L, Gall J, Chin TJ, Sato I, Chellappa R, editors. Computer Vision – ACCV 2022. Cham: Springer Nature Switzerland; 2023. p. 699–715. https://doi.org/10.1007/978-3-031-26351-4_42
- 34.
Zhou Z, Ye L, Wang J, Wu K, Lu K. HiVT: hierarchical vector transformer for multi-agent motion prediction. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2022. p. 8813–23. https://doi.org/10.1109/CVPR52688.2022.00862
- 35. Wang Z, Guo J, Hu Z, Zhang H, Zhang J, Pu J. Lane transformer: a high-efficiency trajectory prediction model. IEEE Open J Intell Transp Syst. 2023;4:2–13.
- 36.
Ngiam J, Caine B, Vasudevan V, Zhang Z, Chiang HT, Ling J, et al. Scene transformer: a unified multi-task model for behavior prediction and planning. 2021. https://doi.org/10.48550/arXiv.2106.08417
- 37. Mo X, Huang Z, Xing Y, Lv C. Multi-agent trajectory prediction with heterogeneous edge-enhanced graph attention network. IEEE Trans Intell Transp Syst. 2022;23(7):9554–67.
- 38. Liang Y, Zhao Z. NetTraj: a network-based vehicle trajectory prediction model with directional representation and spatiotemporal attention mechanisms. IEEE Trans Intell Transp Syst. 2022;23(9):14470–81.
- 39. Hasan F, Huang H. MALS-Net: a multi-head attention-based LSTM sequence-to-sequence network for socio-temporal interaction modelling and trajectory prediction. Sensors. 2023;23(1).
- 40.
Tang X, Kan M, Shan S, Ji Z, Bai J, Chen X. HPNet: dynamic trajectory forecasting with historical prediction attention. 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2024. p. 15261–70. https://doi.org/10.1109/CVPR52733.2024.01445
- 41. Cheng D, Gu X, Qian C, Du C, Wang J. Vehicle trajectory prediction with interaction regions and spatial–temporal attention. IEEE Access. 2023;11:130850–9.
- 42. Gao Y, Yang K, Yue Y, Wu Y. A vehicle trajectory prediction model that integrates spatial interaction and multiscale temporal features. Sci Rep. 2025;15(1):8217. pmid:40065111
- 43.
Luo Y, Chen K, Zhu M. GRANP: a graph recurrent attentive neural process model for vehicle trajectory prediction. 2024 IEEE Intelligent Vehicles Symposium (IV); 2024. p. 370–5. https://doi.org/10.1109/IV55156.2024.10588741
- 44. Li H, Ren Y, Li K, Chao W. Trajectory prediction with attention-based spatial–temporal graph convolutional networks for autonomous driving. Appl Sci. 2023;13(23):12580.
- 45. Wang M, Zhang L, Chen J, Zhang Z, Wang Z, Cao D. A hybrid trajectory prediction framework for automated vehicles with attention mechanisms. IEEE Trans Transp Electrif. 2024;10(3):6178–94.
- 46. Cui H, Xiao K, Wang H, Zhang X. Lane-flow-learning based autonomous vehicle trajectory prediction using spatial–temporal fusion attention. Inf Sci. 2026;737:123134.
- 47. Gu Z, Xu X, Shen J, Liu R, Zheng J. Vehicle trajectory prediction in highway weaving areas considering interaction-transfer characteristics. IEEE Trans Intell Transp Syst. 2026;27(6):6437–50.
- 48. Baicang G, Hao L, Xiao Y, Yuan C, Lisheng J, Yinlin W. Multi-modal information fusion for multi-task end-to-end behavior prediction in autonomous driving. Neurocomputing. 2025;634:129857.
- 49. Rashid S, Khattak MAK, Akram MU, Ahmad A, Alsagri HS, Alhakbani HAA. Optimizing multi-layer LSTM based on non-local spatiotemporal interactions for vehicle trajectory prediction. IEEE Access. 2025;13:147690–702.
- 50.
Krajewski R, Bock J, Kloeker L, Eckstein L. The highD dataset: a drone dataset of naturalistic vehicle trajectories on German highways for validation of highly automated driving systems. 2018 21st International Conference on Intelligent Transportation Systems (ITSC); 2018. p. 2118–25. https://doi.org/10.1109/ITSC.2018.8569552
- 51.
Song H, Ding W, Chen Y, Shen S, Wang MY, Chen Q. PiP: planning-informed trajectory prediction for autonomous driving. Computer Vision – ECCV 2020: 16th European Conference; 2020 Aug 23–28; Glasgow, UK, Proceedings, Part XXI. Berlin, Heidelberg: Springer-Verlag; 2020. p. 598–614. Available from: https://doi.org/10.1007/978-3-030-58589-1_36
- 52.
Messaoud K, Yahiaoui I, Verroust-Blondet A, Nashashibi F. Non-local social pooling for vehicle trajectory prediction. 2019 IEEE Intelligent Vehicles Symposium (IV); 2019. p. 975–80. https://doi.org/10.1109/IVS.2019.8813829
- 53.
Shi L, Wang L, Long C, Zhou S, Zhou M, Niu Z, et al. SGCN:sparse graph convolution network for pedestrian trajectory prediction. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); 2021. p. 8990–9. https://doi.org/10.1109/CVPR46437.2021.00888
- 54.
Alahi A, Goel K, Ramanathan V, Robicquet A, Fei-Fei L, Savarese S. Social LSTM: human trajectory prediction in crowded spaces. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); 2016. p. 961–71. https://doi.org/10.1109/CVPR.2016.110
- 55.
Deo N, Trivedi MM. Convolutional social pooling for vehicle trajectory prediction. 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW); 2018. p. 1549–98. https://doi.org/10.1109/CVPRW.2018.00196
- 56. Yu D, Lee H, Kim T, Hwang S-H. Vehicle trajectory prediction with lane stream attention-based LSTMs and road geometry linearization. Sensors (Basel). 2021;21(23):8152. pmid:34884152
- 57.
Zhan W, Sun L, Wang D, Shi H, Clausse A, Naumann M, et al. Interaction dataset: an international, adversarial and cooperative motion dataset in interactive driving scenarios with semantic maps; 2019. https://doi.org/10.48550/arXiv.1910.03088
- 58.
Caesar H, Bankiti VKR, Lang A, Vora S, Liong VE, Xu Q, et al. nuScenes: A multimodal dataset for autonomous driving; 2019. https://doi.org/10.48550/arXiv.1903.11027
- 59. Li Y, Ma D, An Z, Wang Z, Zhong Y, Chen S, et al. V2X-Sim: multi-agent collaborative perception dataset and benchmark for autonomous driving. IEEE Robot Autom Lett. 2022;7(4):10914–21.
- 60. Khan D, Aslam S, Chang K. Vehicle-to-infrastructure multi-sensor fusion (V2I-MSF) with reinforcement learning framework for enhancing autonomous vehicle perception. IEEE Access. 2025;13:50122–36.