Figures
Abstract
How can we identify stable patterns to accurately predict stock price movements amidst severe market noise? Predicting stock price movements is a crucial problem in financial data mining, and has attracted significant attention from researchers and financial institutions. Although there have been several works on learning patterns from quantitative data to predict stock prices, they encounter critical challenges from data uncertainty and model limitations. Data uncertainty arises in financial data because asset prices are determined by the market. Different participants hold distinct valuations for the same asset, which causes prices to fluctuate. This inherent randomness makes it difficult to identify consistent patterns in the data. Additionally, because of this inherent randomness, existing stock movement prediction models that repeatedly stack more recurrent layers end up propagating the randomness forward. It thus becomes more difficult to capture the complex patterns of stock prices. In this work, we propose Craft (Centroid-based Randomness Smoothing Approach for Stock Forecasting with Transformer Architecture), an accurate stock movement prediction model designed to extract clear patterns in stock data and perform refined forecasting through a Transformer-based architecture. By applying an effective randomness smoothing process, Craft uncovers meaningful and consistent patterns that facilitate accurate predictions. We evaluate Craft on fourteen test settings spanning six real-world datasets and three market regimes. Craft achieves the highest prediction accuracy in twelve of these settings and ranks the second in the other two. Craft also improves the annualized Sharpe ratio by up to 1.0 while reducing the relative maximum drawdown by up to 7.0% points over the best competitor.
Citation: Soun Y, Lee H, Kang U (2026) Accurate stock movement prediction via centroid-based randomness smoothing. PLoS One 21(10): e0345252. https://doi.org/10.1371/journal.pone.0345252
Editor: Mohammad Enamul Hoque, BRAC University, BANGLADESH
Received: February 27, 2026; Accepted: September 10, 2026; Published: October 8, 2026
Copyright: © 2026 Soun et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The minimal data set and accompanying code are publicly available without restriction. The four stock-market datasets generated for this study are deposited at Zenodo under a CC BY 4.0 license and are available via https://doi.org/10.5281/zenodo.21506172. The deposit contains daily OHLCV panels with split- and dividend-adjusted close prices for SP500 (USA.csv), CSI300 (CHN.csv), NI225 (JPN.csv), and EURO50 (EUR.csv). Two further datasets were not generated by this study and are distributed by their original authors: CIKM24, available from the MATCC repository at https://github.com/caozhiy/MATCC, and AAAI24, available from the MASTER repository at https://github.com/SJTU-DMTai/MASTER. All code required to reproduce the reported results, including preprocessing, training, evaluation, and the investment simulation, is available at https://github.com/snudatalab/CRAFT.
Funding: This work was supported by the National Research Foundation of Korea, funded by the Ministry of Science and Information and Communications Technology (RS-2026-25483234 to U.K.), and by the Institute of Information and Communications Technology Planning and Evaluation, funded by the Korea government, Ministry of Science and Information and Communications Technology, under the XVoice: Multi-Modal Voice Meta Learning project (2022-0-00641 to U.K.), the Global Artificial Intelligence Frontier Lab project (RS-2024-00509257 to U.K.), the Artificial Intelligence Graduate School Program at Seoul National University (RS-2021-II211343 to U.K.), and the Artificial Intelligence Star Fellowship Support Program at Seoul National University (RS-2025-25442338 to U.K.). The Institute of Engineering Research and the Institute of Computer Technology at Seoul National University provided research facilities for this work. U. Kang is the corresponding author. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Given price data with inherent randomness, how can we identify consistent patterns and make precise predictions of price movements? Stock movement prediction is one of the fields with the smallest gap between fundamental research and economic benefits, making it a topic of great interest to many researchers and financial institutions. Accurately predicting stock movements has a significant impact on financial investments.
However, stock movement prediction entails challenges commensurate with its economic utility due to the following three main limitations. First, stock prices exhibit fluctuations and a significant level of unexplainable fluctuations due to the diverse range of entities involved in trading. Second, the movement of an individual stock is difficult to explain solely by its own dynamics due to the influence of numerous external variables beyond the company’s inherent factors [1,2]. Third, recent multi-layer time-series approaches can intensify the inherent randomness in stock data by propagating noise through increasingly complex hidden representations, exacerbating uncertainty in the final predictions.
The main limitation of existing stock prediction methods lies in their excessive emphasis on explaining the inherent randomness of stock price data. They often incorporate external variables such as technical indicators [3,4], fundamental factors [5,6], or macroeconomic data [7]—leading to potential overfitting when the feature space grows faster than the available training instances. Others adopt deeper Transformer-based architectures on raw price features [8,9], but random spikes frequently overshadow the model’s capacity to isolate meaningful patterns. When multiple self-attention layers are stacked in the presence of significant data uncertainty, the resulting network can amplify noise rather than filtering it out, making it difficult to distinguish true signals from volatile fluctuations. Although Transformer-based approaches have shown promise in capturing both local and global dependencies, they typically rely on unprocessed raw price inputs. Thus, the deeper these models go, the more likely they propagate spurious variations, ultimately limiting their effectiveness in highly stochastic financial settings.
In this work, we propose Craft (Centroid-based Randomness Smoothing Approach for Stock Forecasting with Transformer Architecture), an accurate stock movement prediction model designed to extract generalized patterns from the randomness of stock movements (see Fig 1). Craft reduces randomness and identifies clear patterns from stock data with the following three main ideas. First, Craft reduces stock-specific randomness in price data by applying centroid-based clustering methods to stock return data and using the centroids of tokens as token embedding. Second, Craft forms a “glocal” (global + local) context by combining each stock’s local return sequence with a global context built from stock-wise correlations at each time step. The model aggregates local patterns of highly correlated stocks by a weighted average and applies time-axis self-attention for precise next-token predictions. Third, Craft leverages the time-axis glocal representation within an end-to-end framework by adding a stock-axis attention module. This dual-axis design not only retains the local next-token patterns, but also captures asymmetric and dynamic cross-stock correlations, thus enabling Craft to accurately predict both the next token and the next price movement for each stock.
By converting each stock’s short-term relative returns into discrete tokens, Craft alleviates inherent randomness and highlights short-term patterns. For each time step, a dynamic global context sequence is constructed from stock-level correlations and then combined with the local token sequence of each stock, yielding refined next-token features. Finally, the Stock-axis Multi-Head Attention Layer applies attention over these features across all stocks, learning which stocks contribute the most to each target prediction. This design enables Craft to capture inter-stock influences more precisely and to accurately predict each stock’s subsequent movement.
Our contributions are summarized as follows:
- Method. We propose Craft, an accurate stock prediction framework. Recent successes in large language models (e.g., GPT) show that a decoder operating on coherent, tokenized sequences can capture generalized patterns more effectively. Motivated by this insight, Craft transforms raw returns into stable tokens, allowing the decoder to focus on generalized patterns in stock movements rather than being confounded by high-frequency noise.
- Accuracy. We conduct extensive evaluations using major real-world stock market datasets (from the United States, China, Europe, and Japan). Craft achieves the highest accuracy in twelve of the fourteen test settings and ranks the second in the other two, improving accuracy by up to 4.1%p and the Matthews correlation coefficient by up to 7.8 points relative to existing baselines.
- Case study. Through visual inspection, we verify that the temporal stock embeddings accurately reflect real-world sector relationships, enabling stocks with similar characteristics at each time step to reference one another. This mechanism allows Craft to incorporate a robust temporal global context, leveraging time-varying cross-stock correlations for improved modeling.
The symbols used in this paper is in Table 1. The code is available at https://github.com/snudatalab/CRAFT, and the datasets are deposited at https://doi.org/10.5281/zenodo.21506172.
Related works
We introduce related works on stock movement prediction in two categories based on how they handle inherent randomness of stock data. The first category is data-centric, encompassing multi-modal integration and feature engineering. The second category is model-centric, which emphasizes inference architectures tailored for financial time-series forecasting or exploiting financial market characteristics [10,11].
Data augmentation
Data augmentation has been widely used to improve predictive accuracy with additional information. In stock forecasting, such augmentation is motivated by the challenge of predicting price movements from limited or partially informative historical data.
One widely adopted technique is to enhance input features of price series. Zhang [12] and Khandelwal et al. [13] combine ARIMA [14] with neural networks, but purely autocorrelation-based methods cannot fully capture complex market dynamics. Another frequent approach is engineering technical indicators [2] or factor modeling using fundamental ratios [1,15], which can be effective but often rely on linear assumptions that fail under nonstationary conditions [16].
More recently, researchers have actively integrated alternative modalities to capture broader market contexts. For instance, pre-trained language models (LLMs) are increasingly utilized to extract rich representations for general time-series forecasting [17,18]. Concurrently, graph-based augmentations employ concept-oriented networks to explicitly model inter-stock correlations [19], and hypergraph neural networks extend this idea to capture higher-order relationships among stocks [20].
Although these data-centric strategies can sometimes improve short-term forecasts, they inevitably retain much of the randomness intrinsic to stock prices. Moreover, if the augmented feature space grows faster than the available data, the curse of dimensionality may obscure patterns in the data. This persistent challenge motivates our approach, which relies exclusively on historical price data to reduce randomness directly by focusing on short-term patterns, thereby avoiding complexities on external variables.
Model refinement
A complementary line of research aims to address the inherent randomness in stock data by refining the predictive model architecture. Since stock prices form a time series, many studies [21,22] leverage RNNs to capture sequential dependencies. Early GRU/LSTM networks helped mitigate vanishing and exploding gradients, enabling more stable training. However, they still exhibit critical limitations:
- Propagation of randomness. When data contain substantial noise and nonstationary characteristics, more recurrent layers may propagate, rather than reduce, randomness.
- Long-range dependencies. Even with gating mechanisms, it is difficult to capture multiple time scales in financial data.
In recent years, Transformer-based models have garnered attention for stock price forecasting. Although conceptually similar to attention-augmented RNNs, Transformers remove recurrence entirely and rely on self-attention to capture local and global dependencies in parallel. Their positional encoding compensates for the lack of sequence ordering in self-attention [23].
Several works exploit Transformer structures: DTML (Data-axis Transformer with Multi-Level contexts) [11] utilizes attentive contexts across time and stock axes, TEGRU (Transformer Encoder Gated Recurrent Unit) [24] integrates news sentiments, and MASTER (MArket-guided Stock TransformER) [9] models cross-time correlations with feature selection. Beyond movement prediction, Transformer variants such as Informer and Autoformer have been benchmarked for pairs trading on equity and cryptocurrency markets [25]. While these methods often improve robustness, large model sizes and deep layers risk amplifying random fluctuations, making it hard to overcome the volatility of financial markets [26].
In contrast, our proposed Craft reduces noise earlier in the process using centroid-based tokenization and employs an end-to-end architecture that integrates self-attention across both time axis and stock axis. This approach minimizes the impact of random fluctuations while maintaining predictive accuracy. Rather than stacking numerous layers on raw price data, Craft focuses on short-term local patterns and inter-stock correlations, mitigating hidden-representation noise and capturing crucial temporal signals in dynamic markets.
Proposed method
We propose Craft (Centroid-based Randomness Smoothing Approach for Stock Forecasting with Transformer Architecture), an accurate stock movement prediction method. Our formal problem definition is as follows.
Problem 1 (Stock movement prediction.). Given the historical price vector for an individual stock i across multiple stocks over a lookback window w ending at the current time step t, predict the discrete price movement
at day t + 1, where
if
and 0 otherwise.
Overview
Challenges and ideas.
Raw price data often exhibit notable volatility, as different market participants assign varying valuations to the same asset. Inspired by large language models (LLMs) [27,28] which tokenizes text into coherent units for next-token prediction, Craft formulates this forecasting process by transforming continuous, noisy price segments into stable, discrete representations to uncover consistent market dynamics. To effectively solve this task, we identify three main challenges and propose the corresponding solutions:
- C1 How can we extract clear patterns from data with inherent randomness?
Stock prices often fluctuate; this randomness can mask meaningful local patterns crucial for accurate predictions. - C2 How can we consider complex and dynamic inter-stock correlations to enhance predictive power?
Even after identifying local patterns for each individual stock, overall movements can still be influenced by wider market or sector-based trends. A framework is needed to integrate correlated stock signals effectively. - C3 How can we utilize dynamic correlation patterns for accurate price movement prediction?
Historical correlation matrices capture average relationships, but actual cross-stock dependencies shift over time. Leveraging evolving correlations alongside local price patterns remains a key challenge for accurate price movement prediction.
We address the above challenges with the following three ideas.
I1 Centroid-based price embedding (Token generation).
To address C1, Craft clusters noisy return data into short-term “centroid-based” tokens, analogous to how large language models tokenize text. By mitigating high-frequency volatility, we obtain coherent local price segments that clarify time dependencies.
I2 Related context aggregation via dynamic stock embedding (Related context aggregation via dynamic stock embedding).
To address C2, Craft computes correlation-based stock embeddings at each time step and forms a global context that aggregates signals from correlated stocks. This allows Craft to capture inter-stock dependencies and unify local patterns with broader market influences.
I3 Two-stage attention in an end-to-end framework (Inter-stock glocal context aggregation via end-to-end framework).
To address C3, Craft employs a dual attention mechanism. A temporal-axis attention captures local time dependencies and predicts the next token for each stock. A subsequent stock-axis attention integrates cross-stock correlations for refined price movement predictions. By training these attention modules in a joint manner, the model balances temporal and cross-stock patterns.
The overall procedure of Craft (Fig 1) consists of three main modules. First, the Token Generation module takes the raw price vectors as input, applies relative normalization, and uses k-means clustering to convert them into a stable sequence of centroid tokens (I1). Subsequently, the model extracts temporal stock embeddings via SVD to construct a dynamic global context for each stock, which is then processed alongside the local tokens through a time-axis self-attention mechanism to capture temporal dependencies (I2). Finally, the representations of all stocks at the last time step are fed into a stock-axis multi-head self-attention module to exchange information among correlated peers, yielding the final price movement prediction in an end-to-end manner (I3).
Token generation
We transform each short-term price segment into a discrete token. Our objective is to suppress high-frequency randomness and expose stable local patterns to the Transformer. Instead of directly using the noisy raw price returns, our idea is to normalize the short-term price segments across stocks and times, and replace them with representative short-term price patterns where inherent randomness is effectively smoothed.
Relative normalization.
Let be the price of stock i at day t. We choose a small window size w (e.g., fewer than 20 days), and for each
define the j-day relative return:
capturing the ratio between the price at day and the price at day
. This yields a
-dimensional vector of relative returns:
which offsets scale discrepancies and reduces nonstationary effects.
Centroid tokenization.
Stock prices often exhibit significant randomness on top of any baseline trend. A standard idealization is the random walk with drift, where an individual asset price evolves as
To suppress the noise while retaining a small set of recurring short-term patterns, we adopt a vector-quantization scheme analogous to VQ-VAE [29]: each short-term return vector is replaced by its nearest centroid, yielding a discrete token. We apply k-means clustering to the set of vectors
across all stocks i and times t within the training period. The resulting cluster centroids are
which we also treat as token embeddings. Let
assign each vector to its nearest centroid. Hence, for each
, the final token embedding is
We use a grid search over to find the optimal k separately for each market regime, fitting the centroids on the regime’s training split and selecting k on its validation split. The standard elbow [30] and silhouette [31] methods often fail to find stable solutions for noisy stock price data with a large k [32,33]. Consequently, we select the k that yields the highest average validation accuracy for stock movement prediction. The detailed sensitivity analysis and selection process for k are presented in the Experiment section.
Centroid-token sequence.
After assigning to a token embedding
for each stock i and day t, we form a sequence of centroid tokens:
Here, is the sequence length that the Transformer decoder uses as the number of tokens.
Next-token prediction reframing.
Instead of forecasting the raw price , we aim to predict
which parallels “next-token” modeling in language tasks. This lets the Transformer attend to stable local patterns rather than random spikes in the original time series.
In practice, a relatively small window size w (fewer than 20 trading days) further reduces high-frequency fluctuations, so these locally consistent tokens provide robust features for downstream prediction and portfolio selection.
Related context aggregation via dynamic stock embedding
We construct a time-varying global context for each stock from temporal cross-stock correlations. Inter-stock dependencies change with market regimes and sector rotations. Per-stock signals or static factor loadings do not represent the contemporaneous market structure. We obtain spectral embeddings via Eigendecomposition of a rolling cross-stock correlation matrix and use them to aggregate related stocks at each time step. This follows the factor-model tradition of extracting latent factors from return correlations [1,15].
Daily return.
For each stock i at day t, let
be the daily return of i at t which captures the change of the daily closing price.
Temporal stock embedding via SVD.
Let N be the number of stocks in our universe. At each time t, we collect the daily returns of the 60 trading days up to (roughly 3 months) into a matrix
where the i-th row, , corresponds to stock i’s returns over that period. We then construct an
correlation matrix:
where each entry is the sample correlation coefficient between the past 60 daily returns of stocks i and j (days
to
) at time t. Applying SVD to
yields
where contains the top
left singular vectors. We define the stock embedding matrix
, so the embedding for stock i at time t is the i-th row of
, denoted as
.
Global context via embedding similarities.
To incorporate inter-stock relationships, we derive a global context for each stock i by aggregating embeddings of other stocks
. Specifically, let
Here, is the Euclidean distance between embeddings for stocks i and j, and
normalizes these distances so that
. This ensures that related (similar) stocks contribute more strongly to the context of stock i. Both
and
are rows of the same matrix
, so any column-wise sign flip from the SVD affects both rows equally and cancels inside the difference
.
Time-axis self-attention.
We predict future price patterns for each stock i from:
- local token sequence
(from Token generation), and
- global context sequence
We concatenate these along the feature dimension:
Next, we apply a linear projection via a trainable matrix to refine feature representations and unify the dimension:
Craft adopts a decoder-only masked self-attention approach, ensuring that for the -th token in a sequence, subsequent tokens
remain hidden.
Self-attention operation.
We project into queries, keys, and values:
where . The attention output is
Next, we apply a residual connection and an MLP to obtain
Here, d is the dimension used to represent temporal contexts. This representation encodes temporal dependencies at each time step. Moreover, reflects both individual stock’s local patterns and related stocks’ global context, thus providing a “glocal” representation essential for accurate forecasting.
Inter-stock glocal context aggregation via end-to-end framework
We aggregate cross-stock information along the stock axis to produce final price movement predictions from a glocal (global + local) representation. Methods that rely only on per-stock information or assume fixed correlations miss time-varying relations [8,10,34]. Existing cross-stock attention models [9,11] typically focus on exchanging intermediate hidden contexts. However, Craft takes a distinct approach by leveraging a glocal context to directly model realistic market correlations among the final price movement representations. Specifically, we stack the last step vectors of all stocks and apply multi-head self-attention along the stock axis to exchange information among correlated peers. We train this classifier jointly with the time-axis next-token objective in an end-to-end framework.
Next price movement prediction.
Recall that each stock i has a time-axis representation
obtained from Related context aggregation via dynamic stock embedding. We focus on the final time step’s output , which captures the most recent local context for stock i. Collecting these vectors from all N stocks yields
This matrix serves as the input to a stock-axis multi-head self-attention, allowing Craft to refine each stock’s direction forecast by leveraging the correlations among these final-step embeddings.
Stock-axis self-attention.
Let M be the number of heads, and let be the dimension per head. We project
into queries, keys, and values for each head
:
where . The scaled dot-product attention is
We concatenate these M outputs along the feature dimension:
We apply a linear transformation for refinement and add a residual connection:
Let be the i-th row of
. Finally, we map each
to a 2-dimensional logit vector
then apply a softmax:
Here, is the probability that stock i’s next return is predicted to be positive.
Next price movement prediction loss.
Let be the ground-truth upward/downward label of stock i at time t + 1, and
be the predicted class. We define
Total prediction loss.
Combining the time-axis loss Ltime (from Related context aggregation via dynamic stock embedding) with the stock-axis loss gives
where balances next-token prediction and final price movement classification. By minimizing Ltotal, our end-to-end framework learns
for accurate centroid-token forecasting, while
leverages inter-stock attention for robust price movement prediction.
Experiments
We conduct extensive experiments on real-world datasets to provide answers to the following questions.
- Q1. Stock price movement prediction performance (see Stock price movement prediction performance (Q1)). How accurate is Craft compared to previous state-of-the-art competitors?
- Q2. Investment simulation (see Investment simulation (Q2)). Can Craft generate profitable returns, and does it maintain consistent performance across stock markets in multiple countries?
- Q3. Stock embedding quality (see Stock embedding quality (Q3)). Does the stock embedding, derived via SVD for Craft’s dynamic global context, effectively capture correlated stocks so that their price information is reflected accordingly?
- Q4. Ablation study (see Ablation study (Q4)). How do the key modules of Craft contribute to improvements in predictive performance?
Experimental settings
Datasets
We use six datasets evaluated across three distinct market regimes, as summarized in Table 2: the COVID-19 Crisis, the AI Rally, and the Recent Volatile period. Notably, the experiments for the Recent Volatile period comprehensively include both the newly proposed datasets and the existing public benchmarks. For the public datasets, the AAAI24 dataset is employed by Li et al. [9] in a data augmentation approach that leverages external macro data, while the CIKM24 dataset is used by Cao et al. [8] for an incremental learning-based model refinement methodology. We adopt the preprocessed versions directly from their official repositories, retaining the same training, validation, and test splits. SP500, CSI300, NI225, and EURO50 are new public benchmark datasets that we collect from the Yahoo Finance public feed via the yfinance Python library, covering the US, China, Japan, and Europe stock markets, respectively. Downloads were finalized in March 2025. We use the split- and dividend-adjusted closing prices. To avoid potential sector biases or small-cap effects, we select the union of the top 20% of stocks by market capitalization or by trading volume within each sector, restricted to stocks continuously listed from 2020−01 through the download date so that the dataset contains no delisting event, forming our final dataset. This procedure yields a more representative and stable set of stocks for our datasets. The three regime-specific test windows are nested temporal splits of this single fixed universe. This continuous-listing filter introduces survivorship bias, but applies uniformly to all evaluated methods, so the relative performance ranking among models is not affected by this universe choice.
Competitors
We compare Craft with the following baselines for stock movement prediction:
- LSTM [34] is a simple baseline that employs a standard LSTM to predict stock movements.
- ALSTM [35] incorporates temporal attentive contexts into an LSTM, highlighting the most influential time steps for price predictions.
- PatchTST [36] is a general time-series forecasting method that converts the input into short-term patches and leverages a Transformer architecture for effective prediction.
- Informer [37] is a long-sequence forecasting Transformer that uses ProbSparse self-attention to reduce the quadratic attention cost.
- Autoformer [38] replaces dot-product attention with an auto-correlation mechanism and an internal series-decomposition block.
- FEDformer [39] performs attention in the frequency domain and adds a mixture-of-experts decomposition of trend and seasonal components.
- MATCC [8] explicitly extracts overarching market trends and decomposes stock data into trend and fluctuation components, further exploiting cross-time correlations for robust prediction.
- MASTER [9] is a market-guided stock Transformer that captures momentary, cross-time stock correlations and leverages market information for automatic feature selection.
- MA+LSTM, EMA+LSTM, ARIMA+LSTM, Wavelet+LSTM, and Kalman+LSTM apply a denoising or smoothing step (moving average, exponential moving average, ARIMA, wavelet thresholding, or a Kalman filter) to the raw prices before a standard LSTM.
- ARMA, ARIMA-GARCH (the sign of the one-step forecast), Momentum, and Reversal are classical statistical and factor baselines.
Informer, Autoformer, and FEDformer are regression forecasters. To fit the movement-prediction setting, we append a sigmoid output layer that converts their forecasts into up/down probabilities.
The five preprocessing+LSTM baselines isolate the effect of denoising alone, which Craft’s centroid tokenization is designed to surpass. Each smoothing step uses a standard daily-frequency configuration fixed across all datasets: a 5-day simple/exponential moving-average window [40], an ARIMA(1,1,1) specification [14,41], Haar wavelet denoising at two levels with soft universal thresholding [42,43], and a local-level Kalman filter with maximum-likelihood noise variances [44]. Fixing these common defaults rather than tuning per dataset isolates the smoothing mechanism and avoids a per-dataset tuning advantage for the denoising baselines. The four classical baselines (ARMA, ARIMA-GARCH, momentum, and reversal) are deterministic, so each is reported from a single run without a standard deviation.
Evaluation metrics
We report the mean and standard deviation over five random seeds (from 0 to 4) for ACC and MCC. We use four evaluation metrics:
- Accuracy (ACC) is the fraction of daily movements correctly predicted by each model.
- Matthews Correlation Coefficient (MCC) is a balanced metric that takes into account the overall distribution of correct predictions. Given the counts of true positives, true negatives, false positives, and false negatives, MCC is defined as
where tp and tn refer to correctly predicted upward and downward movements, and fp and fn denote incorrect predictions.
- Annualized Sharpe Ratio (ASR) is the annualized ratio of the sample mean of daily test returns to the sample standard deviation of those returns over the test period.
- Relative Maximum Drawdown (RMDD) is the largest percentage drop from a running peak of the cumulative return over the test period.
Hyperparameters.
We search the hyperparameters of Craft as follows: the input data length , the input token sequence length
, the hidden layer size
, the number of attention layers
, the number of attention heads
, the loss balancing parameter
, and the learning rate
. The SVD stock-embedding dimension is fixed at
for all datasets rather than searched. Among these, the hidden layer size d, the loss balancing
, and the learning rate
exert the strongest influence on Craft’s validation performance. S1 Table reports the measured final-epoch magnitudes of the two loss terms, which sit on clearly different scales and are balanced by
. We report their ACC and MCC sensitivity profiles in Figs 2–7. Guided by these profiles, we select per regime a parameter set that performs robustly across the evaluated datasets, favoring general-purpose choices over extreme grid points. For competing methods, we simply use their default settings from the publicly available repositories. MASTER and MATCC further receive their original input features, namely the Alpha158 stock indicators together with the 63-dimensional market-feature block that their architectures consume, computed from each dataset’s own market index. Craft instead uses returns alone. Complete training details and per-dataset selected hyperparameters are documented in the released code at https://github.com/snudatalab/CRAFT.
Each panel corresponds to one dataset. Lines distinguish the three evaluated test regimes for our four new datasets (top two rows). CIKM24 and AAAI24 (bottom row) are evaluated only under the Recent Volatile regime because their fixed test windows do not cover the COVID-19 Crisis or AI Rally periods (cf. Table 2). The remaining hyperparameter-sensitivity figures use the same panel layout.
Implementation.
We train all models for 200 epochs on four NVIDIA GTX 1080 Ti GPUs (11 GB each), with an Intel Xeon Silver 4214 CPU and 503 GB of system memory. Table 3 summarizes Craft’s per-dataset compute budget. Trained in the same environment, MASTER is the lightest competitor at roughly 30–40 minutes per regime, whereas MATCC is comparable to or slightly heavier than Craft at roughly 2–3 hours. Craft’s training cost is therefore in line with existing market Transformers.
Stock price movement prediction performance (Q1)
Table 4 reports the fraction of up-days in each test period. Because the base rates stay near 0.50, an accuracy above 0.50 reflects genuine predictive signal rather than an exploitable label imbalance. The Chinese markets (CSI300, AAAI24) are mildly down-skewed, so a constant up-prediction would score below 0.50.
S2 Table reports the full per-stock up-day rate distribution, whose medians likewise lie near 0.50.
We evaluate the predictive capabilities of Craft across three distinct market regimes to verify its robustness: the COVID-19 Crisis, the AI Rally, and the recent volatile period. Tables 5–7 report the three regimes in turn. Craft attains the highest ACC and MCC in every dataset and regime except the following. In the COVID-19 crisis, MASTER attains a higher ACC and MCC on CSI300. In the recent volatile period, FEDformer attains a higher ACC and MCC on NI225, and MATCC attains a higher MCC on AAAI24, while ALSTM matches rather than trails the MCC of Craft on CIKM24.
By effectively smoothing high-frequency noise through centroid-based tokenization and capturing dynamic inter-stock correlations, our model maintains robust predictive power regardless of extreme macroeconomic shifts.
1) COVID-19 crisis period (2020−01
2020−03).
We analyze performance during the extreme volatility of the pandemic onset. Table 5 summarizes the results. In this panic-driven market, Craft generally demonstrates superior stability, achieving the highest ACC and MCC in the US, Japan, and Europe. Notably, in the Chinese market (CSI300), MASTER achieves a slightly higher ACC (0.552) than Craft (0.545). This suggests that MASTER’s formulation aggressively fits the specific policy-driven dynamics of the Chinese market during this anomaly. However, Craft outperforms MASTER in all other regions, confirming that our approach effectively suppresses high-frequency noise to maintain robust performance during global turbulence.
2) AI rally period (2024−01
2024−06).
Next, we examine the performance during the strong bullish trend driven by the technology sector (Table 6). In this regime, Craft consistently achieves the highest ACC and MCC across all evaluated datasets. By effectively capturing the global context of tech-driven momentum via dynamic stock embeddings, Craft successfully identifies and leverages the upward trends better than competitors.
3) Recent volatile period.
Finally, Table 7 presents the comprehensive performance on all six datasets covering the full test periods, which natively embody recent volatile market conditions. Craft achieves the highest ACC on all evaluated benchmarks except NI225, where FEDformer’s moving-average series decomposition transfers better to this market. Unlike MATCC, which uses a RWKV-based model, Craft combines a token-based Transformer decoder with a dynamically weighted related-stock context. This architecture provides stronger robustness against random spikes, resulting in superior ACC and MCC across most datasets. Because the market-feature block in MASTER and MATCC is engineered for market-specific dynamics, its contribution varies across markets. This variation is the largest for MASTER on the US benchmarks (CIKM24 and SP500), where its MCC drops to near or slightly below zero. Among the general-purpose Transformers, Autoformer and FEDformer apply moving-average series decomposition, which suppresses part of the high-frequency noise and places them above Informer. Because equity prices exhibit weak and unstable seasonality [45], this periodic decomposition transfers only partially to daily movement prediction. Among the preprocessing baselines, the adaptive smoothers EMA+LSTM and Kalman+LSTM improve over a plain LSTM and reach the level of the attention-based ALSTM on most datasets, whereas the simpler moving-average and wavelet variants are more uneven. This confirms that denoising the raw prices yields a measurable but limited benefit, which Craft’s centroid tokenization consistently surpasses. On CIKM24, the MCC of Craft (0.019) ties that of ALSTM, while Craft leads ALSTM by 3.2%p in ACC (0.531 vs. 0.499). We assess the statistical significance of these prediction differences later in this section.
We also investigate the impact of the number k of centroids on model performance. Fig 8 reports the validation-ACC sensitivity to k, selected independently for each market regime on that regime’s validation split. Since the criterion is downstream validation accuracy, k is task-supervised rather than purely unsupervised. The per-dataset optima are spread across the grid, but the cross-dataset mean peaks 1) at k = 200 for the COVID-19 crisis and recent volatile regimes, and 2) at k = 150 for the AI rally regime.
Within each regime, colored curves are the per-dataset validation accuracy and the black curve is the cross-dataset mean. The star marks the k selected on each regime’s validation split: k = 200 for the COVID-19 crisis and recent volatile regimes, and k = 150 for the AI rally regime. Each curve is averaged over five random seeds. The per-regime validation windows are given in Table 2.
Lastly, Fig 9 highlights distinct token-label distributions per market. The Chinese market (CSI300) has a small number of dominant tokens, covering a significant portion of the observations. We observe that Craft naturally adapts to such highly skewed token distributions, leading to the observed regional differences in performance.
Different markets exhibit varying token-label distributions, contributing to observed performance differences.
The ACC values in Tables 5–7 lie in the 0.50–0.56 range, and the MCC values remain close to zero. These magnitudes are typical of daily stock-direction prediction, where small accuracy gains over 0.50 baseline translate into materially different investment outcomes.
Investment simulation (Q2)
To translate our predictive performance into real-world profitability, we simulate investment strategies across the three regimes under realistic trading frictions. We charge a one-way transaction cost of 6 bps on each sell, and to protect profitability under this cost we gate each position by 60 % probability threshold that limits turnover. At the close of day t, a model takes a long position in stock i for day t + 1 only when its predicted exceeds 0.60, and analogously a short position when
exceeds 0.60. Eligible longs and shorts are equal-weighted within each leg. Each position is entered at the closing price of day t and unwound at the closing price of day t + 1, earning the close-to-close return. For CSI300, where unconstrained shorting of the full A-share universe is not realistic, we disable the short leg (long-only). Because the resulting long-short portfolio is cash-neutral, the ASR is computed against a zero risk-free rate. Entering at the same day-t close from which the signal is computed is an idealized fill that ignores intraday execution latency, but this protocol is applied identically to every competing method, so the relative comparison among methods remains fair. As a robustness check, we recompute the annualized Sharpe ratio of every learned method under a stricter two-sided cost that also charges 6 bps on buys, over the recent volatile period. Craft retains the highest value on four of the six datasets (S3 Table).
Fig 10 reports the resulting portfolio-value trajectories for all six datasets. Flat segments correspond to periods during which a model’s predicted probabilities do not exceed the 0.60 threshold, so no positions are opened. Craft attains the highest cumulative return on five of the six datasets. For instance, on CIKM24 Craft delivers +98% cumulative return against the Market’s +85%, and on S&P500 + 33% return against the Market’s +26%. NIKKEI225 is the sole exception, where FEDformer’s series decomposition narrowly leads.
Positions follow the 60 % probability-threshold gated long-short strategy described in the text. Craft (proposed) is shown in black. Flat segments indicate windows during which the model’s top-class probability does not exceed the threshold, so no positions are opened. The short leg is disabled for CSI300 to respect A-share short-selling constraints.
1) COVID-19 crisis period.
Table 8 reports the performance during the 2020 market crash. While MASTER achieves a higher ASR exclusively in the Chinese market (1.551 vs. Craft’s 1.342), Craft consistently demonstrates vastly superior stability by maintaining significantly lower RMDD across all evaluated markets. For instance, on the S&P500 dataset, Craft reduces the maximum drawdown to 18.2%, compared to MASTER’s 35.1%. This superior risk management is inherently driven by Craft’s stock-axis attention network, which effectively captures flight-to-safety dynamics and executes conservative strategies during extreme turbulence.
2) AI rally period.
During the recent bullish trend (Table 9), Craft consistently achieves the highest ASR and the lowest RMDD across all datasets. This confirms that our model successfully capitalizes on strong sectoral momentum.
3) Recent volatile period.
Table 10 displays the results for the full test period. Craft achieves the highest ASR on five of the six datasets, with NI225 the sole exception, and outperforms the best competitor by up to 1.0. Relative to the best alternative, Craft reduces RMDD by up to 7.0% points, ensuring robust returns even in volatile markets.
Table 11 reports, per dataset and regime, the Diebold–Mariano (DM) test on daily loss differentials, the Pesaran–Timmermann (PT) test of Craft’s directional skill, and the Jobson–Korkie test with Memmel correction on the Sharpe-ratio difference, each comparing Craft against the strongest baseline. The Diebold–Mariano test uses the 0–1 directional loss aggregated across listed stocks into one value per trading day, with a Newey–West HAC variance at truncation lag . The Pesaran–Timmermann test pools all (stock, day) pairs. The reported p-values are per dataset–regime cell, without adjustment for multiplicity across cells. The DM and PT tests are each significant on 10 of the 14 dataset–regime pairs. The four DM exceptions are the two splits where Craft trails the strongest baseline (CSI300 in the COVID-19 crisis and NI225 in the recent volatile period) and the two where Craft leads by only 0.3%p in accuracy (AAAI24 and EURO50 in the recent volatile period). The JK test reaches significance mainly on the recent-volatile datasets (four of six). The COVID-19 and AI-rally windows are shorter (about three and six months), so their smaller samples rarely yield significant Sharpe-ratio differences.
Stock embedding quality (Q3)
Craft leverages the dynamic global context of each stock by applying SVD to the correlation matrix at every time step. Specifically, at time t, Craft constructs a correlation matrix from the past 60 days of returns for all individual stocks, and then computes each stock’s embedding via SVD. These embeddings incorporate temporal information, since the pairwise correlation among stocks varies over time.
Table 12 shows the five closest neighbours from the dynamic embeddings for TSLA (Tesla) and BAC (Bank of America) at two snapshots (2021 and 2024). We specifically select these two stocks to demonstrate the model’s capacity to capture contrasting temporal dynamics: TSLA represents a case where market correlations shift dramatically over time, while BAC illustrates a highly stable sectoral relationship. In 2021 TSLA is nearest to Big–Tech stocks such as AAPL and AVGO. By late 2024 its closest neighbours shift to AI–focused firms like NVDA and MSFT. This shift illustrates how the embeddings track the market’s AI boom. BAC stays close to JPM and WFC at both times, confirming a stable financial–sector relationship that our method captures over time.
Craft effectively captures the context of stocks that share strong correlations or exhibit similar characteristics, since it forms each stock’s global context by taking a weighted average of embeddings based on pairwise distances. This mechanism enables Craft to reflect meaningful interdependencies among stocks and provides a robust global perspective for subsequent predictions.
Ablation study (Q4)
Fig 11 compares Craft with its ablated variants.
The dashed line indicates the best competitor on each dataset. Removing the centroid-based clustering (-CC) causes the largest accuracy drop in all four markets.
- CRAFT-CC: Craft without centroid-based clustering (CC) module.
- CRAFT-GC: Craft without global context (GC) feature.
- CRAFT-SA: Craft without stock-axis attention (SA) module.
We observe that each omitted module leads to a drop in prediction accuracy, and the full Craft achieves the best performance. This indicates that the centroid–based clustering reduces inherent randomness, the global context leverages correlated stock patterns, and the inter-stock glocal context enhances the overall prediction accuracy.
Conclusion
We propose Craft, an accurate method for stock price movement prediction that attentively integrates glocal patterns from multiple stocks through both time-axis and stock-axis contexts, without requiring prior knowledge. Craft operates in three key phases: 1) converting stock prices into centroid-based tokens to mitigate inherent randomness, 2) generating a multi-level context that combines local token sequences with a dynamic global context via correlation-based embeddings, and 3) unifying time-axis and stock-axis attention in an end-to-end framework, maximizing predictive performance through coordinated learning of local and inter-stock patterns. Experiments on datasets from major world markets show that Craft achieves the highest accuracy in twelve of the fourteen test settings and ranks the second in the other two, improving accuracy by up to 4.1%p and the Matthews correlation coefficient by up to 7.8 points over the best competitors. In addition, Craft improves the annualized Sharpe ratio by up to 1.0 and reduces the maximum drawdown by up to 7.0%p over the best competitor. The learned global context embeddings also provide new perspectives on how stocks interrelate in both local and global domains.
Supporting information
S1 Table. Measured final-epoch training-loss magnitudes of Craft.
https://doi.org/10.1371/journal.pone.0345252.s001
(PDF)
S2 Table. Per-stock up-day rate distribution over the recent-volatile test window.
https://doi.org/10.1371/journal.pone.0345252.s002
(PDF)
S3 Table. Investment performance under a two-sided transaction cost.
https://doi.org/10.1371/journal.pone.0345252.s003
(PDF)
Acknowledgments
The authors used a large language model (Claude, Anthropic) to assist with language editing of the manuscript text. The tool was applied only to wording and phrasing, and the authors verified every edited sentence against the underlying results and sources before accepting it. All ideas, the study design, the proposed method, the analysis, the released code, the figures, the tables, and the conclusions are the authors’ own. The authors take full responsibility for the content of this article.
References
- 1. Fama EF, French KR. Common risk factors in the returns on stocks and bonds. J Finan Econo. 1993;33(1):3–56.
- 2. Peng Y, Albuquerque PHM, Kimura H, Saavedra CAPB. Feature selection and deep neural networks for stock price direction forecasting using technical analysis indicators. Machine Learn Appl. 2021;5:100060.
- 3.
Vargas MR, dos Anjos CEM, Bichara GLG, Evsukoff AG. Deep leaming for stock market prediction using technical indicators and financial news articles. In: 2018 International Joint Conference on Neural Networks (IJCNN), 2018. 1–8. https://doi.org/10.1109/ijcnn.2018.8489208
- 4.
Deng S, Mitsubuchi T, Shioda K, Shimada T, Sakurai A. Combining technical analysis with sentiment analysis for stock price prediction. In: 2011 IEEE Ninth international conference on dependable, autonomic and secure computing. 2011. 800–7. https://doi.org/10.1109/dasc.2011.138
- 5.
Khairi TWA, Zaki RM, Mahmood WA. Stock price prediction using technical, fundamental and news based approach. In: 2019 2nd Scientific Conference of Computer Sciences (SCCS), 2019. 177–81. https://doi.org/10.1109/sccs.2019.8852599
- 6.
Beyaz E, Tekiner F, Zeng X, Keane J. Comparing technical and fundamental indicators in stock price forecasting. In: 2018 IEEE 20th International conference on high performance computing and communications; IEEE 16th International conference on smart city; IEEE 4th International conference on data science and systems (HPCC/SmartCity/DSS), 2018. 1607–13. https://doi.org/10.1109/hpcc/smartcity/dss.2018.00262
- 7.
Turan AC, Kokach V, Kocaçınar B, Şengel Ö, Akbulut FP. Hybrid deep learning framework for stock price prediction incorporating technical and macroeconomic indicators. In: 2024 9th International Conference on Computer Science and Engineering (UBMK), 2024. 1189–94. https://doi.org/10.1109/ubmk63289.2024.10773522
- 8.
Cao Z, Xu J, Dong C, Yu P, Bai T. MATCC: a novel approach for robust stock price prediction incorporating market trends and cross-time correlations. In: Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 2024. 187–96. https://doi.org/10.1145/3627673.3679715
- 9. Li T, Liu Z, Shen Y, Wang X, Chen H, Huang S. MASTER: market-guided stock transformer for stock price forecasting. AAAI. 2024;38(1):162–70.
- 10.
Lin H, Zhou D, Liu W, Bian J. Learning multiple stock trading patterns with temporal routing adaptor and optimal transport. In: Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining, 2021. 1017–26. https://doi.org/10.1145/3447548.3467358
- 11.
Yoo J, Soun Y, Park Y, Kang U. Accurate multivariate stock movement prediction via data-axis transformer with multi-level contexts. In: Proceedings of the 27th ACM SIGKDD conference on knowledge discovery & data mining. 2021. 2037–45. https://doi.org/10.1145/3447548.3467297
- 12. Zhang GP. Time series forecasting using a hybrid ARIMA and neural network model. Neurocomp. 2003;50:159–75.
- 13. Khandelwal I, Adhikari R, Verma G. Time series forecasting using hybrid ARIMA and ANN models based on DWT decomposition. Procedia Comp Sci. 2015;48:173–9.
- 14.
Box GEP, Jenkins GM, Reinsel GC, Ljung GM. Time series analysis: Forecasting and control. Hoboken, NJ: John Wiley & Sons; 2015.
- 15. Rosenberg B, Rudd A. Factor-related and specific returns of common stocks: serial correlation and market inefficiency. J Finance. 1982;37(2):543.
- 16. Daniel K, Moskowitz TJ. Momentum crashes. J Finan Econ. 2016;122(2):221–47.
- 17.
Zhou T, Niu P, Wang X, Sun L, Jin R. One fits all: power general time series analysis by pretrained LM. In: Advances in neural information processing systems 36. 2023. 43322–55. https://doi.org/10.52202/075280-1877
- 18.
Jin M, Wang S, Ma L, Chu Z, Zhang JY, Shi X. Time-llm: Time series forecasting by reprogramming large language models. 2023.
- 19.
Xu W, Liu W, Wang L, Xia Y, Bian J, Yin J. Hist: A graph-based framework for stock trend forecasting via mining concept-oriented shared information. 2021.
- 20.
Alaygut T, Sefer E. Hypergraph neural networks to predict stock movements by exploring higher-order relationships. In: Proceedings of the 6th ACM International Conference on AI in Finance, 2025. 700–8. https://doi.org/10.1145/3768292.3770389
- 21. Chong E, Han C, Park FC. Deep learning networks for stock market analysis and prediction: Methodology, data representations, and case studies. Expert Systems with Applications. 2017;83:187–205.
- 22.
Selvin S, Vinayakumar R, Gopalakrishnan EA, Menon VK, Soman KP. Stock price prediction using LSTM, RNN and CNN-sliding window model. In: 2017 International Conference on Advances in Computing, Communications and Informatics (ICACCI), 2017. 1643–7. https://doi.org/10.1109/icacci.2017.8126078
- 23.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, et al. Attention Is All You Need. arXiv preprint. 2017. https://doi.org/10.48550/arXiv.1706.03762
- 24. Haryono AT, Sarno R, Sungkono KR. Transformer-gated recurrent unit method for predicting stock price based on news sentiments and technical indicators. IEEE Access. 2023;11:77132–46.
- 25. Yilmaz S, Sefer E. Pairs trading with time-series deep learning models. J Finance and Data Sci. 2025;11:100177.
- 26. Wu H. Revisiting attention for multivariate time series forecasting. AAAI. 2025;39(20):21528–35.
- 27.
Radford A, Wu J, Child R, Luan D, Amodei D, Sutskever I. Language models are unsupervised multitask learners. OpenAI Technical Report. 2019.
- 28. Brown T, Mann B, Ryder N, Subbiah M, Kaplan J, Dhariwal P, et al. Language models are few-shot learners. Adv Neural Inform Process Syst. 2020;33:1877–901.
- 29.
van den Oord A, Vinyals O, Kavukcuoglu K. Neural discrete representation learning. Adv Neural Inform Process Systems (NeurIPS). 2017.
- 30. Thorndike RL. Who belongs in the family?. Psychometrika. 1953;18(4):267–76.
- 31. Rousseeuw PJ. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis. J Comp Appl Math. 1987;20:53–65.
- 32. Kriegel HP, Kröger P, Zimek A. Clustering high-dimensional data: A survey on subspace clustering, pattern-based clustering, and correlation clustering. ACM Trans Knowl Discov Data. 2009;3(1):1–58.
- 33. Schubert E. Stop using the elbow criterion for k-means and how to choose the number of clusters instead. SIGKDD Explor Newsl. 2023;25(1):36–42.
- 34.
Nelson DMQ, Pereira ACM, de Oliveira RA. Stock market’s price movement prediction with LSTM neural networks. In: 2017 International Joint Conference on Neural Networks (IJCNN), 2017. 1419–26. https://doi.org/10.1109/ijcnn.2017.7966019
- 35.
Qin Y, Song D, Chen H, Cheng W, Jiang G, Cottrell G. A dual-stage attention-based recurrent neural network for time series prediction. 2017. https://doi.org/10.48550/arXiv.1704.02971
- 36.
Nie Y, Nguyen NH, Sinthong P, Kalagnanam J. A time series is worth 64 words: Long-term forecasting with transformers. 2022. https://doi.org/10.48550/arXiv.2211.14730
- 37. Zhou H, Zhang S, Peng J, Zhang S, Li J, Xiong H, et al. Informer: Beyond efficient transformer for long sequence time-series forecasting. AAAI. 2021;35(12):11106–15.
- 38.
Wu H, Xu J, Wang J, Long M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. In: Advances in neural information processing systems. 2021.
- 39.
Zhou T, Ma Z, Wen Q, Wang X, Sun L, Jin R. FEDformer: frequency enhanced decomposed transformer for long-term series forecasting. In: Proceedings of the 39th International Conference on Machine Learning, 2022.
- 40. Brock W, Lakonishok J, LeBaron B. Simple Technical Trading Rules and the Stochastic Properties of Stock Returns. The Journal of Finance. 1992;47(5):1731–64.
- 41.
Ariyo AA, Adewumi AO, Ayo CK. Stock price prediction using the ARIMA model. In: 2014 UKSim-AMSS 16th International conference on computer modelling and simulation. 2014. 106–12. https://doi.org/10.1109/uksim.2014.67
- 42. Bao W, Yue J, Rao Y. A deep learning framework for financial time series using stacked autoencoders and long-short term memory. PLoS One. 2017;12(7):e0180944. pmid:28708865
- 43. Donoho DL, Johnstone IM. Ideal spatial adaptation by wavelet shrinkage. Biometrika. 1994;81(3):425–55.
- 44.
Harvey AC. Forecasting, structural time series models and the kalman filter. Cambridge, UK: Cambridge University Press; 1989.
- 45. Malkiel BG. The efficient market hypothesis and its critics. J Economic Persp. 2003;17(1):59–82.