Figures
Abstract
Language use, vocal behavior, and facial activity provide complementary indicators of affective state, but multimodal models can be affected by noisy temporal observations and insufficient global guidance before cross-modal interaction. We developed a framework combining bidirectional gated recurrent unit encoders, attention pooling, dynamic hypernode injection, and graph attention fusion. Textual, acoustic, and visual sequences were mapped into a shared latent space and compressed into modality-level representations. A sample-dependent hypernode and a learnable static prior were then injected through gated residual connections before graph propagation. The model was evaluated on CMU-MOSI and CMU-MOSEI using five random seeds and validation-MAE checkpoint selection. On CMU-MOSI, the model obtained an MAE of , a correlation of
, and an Acc-7 of
. On CMU-MOSEI, the corresponding results were
,
, and
. A five-run comparison of pre-GAT, post-GAT, and no injection showed that pre-GAT injection provided the strongest overall regression-oriented balance, although post-GAT was slightly higher on MOSEI Acc-7. Gate, graph-attention, representation, and balanced case diagnostics showed that the injected prior was sample dependent and modality selective. These findings support pre-propagation hypernode injection under complete, word-aligned tri-modal inputs, without establishing universal superiority or robustness to missing and corrupted modalities.
Citation: Cao W, Xu J, Li Y, Yan H, Hu Z, Chen W (2026) Dynamic hypernode injection and graph attention fusion for multimodal sentiment analysis of behavioral signals. PLoS One 21(9): e0359425. https://doi.org/10.1371/journal.pone.0359425
Editor: Shuai Liu, Hunan Normal University, CHINA
Received: July 22, 2026; Accepted: September 14, 2026; Published: September 28, 2026
Copyright: © 2026 Cao et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The numerical values reported in the results tables are provided in S1 Data. The training and analysis code supporting the findings of this study is publicly available from Zenodo (version DOI: 10.5281/zenodo.22218345). The CMU-MOSI and CMU-MOSEI benchmark datasets are publicly available through the official CMU MultimodalSDK (https://github.com/CMU-MultiComp-Lab/CMU-MultimodalSDK). The datasets are not redistributed in the code archive.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Observable language use, vocal patterns, and facial behavior provide complementary information about affective state and interpersonal expression. Computational models that integrate these signals are increasingly used in affective computing and behavior-aware human-computer interaction. The affective computing framework [1] laid the theoretical foundation for this field. Subsequent work showed that jointly modeling facial expressions, body movements, and speech information can improve emotion recognition [2]. More recently, a comprehensive review [3] surveyed multimodal emotion recognition across theoretical foundations, datasets, modality-specific feature extraction, cross-modal alignment, and fusion strategies, and summarized the strengths, limitations, challenges, and future directions of the field.
CMU-MOSI [4] and CMU-MOSEI [5] are standard video-based benchmarks for multimodal sentiment analysis. Here they are used to examine a modeling question rather than as the motivation by themselves: how sample-level global information can guide structured cross-modal propagation while modality-specific temporal evidence is retained.
As multimodal fusion methods have continued to evolve, tensor fusion has been used to explicitly model higher-order interactions among text, acoustic, and visual modalities [6], multi-view memory mechanisms have been applied to capture cross-modal dynamics [7], and low-rank decomposition has been employed to reduce the computational cost of fusion [8]. Transformer-based and representation-learning approaches have further improved cross-modal interaction, including directional cross-modal attention for unaligned sequences [9] and modality-invariant and modality-specific representation learning [10]. Cross-modal enhancement networks have been designed to enrich textual representations with non-textual information [11], and text-enhanced transformer fusion has been developed to emphasize the guiding role of language in sentiment prediction [12]. Recent studies have also examined sequential fusion of text-close and text-far representations [13], multimodal adapter mixtures [14], invariant multimodal representations [15], and directed pairwise cross-modal GRU attention [16]. Complementarity- and importance-aware fusion has also been proposed for multimodal emotion recognition [17], while fuzzy cognition-based dynamic fusion has been used to integrate modality-specific sentiment scores for multimodal sentiment analysis [18].
Beyond conventional fusion, structured modeling has become an important direction for capturing complex multimodal dependencies. DialogueGCN [19] and COGMEN [20] obtain contextual information from utterance-level conversational graphs, speaker relations, and message passing. ALMT [21] instead constructs a language-guided adaptive hyper-modality during staged cross-modal interaction. These mechanisms are effective for their respective settings, but they do not explicitly construct a sample-specific tri-modal prior and inject it into every modality node before the first graph-attention layer. The proposed pipeline makes this ordering explicit while retaining a compact three-node segment graph. Table 1 summarizes these structural differences.
Despite these advances, three modeling gaps remain relevant. First, heterogeneous modality features may enter fusion without a sufficiently comparable latent representation. Second, sample-level global information is not always used to condition all modality nodes before structured propagation begins. Third, redundant temporal observations can enter graph reasoning unless sequence evidence is compressed before cross-modal interaction.
To address these issues, we propose a multimodal sentiment analysis framework that integrates BiGRU-based intra-modal temporal modeling, dynamic hypernode injection, and graph attention fusion. Specifically, the proposed method first maps textual, acoustic, and visual features into a shared latent space through a text encoder and modality-specific linear projections. It then employs BiGRU and attention pooling to obtain modality-level node representations that preserve temporal dynamics while suppressing irrelevant time-step noise. On this basis, a dynamic hypernode is constructed from the global tri-modal context and combined with a learnable static hyper-prior. The resulting global semantic prior is adaptively injected into each modality node through a gated residual mechanism before cross-modal propagation. Finally, the enhanced modality nodes are organized as a fully connected tri-modal graph, and graph attention fusion is used to model structured inter-modal dependencies for continuous sentiment intensity prediction.
By explicitly introducing global semantic guidance before graph propagation and combining it with modality-aware temporal compression, the proposed framework provides a structurally distinct fusion pipeline. Experiments on CMU-MOSI and CMU-MOSEI evaluate this design across regression and classification metrics, five random seeds, three injection positions, component ablations, and direct mechanism diagnostics. The contribution is architectural and empirical within complete, aligned tri-modal inputs; it is not a claim of universal robustness.
Materials and methods
Study design and data sources
This computational benchmark study evaluates a multimodal sentiment model using two public datasets, CMU-MOSI and CMU-MOSEI. CMU-MOSI [4] contains 2,199 opinion-level video segments, and CMU-MOSEI [5] contains 22,856 segments with a broader speaker distribution and more diverse scenarios. Both benchmarks use continuous sentiment scores in the range ; these scores also support the seven ordered classes used for Acc-7. Table 2 summarizes the official training, validation, and test partitions used in this study.
Both datasets provide textual, acoustic, and visual modalities and are used here as established benchmarks of affect-related behavioral signal modeling.
Overall framework
For continuous sentiment-intensity prediction, let the text, acoustic, and visual inputs be ,
, and
, respectively. Here,
is the sequence length of modality m,
is its input dimension, and
is the sentiment label. The objective is to learn the following end-to-end mapping:
where denotes the predicted sentiment intensity.
To jointly model the complementary information conveyed by temporal dynamics, global context, and cross-modal dependencies, we propose a multimodal sentiment analysis framework based on dynamic hypernode injection and graph attention fusion. The overall pipeline consists of five stages. First, a pretrained language model and modality-specific linear projections are used to map the three modalities into a shared latent space. Second, bidirectional gated recurrent units [22] are employed to model intra-modal temporal dependencies, and attention pooling is used to obtain modality-level node representations. Third, a dynamic hypernode is constructed from the global tri-modal context and injected into each modality node through a gated residual mechanism. Fourth, multi-layer graph attention networks [23] are applied to a cross-modal graph composed of text, acoustic, and visual nodes in order to perform structured relation propagation. Finally, the graph-fused representation is fed into a regression head to predict sentiment intensity.
As illustrated in Fig 1, the proposed framework follows a hierarchical modeling process from single-modal representation learning to global interaction and downstream prediction. The three modalities are first aligned in a unified latent space and then compressed into modality-level nodes through temporal encoding and attention pooling. These nodes are subsequently enhanced by dynamic hypernode injection. On this basis, cross-modal dependencies are modeled through graph attention propagation over a fully connected three-node graph, and the final sentiment intensity score is obtained through node-level readout and regression. Table 3 summarizes the principal symbols used in the model formulation.
The figure was created by the authors. The face sequence in the visual-input block is a schematic illustration and does not depict a study participant or any identifiable individual.
Unified modal representation learning
For the text modality, we use a pretrained BERT encoder [24] to extract contextualized semantic representations. Specifically, the textual input consists of input_ids, attention_mask, and token_type_ids, and the final hidden states produced by BERT are taken as the text sequence representation. In implementation, the BERT parameters are fine-tuned end to end to improve task-specific adaptability.
Let the text representation after BERT encoding be . For m = t,
; for m = a,
; and for m = v,
. To reduce modality heterogeneity, we assign each modality an independent projection and map all modalities into a shared latent space of dimension d:
where m = t, for the text modality. After projection, layer normalization and dropout are further applied to improve training stability and generalization:
Through this process, the text, acoustic, and visual modalities are brought into the same representational space, providing comparable inputs for subsequent intra-modal temporal modeling and cross-modal graph fusion.
Intra-modal sequence modeling
To capture dynamic dependencies within each modality, we encode the unified text, acoustic, and visual sequences with bidirectional gated recurrent units. For modality m, given at time step i, the forward and backward hidden states are computed as:
By concatenating the bidirectional outputs, we obtain the temporal feature representation:
Since the bidirectional encoding expands the hidden dimension to 2d, we compress it back to d dimensions through a linear projection, followed by GELU, layer normalization, and dropout:
To compress each aligned sequence into a modality-level node representation, we use linear attention pooling. For the i-th time step of modality m, the attention score is defined as:
The normalized weights are:
The final modality-level node representation is given by:
This BiGRU-plus-attention-pooling design combines sequential encoding with learned temporal weighting. Its contribution is evaluated by the repeated-seed ablation results below.
Dynamic hypernode construction and gated injection
Relying solely on pairwise interactions between modalities often makes it difficult to introduce a global semantic prior before cross-modal propagation. To address this issue, we design a dynamic hypernode module that aggregates the global context of the three modality nodes for the current sample and injects it into each modality representation as prior information.
First, the text, acoustic, and visual modality nodes are averaged to obtain a sample-level context vector:
Next, a two-layer multilayer perceptron performs a nonlinear transformation of the global context and generates the hypernode representation:
Unlike a purely input-driven dynamic representation, we further introduce a learnable global prior vector as a bias term. In the implementation, this prior is derived from the first token of the learnable hypernode parameters. The final dynamic hypernode representation is therefore written as:
To avoid imposing an identical degree of global injection on all modalities, we assign each modality an individual scalar gating coefficient to control the influence of the hypernode on that modality. For modality m, the gating value is defined as
where and
are learnable gate parameters and
denotes the sigmoid function. Finally, the global prior is injected through a residual formulation:
Through this design, the model can adaptively inject the global semantic prior jointly formed by the three modalities into each modality representation while preserving modality-specific characteristics, thereby alleviating multimodal information fragmentation and insufficient global context modeling.
Graph-based attention for cross-modal fusion
After dynamic hypernode injection, the hypernode is not directly introduced as an additional graph node. Instead, as illustrated in Fig 2, it is generated once from the tri-modal context together with the static hyper-prior and is used to enhance the three modality nodes before graph propagation. Let denote the enhanced text, acoustic, and visual nodes, respectively. We then construct a fully connected tri-modal graph G = (V, E), where
. This design keeps the graph structure compact while allowing the injected global prior to participate in subsequent cross-modal propagation through the enhanced modality nodes. As a result, the graph explicitly represents the latent dependencies among the three modalities and provides a structured basis for cross-modal information fusion.
The sample-level context is transformed into a hypernode, combined with a static prior, and added to each modality through a learned gate. The enhanced text, acoustic, and visual nodes form a fully connected three-node graph; the hypernode itself is not an additional graph node.
For convenience, the initial node states are defined as:
During graph propagation, we employ stacked graph attention layers to model inter-modal interactions. In the l-th layer, the unnormalized attention coefficient between node i and its neighbor is computed as
where W(l) is a learnable linear transformation matrix, a(l) is the attention vector, and denotes vector concatenation. The corresponding normalized attention weight is:
The node update process can be written as:
where denotes the non-linear activation function. Through this layer-wise aggregation mechanism, each modality node can adaptively absorb information from the other modalities according to sample-specific relational strength, thereby enabling structured cross-modal interaction rather than simple feature concatenation.
To model deeper inter-modal dependency propagation, the graph fusion block stacks graph attention layers. Each non-final layer is followed by ELU and Dropout to enhance non-linear expressiveness and reduce overfitting, whereas LayerNorm is applied after the final layer to stabilize node representations after deep graph propagation. After the final layer, we obtain the node set
. We then perform node-level attention pooling for graph readout. Specifically, an importance score is assigned to each modality node as:
and the corresponding readout weight is computed by
The final fused representation is given by
This graph attention fusion mechanism enables the model to propagate and aggregate cross-modal information according to the contribution of each modality in a given sample, thereby producing a fused representation that is both globally coherent and cross-modally complementary.
Sentiment regression prediction and model training
Following graph readout, the fused graph-level representation is first regularized by layer normalization and dropout:
A two-layer multi-layer perceptron is then used as the regression head to predict continuous sentiment intensity. The prediction process is formulated as:
where are the two regression-layer parameters, o is the hidden prediction vector, and
is the scalar sentiment score. This design is consistent with the regression branch shown on the right-hand side of Fig 2, where the graph-level fused representation is transformed into a scalar sentiment-intensity prediction.
For the continuous sentiment intensity prediction tasks on CMU-MOSI and CMU-MOSEI, the model is optimized using mean squared error:
where N is the number of samples in a mini-batch and is the ground-truth sentiment label of the n-th sample.
During training, the entire framework is optimized end to end with AdamW. Unified modal representation learning, intra-modal temporal modeling, dynamic hypernode injection, graph-based cross-modal fusion, and sentiment regression are jointly learned within a single objective.
Implementation details
The model is implemented in PyTorch [25]. The archived environment uses Python 3.10.0, PyTorch 2.4.1 + cu118, CUDA 11.8, cuDNN 90100, and one NVIDIA RTX 4080 SUPER GPU. The ALMT training program loads word-aligned serialized feature artifacts named aligned_50.pkl. Text is encoded with bert-base-uncased and has an output dimension of 768; the aligned acoustic and visual features have 25 and 171 dimensions, respectively. Fixed sequence lengths are 45 for MOSI and 75 for MOSEI. The exact CMU-MultimodalSDK release/commit and original acoustic/visual extractor versions are not preserved in the supplied records and are therefore reported as unavailable rather than inferred. The training and analysis code used for this study is publicly archived at Zenodo (https://doi.org/10.5281/zenodo.22218345) [26].
At runtime, the text tensor has three channels for input IDs, attention mask, and token-type IDs. Acoustic negative-infinity values are replaced with zero, no additional feature normalization is applied, and all tensors are converted to float32.
The model is trained with mean squared error and AdamW [27]. The fixed settings are batch size 64, learning rate 10–4, weight decay 10–4, and a maximum of 200 epochs. Early stopping uses a patience of 40 epochs, and the checkpoint with the lowest validation MAE is selected. Seven graph-attention layers are used for MOSI and four for MOSEI. Five random seeds are run for the main model, comparison methods, component ablations, and each injection-position condition. For these experiments, the reported rows give means and sample standard deviations across five runs, and every metric within a run comes from the same validation-selected checkpoint. Table 4 consolidates the principal feature, training, and runtime settings. Except for the fixed sequence length and the graph depth selected from the development-stage depth sweep, the listed settings were kept the same for MOSI and MOSEI.
Evaluation protocol and metrics
We evaluate continuous and categorical sentiment prediction using mean absolute error (MAE), Pearson correlation (Corr), five-class accuracy (Acc-5), seven-class accuracy (Acc-7), and binary accuracy under Has0 and Non0 conventions. Has0 retains zero-label samples and separates negative from non-negative instances, whereas Non0 excludes zero-label samples and separates negative from positive instances. Acc-7 is calculated by clipping scores to and rounding to the nearest ordered class. Test data are used only for final evaluation; checkpoint selection uses validation MAE. All displayed metrics for a run are taken from the same selected checkpoint.
Ethics statement
This study involved secondary analysis of publicly available benchmark datasets and did not involve recruitment of participants or collection of new human-subject data. No additional ethics approval or informed consent was required for this study. The person shown schematically in Fig 1 is an author-created illustration and does not depict a study participant or any identifiable individual.
Results
Comparative experimental results and analysis
We compare the proposed framework with TFN [6], LMF [8], MulT [9], MISA [10], CENet [11], LF-DNN as included in the CH-SIMS benchmark framework [28], and ALMT [21]. The supplied experiment package reports five-run means and sample standard deviations under the same metric implementation. Because paired sample-level predictions were not available for every baseline, numerical rank differences are treated as descriptive rather than as formal superiority tests. Tables 5 and 6 report the corresponding five-run results for CMU-MOSI and CMU-MOSEI, respectively.
The quantitative comparisons show where the numerical gains are concentrated and where the margins remain small. On CMU-MOSI, the proposed model has a mean MAE of 0.7290, which is 0.0364 lower than the next-lowest value (CENet, 0.7654), and a Corr of 0.7880, which is 0.0099 higher than MISA (0.7781). Its Acc-7 is 0.4561 versus 0.4490 for CENet. The margins are narrower for binary accuracy: Has0 Acc-2 is 0.8233 versus 0.8200 for ALMT, and Non0 Acc-2 is 0.8373 versus 0.8354 for MISA. On CMU-MOSEI, the corresponding differences from the strongest baseline mean are also modest for several metrics: MAE is 0.5541 versus 0.5575 for MISA, Corr is 0.7551 versus 0.7515 for MISA, Acc-5 is 0.5450 versus 0.5421 for ALMT, and Acc-7 is 0.5269 versus 0.5211 for MulT. Thus, the proposed method is numerically favorable across the displayed mean metrics, but several margins are small. Because matched sample-level predictions are unavailable for every baseline, these differences are treated as descriptive comparisons rather than evidence that the method statistically outperforms existing approaches.
The earlier single-run comparison suggested a different MOSI–MOSEI pattern. Under the uniform five-run results supplied for this revision, that pattern did not persist. We therefore report the revised statistics and avoid a post hoc causal explanation based only on dataset scale or heterogeneity.
Analysis of the training process
Fig 3 summarizes the five-seed training trajectories. The main decrease in training loss and validation MAE occurs early on both datasets. MOSI validation Corr and Acc-5 then approach a broad plateau, with occasional seed-specific spikes. MOSEI also converges early, but its validation Corr and Acc-5 show a mild later decline while training loss continues to decrease. This train–validation separation is consistent with later-stage overfitting, especially on MOSEI, and supports validation-based early stopping.
Panels show training mean-squared error, validation MAE, validation correlation, and validation Acc-5. Thin colored lines represent seeds, the black line is their mean, the gray band is one sample standard deviation, and circles mark the validation-MAE-selected checkpoint for each seed.
Ablation studies
To examine the contribution of individual components, we removed attention pooling, BiGRU encoding, and the static hyper-prior. Table 7 reports five-run means and sample standard deviations. Removing attention pooling produced the largest MOSI Corr decrease (0.7880 to 0.6735) and a substantial MOSEI Acc-7 decrease (0.5269 to 0.4130), consistent with its temporal-selection role. Removing BiGRU had a smaller but consistent effect on the regression-oriented metrics. Removing the static prior modestly affected MOSI MAE and Corr but more strongly affected MOSEI MAE and Acc-7. These patterns are consistent with complementary roles for temporal selection, sequential encoding, and prior conditioning; they are component-level associations rather than causal localization.
Injection-position validation
Table 8 directly compares injection before GAT propagation, injection after the GAT stack but before readout, and no injection. On MOSI, pre-GAT injection had higher Corr and Acc-7 than both controls, while its MAE was nearly identical to post-GAT and lower than no injection. On MOSEI, pre-GAT had the lowest MAE, highest Corr, and strongest binary accuracies, while post-GAT had a slightly higher Acc-7 (0.5275 versus 0.5269). Thus, pre-propagation injection provided the strongest overall regression-oriented balance but was not uniformly best on every metric. The supplied pre-GAT and control runs used different seed identities, so this comparison is not treated as a paired significance test.
Mechanism diagnostics and representative cases
The gate analysis pooled the complete test set across five pre-GAT seeds. Mean text/acoustic/visual gates were 0.369/0.262/0.216 on MOSI and 0.314/0.246/0.220 on MOSEI. This ordering indicates modality-selective rather than identical injection. In the first GAT layer, pooled attention mass was higher for text-related units (0.515 on MOSI and 0.489 on MOSEI) than for acoustic- or visual-related units. From the second layer onward, the pooled means approached one third. These values describe aggregate tendencies and do not imply uniform attention for every sample or head.
The dynamic hypernode also varied across samples. Its first two principal components explained 72.65% of variance on MOSI and 65.25% on MOSEI. For the final fused representation, the corresponding totals were 99.58% and 97.18%. These variance profiles establish that the hypernode is not a constant vector, but they do not by themselves establish sentiment-class separation or causal mediation. Fig 4 summarizes these results. A separate temporal-attention audit found that the current pooling implementation did not explicitly mask padded text positions; temporal-attention mass is therefore not used here as affirmative mechanism evidence.
(a) Mean modality-specific pre-GAT gate values; error bars show sample standard deviations across the pooled test-set values. (b) Mean graph-attention mass pooled over samples, seeds, heads, and associated directed units; the dotted line marks uniform mass (1/3), and the MOSEI trajectories stop at layer 4 because the final MOSEI model uses four GAT layers. (c) Mean Pearson correlation across five runs for pre-GAT, post-GAT, and no injection; the corresponding sample standard deviations are reported in Table 8. (d) Variance explained by the first two principal components of the dynamic hypernode and fused representation. Aggregate attention patterns and PCA variance do not imply identical sample-level behavior or causal mediation.
Balanced case selection included examples where pre-GAT helped, where it hurt, and where all variants failed. Table 9 reports six such cases from the common seed-7 diagnostic run. The mixed pattern shows that injection position changes individual predictions but is not beneficial for every segment.
Effect of the number of graph fusion layers
The original development-stage depth sweep compared one to ten GAT layers. Because the model is trained for continuous sentiment regression and validation MAE is used for checkpoint selection, the final depth was chosen by prioritizing the joint MAE–Corr profile while also checking Acc-7 rather than by maximizing binary F1 alone. On CMU-MOSI, seven layers produced the lowest MAE (0.684), the highest Corr (0.813), and the highest Acc-7 (49.7%) among the tested depths, although five layers gave higher binary F1. On CMU-MOSEI, four layers produced the lowest MAE (0.539), the highest Corr (0.768), and the highest Acc-7 (53.8%), while other depths produced higher binary F1. These observations motivated the seven-layer MOSI and four-layer MOSEI configurations used in the five-seed main and injection-position experiments. The depth sweep itself was a development-stage sensitivity analysis without repeated-seed evaluation, so these choices should be interpreted as dataset-specific settings for the present experiments, not as evidence of a universally optimal GAT depth. Table 10 summarizes the development-stage graph-depth results for both datasets.
Discussion
The experimental results show that the position of hypernode injection has a measurable effect on multimodal sentiment prediction. Across the five validation-selected runs, injecting the hypernode before graph propagation produced the most consistent overall performance on both CMU-MOSI and CMU-MOSEI, particularly for regression-oriented metrics. Post-GAT injection achieved a slightly higher Acc-7 score on MOSEI, indicating that the advantage of pre-GAT injection is not uniform across all evaluation criteria. A more appropriate interpretation is therefore that early global conditioning provides a favorable balance under the present experimental setting. By introducing the shared multimodal context before message passing, the graph layers can propagate representations that have already been adjusted according to sample-level information, rather than incorporating this information only after graph interaction has taken place.
The diagnostic analyses provide further insight into how the model uses this global conditioning. Text consistently received the largest average injection gate on both datasets, suggesting that the model relies more strongly on textual information when adjusting modality-specific representations. At the same time, the dynamic component of the hypernode varied across samples, which indicates that the injected context is not reduced to a fixed global bias. The first GAT layer also assigned relatively greater attention to text-related units, whereas attention distributions became more balanced in deeper layers. Taken together, these patterns suggest that the model first emphasizes the modality carrying the strongest sentiment signal and then progressively integrates information across modalities during graph propagation. This interpretation is consistent with the known importance of language in multimodal sentiment analysis, while still allowing visual and acoustic cues to modify the final representation.
These diagnostic results should nevertheless be interpreted with some caution. Average gate values and graph-attention weights describe how information is distributed inside the trained model, but they do not by themselves demonstrate causal mediation between particular modalities and the final prediction. Similarly, variation observed in the dynamic hypernode through PCA reflects changes in its representation across samples rather than explicit class separation. The temporal-attention analysis revealed an additional limitation: a noticeable proportion of attention was assigned to padded positions because padding was not explicitly masked during temporal pooling. For this reason, temporal-attention maps are not treated as evidence for the interpretation of model behavior in the present study. A padding-aware pooling mechanism would provide a cleaner basis for such analysis in future implementations.
The current results also define the practical scope of the proposed framework. Accordingly, the conclusions should be interpreted as specific to CMU-MOSI and CMU-MOSEI, the complete word-aligned tri-modal input setting, and the reported training and evaluation protocol. The model is designed for settings in which text, visual, and acoustic streams are simultaneously available and can be aligned reliably at the word or segment level. Under these conditions, the compact graph provides an efficient way to combine modality-specific information with a sample-dependent global context. The present experiments, however, do not examine incomplete modalities, severe corruption, additive noise, multilingual inputs, unaligned streams, cross-domain transfer, or multi-turn conversational structure. These scenarios introduce different sources of uncertainty and may require more explicit mechanisms for modality reliability or temporal context. The current hypernode also uses simple mean-based aggregation of the three pooled modality representations. Although this choice keeps the architecture compact, attention-weighted or uncertainty-aware aggregation may provide a more flexible alternative when modality quality varies substantially across samples.
Several experimental limitations should also be considered when interpreting the reported gains. The available experimental archive did not preserve the exact CMU-MultimodalSDK release or the original versions of the acoustic and visual feature extractors, which restricts complete reconstruction of the preprocessing environment. In addition, the pre-GAT and control configurations were each evaluated over five runs, but they were not generated from an identical set of random seeds. Their comparison therefore reflects repeated-run performance rather than a strictly paired significance test. The graph-depth analysis was also conducted without repeated-seed evaluation, so the observed degradation at greater depth should not be taken as direct evidence of graph oversmoothing. Finally, computational efficiency was not evaluated systematically, and the current experiments do not determine whether similar performance could be obtained with fewer graph layers, reduced hidden dimensions, or sparse attention mechanisms.
Future work should address these limitations through a more tightly controlled evaluation protocol. Positional ablations should be repeated with identical preregistered seeds, and paired statistical tests should be reported when predictions from matched runs are available. Padding-aware temporal pooling is also necessary before temporal attention can be used reliably for interpretation. Beyond these implementation issues, robustness to modality dropout, corrupted inputs, and domain shift deserves particular attention because real multimodal systems rarely receive equally reliable signals from every modality. Comparing mean-based hypernode construction with attention-weighted and uncertainty-aware alternatives may clarify whether the global representation can adapt more effectively to such variation. Lightweight or sparse graph-attention variants could further reduce computational cost. Extensions to multilingual sentiment analysis, multi-turn conversation, and cross-domain transfer would also help establish whether the same global-conditioning mechanism remains useful beyond aligned English-language segment-level benchmarks. In this context, incomplete-modality approaches such as proxy-driven robust multimodal fusion provide a relevant direction for comparison [29].
Conclusions
We proposed a multimodal sentiment framework that combines BiGRU temporal modeling, attention pooling, dynamic hypernode injection, and graph-attention fusion. The sample-specific hypernode and static prior are injected through modality-specific gates before graph propagation.
Across five validation-selected runs, the model achieved an MAE of , Corr of
, and Acc-7 of
on MOSI. On MOSEI, the corresponding values were
,
, and
. The pre/post/no-injection comparison showed the strongest overall regression-oriented balance for pre-GAT injection, although it did not lead on every metric. Gate, graph-attention, PCA, and balanced case analyses further showed that the injected prior was sample dependent and modality selective.
These findings support pre-propagation hypernode injection within complete, aligned tri-modal segment-level data. They do not establish universal superiority, robustness to missing or corrupted modalities, or generalization to multilingual, cross-domain, or multi-turn settings.
Supporting information
S1 Data. Numerical values reported in the results tables.
https://doi.org/10.1371/journal.pone.0359425.s001
(XLSX)
References
- 1.
Picard RW. Affective computing. Cambridge (MA): MIT Press; 1997.
- 2.
Caridakis G, Castellano G, Kessous L, Raouzaiou A, Malatesta L, Asteriadis S, et al. Multimodal emotion recognition from expressive faces, body gestures and speech. Artificial Intelligence and Innovations 2007: From Theory to Applications. Boston (MA): Springer; 2007. p. 375–88. https://doi.org/10.1007/978-0-387-74161-1_41
- 3. Ramaswamy MPA, Palaniswamy S. Multimodal emotion recognition: A comprehensive review, trends, and challenges. WIREs Data Min Knowl. 2024;14(6):e1563.
- 4. Zadeh A, Zellers R, Pincus E, Morency L-P. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages. IEEE Intell Syst. 2016;31(6):82–8.
- 5.
Bagher Zadeh A, Liang PP, Poria S, Cambria E, Morency LP. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Melbourne, Australia: Association for Computational Linguistics; 2018. p. 2236–46. https://doi.org/10.18653/v1/P18-1208
- 6.
Zadeh A, Chen M, Poria S, Cambria E, Morency LP. Tensor fusion network for multimodal sentiment analysis. Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Copenhagen, Denmark: Association for Computational Linguistics; 2017. p. 1103–14. https://doi.org/10.18653/v1/D17-1115
- 7.
Zadeh A, Liang PP, Mazumder N, Poria S, Cambria E, Morency LP. Memory fusion network for multi-view sequential learning. Proceedings of the AAAI Conference on Artificial Intelligence. vol. 32. New Orleans (LA): AAAI Press; 2018. p. 5634–41. https://doi.org/10.1609/aaai.v32i1.12021
- 8.
Liu Z, Shen Y, Lakshminarasimhan VB, Liang PP, Bagher Zadeh A, Morency LP. Efficient low-rank multimodal fusion with modality-specific factors. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Melbourne, Australia: Association for Computational Linguistics; 2018. p. 2247–56. https://doi.org/10.18653/v1/P18-1209
- 9.
Tsai YHH, Bai S, Liang PP, Kolter JZ, Morency LP, Salakhutdinov R. Multimodal transformer for unaligned multimodal language sequences. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics; 2019. p. 6558–69. https://doi.org/10.18653/v1/P19-1656
- 10.
Hazarika D, Zimmermann R, Poria S. MISA: Modality-invariant and -specific representations for multimodal sentiment analysis. Proceedings of the 28th ACM International Conference on Multimedia. Seattle (WA): ACM; 2020. p. 1122–31. https://doi.org/10.1145/3394171.3413678
- 11. Wang D, Liu S, Wang Q, Tian Y, He L, Gao X. Cross-modal enhancement network for multimodal sentiment analysis. IEEE Trans Multimed. 2023;25:4909–21.
- 12. Wang D, Guo X, Tian Y, Liu J, He L, Luo X. TETFN: A text enhanced transformer fusion network for multimodal sentiment analysis. Pattern Recognit. 2023;136:109259.
- 13.
Sun K, Tian M. Sequential fusion of text-close and text-far representations for multimodal sentiment analysis. Proceedings of the 31st International Conference on Computational Linguistics. Abu Dhabi, UAE: Association for Computational Linguistics; 2025. p. 40–9. Available from: https://aclanthology.org/2025.coling-main.4/
- 14.
Chen K, Wang S, Ben H, Tang S, Hao Y. Mixture of multimodal adapters for sentiment analysis. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). Albuquerque (NM): Association for Computational Linguistics; 2025. p. 1822–33. https://doi.org/10.18653/v1/2025.naacl-long.90
- 15.
Zhu A, Hu M, Wang X, Yang J, Tang Y, An N. Multimodal invariant sentiment representation learning. Findings of the Association for Computational Linguistics: ACL 2025. Vienna, Austria: Association for Computational Linguistics; 2025. p. 14743–55. https://doi.org/10.18653/v1/2025.findings-acl.761
- 16. Qin Z, Luo Q, Zang Z, Fu H. Multimodal GRU with directed pairwise cross-modal attention for sentiment analysis. Sci Rep. 2025;15(1):10112. pmid:40128290
- 17. Liu S, Gao P, Li Y, Fu W, Ding W. Multi-modal fusion network with complementarity and importance for emotion recognition. Inf Sci. 2023;619:679–94.
- 18. Liu S, Luo Z, Fu W. Fcdnet: Fuzzy cognition-based dynamic fusion network for multimodal sentiment analysis. IEEE Trans Fuzzy Syst. 2025;33(1):3–14.
- 19.
Ghosal D, Majumder N, Poria S, Chhaya N, Gelbukh A. DialogueGCN: A graph convolutional neural network for emotion recognition in conversation. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). Hong Kong, China: Association for Computational Linguistics; 2019. p. 154–64. https://doi.org/10.18653/v1/D19-1015
- 20.
Joshi A, Bhat A, Jain A, Singh A, Modi A. COGMEN: Contextualized GNN based multimodal emotion recognition. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Seattle (WA): Association for Computational Linguistics; 2022. p. 4148–64. https://doi.org/10.18653/v1/2022.naacl-main.306
- 21.
Zhang H, Wang Y, Yin G, Liu K, Liu Y, Yu T. Learning language-guided adaptive hyper-modality representation for multimodal sentiment analysis. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Singapore: Association for Computational Linguistics; 2023. p. 756–67. https://doi.org/10.18653/v1/2023.emnlp-main.49
- 22.
Cho K, van Merriënboer B, Gülçehre C, Bahdanau D, Bougares F, Schwenk H, et al. Learning phrase representations using RNN encoder–decoder for statistical machine translation. Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP). Doha, Qatar: Association for Computational Linguistics; 2014. p. 1724–34. https://doi.org/10.3115/v1/D14-1179
- 23.
Veličković P, Cucurull G, Casanova A, Romero A, Liò P, Bengio Y. Graph attention networks. International Conference on Learning Representations; 2018. Available from: https://openreview.net/forum?id=rJXMpikCZ
- 24.
Devlin J, Chang MW, Lee K, Toutanova K. BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). Minneapolis (MN): Association for Computational Linguistics; 2019. p. 4171–86. https://doi.org/10.18653/v1/N19-1423
- 25. Paszke A, Gross S, Massa F, Lerer A, Bradbury J, Chanan G, et al. PyTorch: An imperative style, high-performance deep learning library. In: Advances in Neural Information Processing Systems. vol. 32; 2019. p. 8024–35.
- 26.
Cao W. Training and analysis code for the multimodal sentiment analysis study. Zenodo; 2026. Software. https://doi.org/10.5281/zenodo.22218345
- 27.
Loshchilov I, Hutter F. Decoupled weight decay regularization. International Conference on Learning Representations; 2019. Available from: https://openreview.net/forum?id=Bkg6RiCqY7
- 28.
Yu W, Xu H, Meng F, Zhu Y, Ma Y, Wu J, et al. CH-SIMS: A Chinese multimodal sentiment analysis dataset with fine-grained annotation of modality. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics; 2020. p. 3718–27. https://doi.org/10.18653/v1/2020.acl-main.343
- 29.
Zhu A, Hu M, Wang X, Yang J, Tang Y, An N. Proxy-driven robust multimodal sentiment analysis with incomplete data. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for Computational Linguistics; 2025. p. 22123–38. https://doi.org/10.18653/v1/2025.acl-long.1075