Figures
Abstract
Monocular 3D Human Pose Estimation (3D HPE) is a core task in computer vision, which provides critical technical support for applications including VR/AR, immersive human-computer interaction, digital human animation, as well as sports motion capture, athletic technique diagnosis, physical fitness assessment and sports rehabilitation training. However, existing methods still face three key limitations: diffusion-based methods lack explicit human anatomical prior constraints in the diffusion process; disentangled-based methods suffer from severe hierarchical error accumulation along the skeletal kinematic tree; and most mainstream methods fail to fully mine the fine-grained hierarchical spatio-temporal correlations of the human skeleton, leading to increased estimation error for high-hierarchy joints. To address these challenges, we propose a novel Hierarchy-Aware Disentangled Diffusion framework with Spatio-Temporal Denoising (HADD) for monocular 3D HPE, which deeply integrates disentanglement strategy with the diffusion model and embeds skeletal hierarchical information into the full diffusion pipeline. In the forward diffusion process, HADD disentangles 3D pose into bone length and bone direction features, and injects Gaussian noise into the two features separately. In the reverse denoising process, we design a Hierarchical Spatial and Temporal Denoising (HSTD) module to accurately model the hierarchical spatio-temporal relationships of the human skeleton and alleviate error amplification of high-hierarchy joints. Meanwhile, a hybrid loss function combining 3D disentanglement loss and 3D pose loss is constructed to realize dual supervision at the bone and joint levels. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets show that HADD achieves consistent performance improvement over state-of-the-art methods. On Human3.6M, it reduces average MPJPE by up to 12.9% compared with traditional disentangled-based methods, and outperforms mainstream diffusion-based probabilistic methods under lightweight inference configurations. On MPI-INF-3DHP, HADD achieves an MPJPE of 29.2 mm, a PCK of 98.5%, and an AUC of 78.1 under single-hypothesis inference, reaching new state-of-the-art performance on this dataset under the single-hypothesis inference setting with only ground-truth 2D joint sequences as input. This work provides an effective solution to the core limitations of existing 3D HPE methods, and also offers reliable technical support for fine-grained motion analysis and quantitative assessment in sports training and rehabilitation scenarios.
Citation: Jiang J, Jian Z, Zheng H, Liang M, Shi Y, Zhang L (2026) HADD: A hierarchy-aware disentangled diffusion framework with spatio-temporal denoising for 3D human pose estimation. PLoS One 21(9): e0359321. https://doi.org/10.1371/journal.pone.0359321
Editor: Gennady S. Cymbalyuk, Georgia State University, UNITED STATES OF AMERICA
Received: March 24, 2026; Accepted: September 12, 2026; Published: September 28, 2026
Copyright: © 2026 Jiang et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All data used in this study are derived from publicly available academic benchmark datasets, and no new original data were generated. Specifically, the Human3.6M dataset is available upon request via its official website (http://vision.imar.ro/human3.6m/description.php), and the MPI-INF-3DHP dataset is available upon request via its official website (https://vcai.mpi-inf.mpg.de/3dhp-dataset/). Both datasets are published by third-party institutions and are subject to their respective user license agreements. This study strictly adheres to the terms of use and data distribution regulations of both datasets. The full code of this study has been open-sourced in the GitHub repository (https://github.com/ygp1234/hadd), including core source code, evaluation scripts, experimental configurations, and etc (Zenodo DOI: https://doi.org/10.5281/zenodo.21786889).
Funding: This work is partially supported by Young Scientific and Technological Talents Program for Higher Education Institutions of Inner Mongolia (NJYT24061); Fundamental Research Funds Program for Colleges and Universities Directly Affiliated to Inner Mongolia Autonomous Region (JY20220249). The funding covered research activities but did not provide support for open access publication fees.
Competing interests: The authors have declared that no competing interests exist.
1. Introduction
3D Human Pose Estimation (3D HPE) is a core research direction in the field of computer vision, with its primary goal being to accurately regress the 3D coordinates of human joints in 3D space from 2D human pose sequences input via monocular vision. This technology provides critical technical support for a wide spectrum of practical applications, including virtual and augmented reality (VR/AR) [1], immersive human-computer interaction [2], digital human animation [3], intelligent video surveillance [4], as well as competitive sports motion capture, athletic technique diagnosis and optimization and public physical fitness assessment. In these scenarios, the precise capture and analysis of human 3D skeletal motion can realize realistic real-time driving of virtual avatars, construction of natural non-contact interactive systems, fine-grained recognition of complex human actions, quantitative assessment of motor function in rehabilitation scenarios, and standardized quantitative evaluation of sports technical movements. Especially in competitive sports training and youth fitness testing, 3D pose estimation breaks through the limitations of traditional manual observation and 2D plane analysis, enabling three-dimensional refined analysis of dynamic movements such as gymnastics, track and field, racket sports and martial arts, assisting coaches in accurately identifying technical defects, formulating personalized training schemes, and effectively reducing the incidence of sports injuries. Current mainstream 3D HPE methods [5,6] generally adopt a two-stage processing strategy: (1) first, high-precision 2D joint coordinates are detected from raw RGB images using mature 2D human pose estimation algorithms; (2) then, a 2D-to-3D lifting operation is performed to compensate for the lack of depth dimension information, ultimately yielding the estimation results of 3D human poses.
In recent years, monocular 3D HPE technology has witnessed rapid development, and researchers have proposed a variety of solutions to address the core challenge of depth ambiguity, which can be mainly divided into two research veins. (1) The first is the spatiotemporal information mining direction: such methods leverage Convolutional Neural Networks (CNNs) [7] or Transformer architectures [8] to fully explore temporal information and spatial context information in video sequences, thereby compensating for the depth information loss in the 2D-to-3D mapping. For example, Bian et al. [9] proposed a novel human pose estimation framework that adopts a task-specific backbone network and an uncertainty-aware regression head, integrating fine-grained local details and global contextual information to improve the accuracy and robustness of pose estimation. These methods significantly improve the continuity and accuracy of pose estimation through the modeling of sequence information. (2) The second is the human pose prior introduction direction: such methods [10,11] constrain the generation space of 3D poses by incorporating prior knowledge such as human anatomical structures and pose distributions, alleviating the estimation uncertainty caused by depth ambiguity. Among them, diffusion model-based methods [12,13] have achieved the current State-of-the-Art (SOTA) performance in this field by directly regressing 3D joint coordinates. For example, Shi et al. [14] proposed the motion-semantics-based diffusion (MS-Diff) framework for monocular video-based 3D human pose estimation, which integrates high-level motion semantics via a Multimodal Diffusion Interaction (MDI) module and applies spectral feature regularization through a Spectral Convolutional Regularization (SCR) module to reduce noise, resolve pose ambiguities, and handle occlusions effectively. And disentangled-based methods [15,16] have taken a different approach by decomposing the 3D pose estimation task into subtasks of bone length and bone direction prediction, and explicitly incorporating human pose prior constraints based on the forward kinematic structure of the human skeleton. For example, Liu et al. [17] proposed a novel framework that achieves viewpoint invariance for 3D human pose estimation by explicitly disentangling motion and view features through a lightweight view estimator and a view-aware inference pipeline.
In recent years, monocular 3D HPE technology has witnessed rapid development, and researchers have proposed a variety of solutions to address the core challenge of depth ambiguity, which can be mainly divided into two research veins. (1) The first is the spatiotemporal information mining direction: such methods leverage Convolutional Neural Networks (CNNs) [7] or Transformer architectures [8] to fully explore temporal information and spatial context information in video sequences, thereby compensating for the depth information loss in the 2D-to-3D mapping. For example, Bian et al. [9] proposed a novel human pose estimation framework that adopts a task-specific backbone network and an uncertainty-aware regression head, integrating fine-grained local details and global contextual information to improve the accuracy and robustness of pose estimation. These methods significantly improve the continuity and accuracy of pose estimation through the modeling of sequence information, and can better adapt to the analysis requirements of continuous high-dynamic movements in sports scenarios. (2) The second is the human pose prior introduction direction: such methods [10,11] constrain the generation space of 3D poses by incorporating prior knowledge such as human anatomical structures and pose distributions, alleviating the estimation uncertainty caused by depth ambiguity. Among them, diffusion model-based methods [12,13] have achieved the current State-of-the-Art (SOTA) performance in this field by directly regressing 3D joint coordinates. For example, Shi et al. [14] proposed the motion-semantics-based diffusion (MS-Diff) framework for monocular video-based 3D human pose estimation, which integrates high-level motion semantics via a Multimodal Diffusion Interaction (MDI) module and applies spectral feature regularization through a Spectral Convolutional Regularization (SCR) module to reduce noise, resolve pose ambiguities, and handle occlusions effectively. And disentangled-based methods [15,16] have taken a different approach by decomposing the 3D pose estimation task into subtasks of bone length and bone direction prediction, and explicitly incorporating human pose prior constraints based on the forward kinematic structure of the human skeleton. This idea is highly consistent with the kinematic analysis logic in sports biomechanics, and has natural application potential in the field of sports technology analysis. For example, Liu et al. [17] proposed a novel framework that achieves viewpoint invariance for 3D human pose estimation by explicitly disentangling motion and view features through a lightweight view estimator and a view-aware inference pipeline.
Challenges. Despite the remarkable progress made by existing methods, both diffusion model-based methods and disentangled-based methods still have their own limitations, and neither of the two types of methods has fully explored the hierarchical structural information of the human skeleton, resulting in an insurmountable bottleneck in model performance. Specifically, there are three core challenges: (1) Existing diffusion model-based methods [12–14] directly add Gaussian noise to the original 3D pose as a whole without performing structural decomposition on the 3D pose, making it impossible to effectively learn explicit prior knowledge such as the temporal consistency of bone length and the variation rules of joint angles in the human skeleton. The learning process of the model lacks the constraints of human body structure, resulting in an insufficiently refined modeling of 3D poses; (2) Although disentangled-based methods [15–17] can explicitly incorporate the anatomical priors of the human skeleton through the decomposition of bone length and bone direction, affected by the tree-like hierarchical structure of the human skeleton, the prediction errors of bone length and bone direction will continuously propagate and amplify along the skeletal hierarchy, forming a severe hierarchical error accumulation effect. This ultimately leads to their overall performance being far lower than that of mainstream diffusion model-based methods, failing to exert the practical effect of human pose priors. In sports scenarios, this error accumulation will be more prominent in terminal joints such as hands and ankles that are critical for technical movement evaluation, directly affecting the reliability of sports motion analysis; (3) Whether it is Transformer-based spatiotemporal modeling methods [8,9], or diffusion model-based and disentangled-based 3D pose estimation methods [18–20], all ignore the fine-grained hierarchical correlations of the human skeleton. The human skeleton can be divided into multiple hierarchies according to the depth of the kinematic tree, and the spatial position and temporal variation of each joint are strongly correlated with its parent and child joints. With the increase of joint hierarchy, the estimation error of existing methods rises significantly, and hierarchical error accumulation has become a key factor restricting model performance.
To address the above core issues, we propose a Hierarchy-Aware Disentangled Diffusion Framework with Spatio-Temporal Denoising (HADD), which innovatively combines the disentanglement strategy with the diffusion model and deeply integrates the hierarchical information of the human skeleton into the forward and reverse processes of the diffusion model. HADD realizes the dual optimization of structured prior learning and hierarchical spatiotemporal modeling.
The core design ideas of HADD are as follows: (1) on the one hand, the 3D pose is disentangled into bone length and bone direction in the forward process of the diffusion model, and we add noise to the disentangled bone length and bone direction respectively, allowing the model to focus on the temporal consistency of bone length and the spatial variation rules of bone direction separately. HADD can learn the explicit priors of human pose more efficiently, while avoiding the hierarchical error accumulation caused by direct disentangled prediction; (2) on the other hand, in the reverse denoising process of the diffusion model, we design a targeted hierarchical spatiotemporal denoiser to strengthen the spatial and temporal correlations between each joint and its parent and child joints, accurately modeling the hierarchical structural information of the human skeleton and alleviating the estimation error of higher-hierarchy joints.
The contributions of our work can be summarized as the following three points:
- We deeply integrate disentangled representation strategy with the diffusion model, and fully embed the hierarchical structural information of the human skeleton into the full pipeline of the diffusion model. This design alleviates two core limitations of existing methods, namely the lack of explicit human anatomical prior constraints in diffusion-based methods and severe hierarchical error accumulation in disentangled-based methods.
- Based on the hierarchical kinematic structure of the human skeleton, we disentangle the 3D pose into two independent feature dimensions: bone length and bone direction. And we add Gaussian noise to each of them in the diffusion forward process. Meanwhile, we design a disentanglement loss function to guide model training, which enables the model to efficiently learn the explicit prior knowledge of human pose.
- For the reverse denoising process, we design a Hierarchical Spatial and Temporal Denoising (HSTD) module, which consists of HSDM and HTDM sub-modules to accurately model the hierarchical spatio-temporal relationships of the human skeleton and mitigate error amplification of high-hierarchy joints. Meanwhile, we design a hybrid loss function to realize dual supervision at the bone level and joint level, further improving the accuracy and structural rationality of pose estimation.
2. Related work
We focus on diffusion-based 3D human pose estimation integrated with the core ideas of disentangled representation and hierarchical spatiotemporal modeling. Relevant research is mainly divided into two directions: 3D human pose estimation and diffusion models and their applications in pose estimation. Existing studies still have room for optimization in terms of human prior fusion, hierarchical information modeling, and diffusion model adaptation. The research progress and limitations of the two directions are reviewed.
2.1. 3D human pose estimation
3D Human Pose Estimation (3D HPE) is one of the core tasks in the field of computer vision, which strongly supports downstream applications such as virtual reality, human action recognition, and human-computer interaction. Its core goal is to regress the 3D spatial coordinates of human joints from monocular 2D pose sequences. Current mainstream methods can be categorized into one-stage methods and two-stage methods: One-stage methods directly extract features from raw RGB images and regress 3D poses. Typical works [21,22] use convolutional neural networks to model image feature volumes for end-to-end pose prediction, yet the estimation accuracy in the depth dimension is difficult to guarantee due to factors such as image background and human occlusion. Two-stage methods [23,24] first detect 2D keypoints in images using mature 2D pose estimation algorithms, and then complete 3D pose inference through 2D-to-3D lifting. This method has more advantages in accuracy and robustness as it fully leverages the mature technology of 2D pose estimation, and thus has become the current mainstream research paradigm.
To alleviate the depth ambiguity problem in the 2D-to-3D mapping process, existing two-stage methods are mainly optimized from two perspectives: spatiotemporal information fusion and introduction of human pose priors. In terms of spatiotemporal information fusion, early studies used convolutional networks to mine temporal information to make up for the information loss in the depth dimension. Subsequent Transformer-based methods further fuse spatial and temporal context information and improve the accuracy of 3D pose estimation by capturing the spatiotemporal correlation of keypoints, with representative works including MixSTE [19], PoseFormer [25], and PoseFormerV2 [26]. For example, Zhang et al. [19] proposed MixSTE (Mixed Spatio-Temporal Encoder), which consists of alternating temporal and spatial transformer blocks to separately model each joint’s temporal motion and inter-joint spatial correlation, and extends the output to the entire input video sequence for better coherence. Zheng et al. [25] proposed a purely transformer-based model for 3D human pose estimation in videos (PoseFormer), which designs a spatial-temporal transformer structure to model intra-frame human joint relations and inter-frame temporal correlations without using any convolutional architectures. Although such methods exhibit strong capabilities in temporal modeling, they generally pay insufficient attention to the fine-grained hierarchical information of the human skeleton. The human skeleton presents a tree-like kinematic structure, and keypoints can be divided into different hierarchies according to kinematic depth. Existing methods ignore such hierarchical correlation, leading to the continuous accumulation of errors with the increase of hierarchies and ultimately affecting the overall estimation performance.
In terms of introducing human pose priors, a class of disentangled methods has become an important research branch. These methods decompose 3D pose estimation into independent prediction tasks of bone length and bone direction, and then reconstruct 3D joint coordinates based on the forward kinematics of the human skeleton, which can explicitly incorporate human anatomical priors such as temporal consistency of bone length, joint angle limits, and human symmetry loss constraints, with representative works including Anatomy3D [18], DKA [27], and Virtual Bones [28]. For example, Chen et al. [18] proposed a novel 3D human pose estimation model that decomposes the task into bone direction prediction and bone length prediction inspired by human skeleton anatomy. Xu et al. [27] proposed a deep kinematics analysis (DKA) pipeline that systematically integrates 2D kinematics structure optimization, human-topology-based articulated motion decomposition, and a temporal refinement module to model both static and dynamic structures for monocular 3D pose estimation. Although disentangled methods can effectively fuse human priors, they have a key problem of hierarchical error amplification: since the 3D joint coordinates are derived layer by layer based on the bone parameters of parent and child joints, the prediction errors of bone length and direction will accumulate continuously in hierarchical propagation, ultimately resulting in their performance being significantly lower than that of current mainstream non-disentangled methods, which also becomes a core bottleneck for the practical application of disentangled methods.
2.2. Diffusion models
As a class of generative models, diffusion models have achieved breakthrough results in image generation, super-resolution, semantic segmentation, multi-modal tasks and other fields by virtue of their strong probabilistic modeling and denoising capabilities. The core framework of the model includes a forward diffusion process and a reverse denoising process: the forward process gradually adds Gaussian noise to the original data until the data eventually degenerates into a random noise distribution; the reverse process recovers the original data from the noisy data by training a denoising model. Classic works such as DDPM (Denoising diffusion probabilistic models) [29] and DDIM (Denoising Diffusion Implicit Models) [30] have simplified and accelerated the training and sampling process of diffusion models, laying a foundation for their applications in various fields.
In recent years, researchers have begun to explore the application of diffusion models in the 3D human pose estimation task, aiming to solve the uncertainty problem in 3D pose estimation by using their probabilistic modeling capabilities. Most existing diffusion-based 3D HPE methods [31,32] take the original 3D pose as the diffusion object, directly add noise to the 3D joint coordinates in the forward process, and then recover the real pose from the noisy 3D pose through the reverse denoising process. For example, Liu et al. [31] proposed a novel Uncertainty-Guided Diffusion Model for 3D Human Pose Estimation (UDM-HPE) that uses uncertainty-aware probabilistic guidance to improve the denoising process. Some works also combine strategies such as multi-hypothesis aggregation, gradient field prior, and heatmap distribution initialization to improve performance, with representative works including D3DP [20], GFPose [33], and DiffPose [34]. For example, Shan et al. [20] proposed a novel probabilistic 3D human pose estimation method named D3DP (Diffusion-based 3D Pose estimation) combined with JPMA (Joint-wise re-Projection-based Multi-hypothesis Aggregation), where D3DP generates multiple 3D pose hypotheses via a diffusion model conditioned on 2D keypoints, and JPMA aggregates these hypotheses joint-by-joint using reprojection errors to produce the final pose. Although such methods can use the advantages of diffusion models to improve the robustness of pose estimation, they have a key problem of insufficient fusion of explicit human priors: directly diffusing 3D joint coordinates without combining the disentangled representation of the human skeleton makes it difficult to explicitly incorporate human anatomical priors such as bone length and direction, resulting in the model being unable to fully utilize the inherent constraints of the human structure and limiting the further improvement of estimation accuracy.
At the same time, existing diffusion-based 3D HPE methods still inherit the problems of traditional Transformer methods, i.e., insufficient exploration of hierarchical spatiotemporal information of the human skeleton. The design of the denoising model does not consider the hierarchical correlation of keypoints and cannot effectively alleviate the problem of hierarchical error accumulation. In addition, the error accumulation problem of disentangled methods and the insufficient prior fusion problem of diffusion models are independent of each other.
3. Methodology
This section elaborates on the technical details of the proposed HADD, which integrates disentanglement strategies into diffusion models and leverages hierarchical spatio-temporal modeling to mitigate error accumulation. The overall pipeline of HADD is illustrated in Fig 1, which consists of four core components: a 3D pose disentanglement strategy that decomposes joint coordinates into bone length and direction, a forward diffusion process with dimension-wise noise injection, a reverse diffusion process empowered by a Hierarchical Spatial and Temporal Denoising Module (HSTD), and a loss function that unifies disentanglement supervision and pose regression constraints. All mathematical formulations are rigorously derived and explained, with notations and definitions standardized for consistency with the state-of-the-art literature in diffusion models and 3D HPE.
Ethics statements. This study only uses publicly available and strictly de-identified academic benchmark datasets (Human3.6M and MPI-INF-3DHP), and no new human subject data collection was conducted. Both datasets have passed the ethical review of their original publishing institutions, all subjects have signed informed consent forms, and the data do not contain any personally identifiable information. The use of the datasets in this study strictly follows the ethical requirements and usage agreements at the time of their release, and complies with the academic ethical norms in the field of computer vision. Therefore, no additional ethical review approval is required.
3.1. Preliminaries
Definition 1 (Joint hierarchy division rule). We select 17 human core joint points as the pose representation units, based on the following rationale [10]: this joint set constitutes the minimal necessary set of core joints in line with human anatomy and kinematics, which can fully cover the key nodes of the motion chain of the trunk, limbs and head. In addition, the 17 joint points can be naturally decomposed into 16 bone segments, which is well adapted to the feature decomposition of bone length and bone direction in the disentangled diffusion strategy. We divide the 17 human joints into 6 hierarchies according to the depth from the root joint of the kinematic tree, with the following specific division, as shown in Table 1:
The division strictly abides by the inherent tree-like kinematic structure of the human skeleton, taking the pelvic root node as the root and classifying the joints step by step from the proximal trunk to the terminal limbs according to their depth in the kinematic tree, which accurately matches the parent-child dependency among human joints.
Definition 2 (Pinhole Camera Projection Model). We adopt the standard pinhole camera model for 3D-to-2D projection, which is fully compatible with the camera calibration parameters provided by the Human3.6M and MPI-INF-3DHP datasets. For the 3D joint coordinate (
is the frame index,
is the joint index), its 2D projection coordinate
on the image plane is jointly determined by the camera intrinsic and extrinsic parameters. The projection formula is as follows:
where is the camera intrinsic matrix provided by the dataset (containing focal length and principal point coordinates),
and
are the camera extrinsic parameters (rotation and translation matrices).
is a conversion function from homogeneous coordinates to inhomogeneous coordinates, defined as:
For a complete 3D pose sequence containing N frames and J joints, the projected 2D pose sequence is
, which is completely consistent in dimension with the input 2D pose sequence
, providing a unified dimensional basis for subsequent reprojection error calculation.
We have established a unified key symbol definition system for the full text (in Table 2). It should be mentioned that the explanations of temporary variables are provided at their first appearance in the text.
3.2. 3D pose disentanglement strategy
The fundamental motivation of the disentanglement strategy is to decouple the high-dimensional and densely correlated 3D joint coordinate optimization problem into low-dimensional, physically interpretable subproblems of bone length and bone direction prediction. This decomposition not only simplifies gradient-based optimization but also naturally embeds human anatomical priors (e.g., temporal consistency of bone length, kinematic constraints of bone direction) into the diffusion model, addressing the limitation of vanilla diffusion-based 3D HPE methods that fail to learn explicit human pose priors due to direct noise injection on 3D joint coordinates.
For a human skeletal structure with J joints, the 3D pose of a single frame is represented as , where each element denotes the 3D coordinate of a joint in the Cartesian space. For a video sequence with N frames, the ground-truth 3D pose sequence is defined as Y0. Each bone is uniquely associated with a parent joint
and a child joint
. For the i-th bone, its ground-truth length
and unit direction vector
are mathematically formulated as:
where denotes the Euclidean L2 norm, which quantifies the physical length of the bone between the parent and child joint. The
,
represent the 3D coordinates of the parent and child joints of the i-th bone in the ground-truth pose. The unit direction vector is normalized to have a magnitude of 1, thus isolating the directional information of the bone from its length. For the entire video sequence, the disentangled ground-truth bone length sequence and bone direction sequence are denoted as L0 and D0, respectively.
The disentanglement of 3D poses into L0 and D0 confers two critical properties for diffusion-based 3D HPE. Properties of the disentangled representation are as follows: (1) The original optimization space of is split into two independent spaces of
(bone length) and
(bone direction), reducing the number of trainable parameters and alleviating the problem of vanishing gradients in high-dimensional space; (2) Bone length is a time-invariant scalar for a given human subject (i.e.,
remains constant across video frames), while bone direction is constrained by physiological joint angle limits. These priors can be explicitly enforced via loss functions, ensuring the physical plausibility of the predicted 3D pose.
From a theoretical perspective, the rationality of confining disentanglement to the forward diffusion process is rooted in the inherent separability and heterogeneity of human skeletal motion features: (1) Human skeletal motion follows a tree-like hierarchical kinematic chain, where bone length (a time-invariant scalar) and bone direction (a spatiotemporally dynamic 3D vector) exhibit distinct statistical distributions. Bone length remains nearly constant across frames for a specific individual, while bone direction is constrained by physiological joint angle limits and exhibits motion-pattern-dependent variation. This inherent separability ensures that disentangled features can be independently modeled without information loss; (2) According to the principle of disentangled representation in information theory, decomposing the joint coordinate space into bone length and bone direction spaces minimizes the mutual information between features, reducing the model’s learning complexity from optimizing a high-dimensional joint distribution p(Y|U) to two low-dimensional marginal distributions p(L|U) and p(D|U) (U denotes input 2D pose sequences). This dimensionality reduction alleviates vanishing gradients in high-dimensional optimization and accelerates convergence.
3.3. Forward diffusion process
The forward diffusion process of HADD follows the Markov chain formulation of Denoising Diffusion Probabilistic Models (DDPM) [29], and we apply the dimension-wise Gaussian noise injection to the disentangled bone length L0 and bone direction D0 instead of the original 3D joint coordinates. This operation enables the diffusion model to learn the temporal dynamics of bone length and direction separately, preserving human anatomical priors during the noise addition process.
For a generic data sample x0, the forward diffusion process gradually adds independent Gaussian noise to x0 over T steps, generating a noisy sample at step t (
). The process is defined as a Markov chain with a fixed noise schedule (the formula format consistent with the DDPM), where the conditional probability distribution of
given x0 is:
where denotes the multivariate Gaussian distribution with mean
and covariance matrix
;
is the cumulative product of
from step 0 to step t;
is the noise schedule coefficient35 that satisfies
;
is the identity matrix of appropriate dimension, ensuring noise is independently added to each dimension of the data.
At step t, the noisy sample can be directly derived from x0 (bypassing intermediate steps) using the reparameterization trick:
where is a standard Gaussian noise sample independent of x0.
Applying Equation 5 to the disentangled bone length sequence L0 and bone direction sequence D0, the noisy bone length and noisy bone direction
at step t are obtained as:
where and
are independent standard Gaussian noise samples for bone length and bone direction, respectively. The use of separate noise samples ensures that the noise dynamics of bone length (a scalar) and bone direction (a 3D vector) are modeled independently, avoiding cross-dimensional interference that would otherwise obscure anatomical priors. As
,
and
, this means
and
converge to pure Gaussian noise. This property is consistent with DDPM and ensures the reverse diffusion process can be trained to recover the original disentangled features from random noise.
3.4. Reverse diffusion process
The reverse diffusion process is the inverse of the forward process, aiming to recover the original disentangled bone length L0 and bone direction D0 from the noisy samples and
(and ultimately from pure Gaussian noise at inference time). Unlike the forward process, which has a fixed closed-form distribution, the reverse process is a learned conditional distribution
, where
denotes the trainable parameters of the model and U is the input 2D pose sequence (the conditional context for 3D HPE). The HSTD (Hierarchical Spatial and Temporal Denoising) module is the core of
, which models the hierarchical spatio-temporal correlations of human joints to denoise
and
effectively. Our Hierarchical Spatial and Temporal Denoising (HSTD) module, combined with the regression head, explicitly predicts the clean 3D joint coordinate sequence
as its only direct output. It does not predict the Gaussian noise term as in vanilla DDPM, nor does it directly predict the clean disentangled bone length
and bone direction
.
In the training stage, the reverse process takes the concatenated noisy features and the input 2D pose sequence U as inputs, and outputs the denoised 3D joint sequence
via the HSTD module and a regression head.
is the final direct output of the denoising network, while
and
are derived by secondary decomposition of
using Equation 2. The conditional distribution of the reverse denoising process is formulated as:
where denotes the HSTD network with trainable parameters
, and
is the standard deviation of the additional Gaussian noise introduced in the DDIM reverse sampling procedure, with
being its corresponding variance. This parameter controls the stochasticity of the sampling process and the variance of the predicted 3D pose distribution. t denotes the current diffusion timestep.
Remark. Camera intrinsic and extrinsic parameters (K, R, g) are not conditions of the denoising network. They are only introduced in the multi-hypothesis selection stage at inference time to compute 2D reprojection errors for hypothesis ranking, and do not participate in the reverse diffusion denoising computation.
The ground-truth 3D pose sequence is denoted as Y0, and the last dimension corresponds to the 3D Cartesian coordinates (x, y, z) of each joint. The denoised 3D pose sequence predicted by the model is denoted as . For both Y0 and
, we perform disentanglement based on the human skeletal kinematic tree to obtain the ground-truth bone length L0, ground-truth bone direction D0 and predicted bone length
, predicted bone direction
, respectively. The disentanglement operation for the i-th bone (connecting parent joint
and child joint
) follows the Euclidean norm and unit vector normalization as defined:
where ,
are the corresponding predicted coordinates of the 3D coordinates of the parent and child joints of the i-th bone in the ground-truth pose. The disentangled bone length is a scalar (L0,
) and the bone direction is a 3D unit vector (D0,
).
In the inference stage of the reverse diffusion process, to address the inherent depth ambiguity problem in monocular 2D-to-3D human pose lifting and further improve the accuracy and robustness of 3D pose prediction, our method draws on the multi-hypothesis sampling strategy proposed in D3DP [20], and combines the fast-sampling advantage of DDIM [30] to optimize the entire inference pipeline. Different from the fixed one-to-one denoising mode in the training stage, the inference stage adopts a multi-hypothesis iterative sampling mechanism for the disentangled bone length and direction features, which fully explores the multi-solution space of 2D-to-3D lifting and effectively avoids the prediction deviation caused by single random noise sampling.
Specifically, we first sample H independent hypothesis samples of initial noisy bone length and initial noisy bone direction from the standard Gaussian distribution at the same time, where the number of hypotheses H is a tunable hyper-parameter that determines the coverage of the solution space. These H groups of initial noisy bone length and direction features are then input into the pre-trained HSTD for the first round of denoising processing. Through the hierarchical spatio-temporal feature modeling and denoising capability of the HSTD, we obtain the denoised bone length set
and denoised bone direction set
corresponding to the H hypotheses. Following the standard DDIM sampling paradigm [30], we derive the timestep transition from the current step t to the next sampling step
(
) for disentangled bone features. For each hypothesis h, we first compute the noise estimates from the current noisy features and the predicted clean bone features:
where and
are the predicted clean bone length and bone direction decomposed from
via Equation 2. The DDIM update rule for noisy bone features from timestep t to the next subsampled timestep
is then given by:
where the stochastic noise standard deviation is defined as:
and ,
are independent fresh standard Gaussian noise samples for each hypothesis. For deterministic DDIM sampling, we set
and omit the stochastic noise term.
We adopt the pinhole camera model defined in Definition 2 for 3D-to-2D projection. The projected 2D coordinate is dimensionally consistent with the input 2D pose U. For H independently generated 3D pose hypotheses
, we first compute the average 2D reprojection error of each hypothesis relative to the input 2D pose sequence U:
where denotes the 2D projection coordinate of the j-th joint in the n-th frame of the h-th hypothesis, calculated by Equation 1. The hypothesis with the minimum reprojection error is selected as the final output:
This selection process relies solely on the input 2D joint sequence and dataset-provided camera calibration parameters, without using any 3D ground truth information, and represents a practical inference pipeline that can be directly deployed.
In summary, the specific implementation in HADD is as follow:
- Model Forward Pass (Step 1): Input the initial noisy features
,
and the 2D pose sequence U. The HSTD outputs the predicted clean 3D joint coordinates
;
- Secondary Disentanglement (Step 2): Decompose
to obtain the predicted bone length
and bone direction
, as shown in Equation 8;
- Noise Estimate Derivation (Step 3): According to Equation 6, we substitute
,
and
,
into the general DDPM formula respectively to obtain the intermediate noise estimates
and
;
- Next DDIM Sampling Step (Step 4): According to DDIM formula, we use the derived
and
to generate the next-step noisy features (Equation 10).
This closed-loop design of “forward disentanglement for noise injection reverse prediction of joint coordinates
secondary disentanglement for supervision” ensures the consistency of disentangled features throughout the entire diffusion pipeline.
3.5. Hierarchical spatial and temporal denoising module (HSTD)
The Hierarchical Spatial and Temporal Denoising module (HSTD) serves as the core component of the HADD framework for the reverse diffusion process, which is designed to explicitly model the hierarchical spatial-temporal correlations of human skeletal joints and mitigate the hierarchical error accumulation inherent in disentangled 3D human pose estimation. Unlike conventional denoisers that treat joint features as independent tokens or only capture global spatial-temporal context, HSTD leverages the tree-like kinematic structure of the human skeleton to enhance the attention weight of hierarchical adjacent joints (parent-child joints) in both spatial and temporal dimensions. It consists of two complementary sub-modules: the Hierarchical Spatial Denoising Module (HSDM) and the Hierarchical Temporal Denoising Module (HTDM), which are alternately stacked for d iterative loops to refine joint feature representations. The overall architecture of HSTD is illustrated in Fig 2, with the core computational logic of HSDM and HTDM formalized in this section.
3.5.1. Input feature embedding and initialization.
The input to HSTD is the concatenated feature tensor of noisy bone length , noisy bone direction
(generated from the forward diffusion process) and 2D pose sequence U, where N denotes the number of input video frames, J is the total number of human skeletal joints, and
corresponds to the number of bone segments (parent-child joint pairs) in the kinematic tree. First, a linear projection layer
is applied to map the concatenated raw features to a high-dimensional latent space with dimension C, yielding the initial joint feature tensor
:
where are
the zero-padded versions of
and
to match the joint dimension J, and
denotes the feature concatenation operation along the channel dimension. To explicitly encode the hierarchical and positional information of joints, two dedicated embedding modules are introduced, i.e., Hierarchical Spatial Position (HSP) Embedding
and Temporal Position (TP) Embedding
.
The prediction error of the parent joint will be continuously propagated and amplified along the kinematic tree to the terminal joints, making Hierarchy 5 the most severely affected by hierarchical error accumulation in existing methods. Joints in the same hierarchy share the same embedding vector, which is learned during training and encodes both spatial position and hierarchical affiliation of joints. We adopt a sinusoidal positional embedding to encode the temporal order of video frames, capturing the sequential nature of human pose sequences. The embeddings are added to the initial feature tensor to obtain the augmented feature tensor for subsequent transformer processing:
where ⊕ denotes the broadcast addition operation to align the dimensions of spatial and temporal embeddings with Finit. We apply a base spatio-temporal transformer block (Spatio-Temporal Encoder) [19] to F to capture global spatial-temporal context, generating the preprocessed joint feature tensor , which serves as the input to HSDM and HTDM. We do not use flattened joint-frame space for mixed attention calculation. Spatial and temporal attention are executed alternately and strictly separated.
3.5.2. Hierarchical spatial denoising module (HSDM).
HSDM is designed to model the hierarchical spatial dependency between parent and child joints, which is the fundamental spatial characteristic of the human skeletal kinematic tree. For any joint j, its spatial position is inherently determined by the position of its parent joint and the bone segment (length and direction) connecting them. Conventional spatial self-attention fails to emphasize this parent-child spatial correlation, leading to insufficient modeling of skeletal structural priors. HSDM addresses this by amplifying the attention weight of parent joints on child joints in the spatial attention matrix, formalizing the kinematic influence of parent joints on child joints in the attention mechanism. Fig 3 shows the architecture of HSDM.
For the preprocessed feature tensor f, the spatial dimension (joint dimension) is decoupled from the temporal dimension (frame dimension) by reshaping f to , where each row represents the feature of a single joint within one frame. We construct a frame-index-based binary mask. The attention scores between joints from different frames are set to
, so that HSDM only calculates valid spatial attention for joints within the same frame. The query Q, key K, and value V matrices for spatial self-attention are generated via linear projections with weight matrices
:
The initial unnormalized spatial attention matrix is computed as:
where is the dimension of each head in the multi-head attention (we set h = 8 for multi-head split), and
is the scaling factor to mitigate the gradient vanishing problem caused by high-dimensional dot products.
Based on the predefined human skeletal kinematic tree, we define the set of hierarchical-related joint triplets as , where
is the set of parent joints, J is the set of target joints, and
is the set of child joints. For each triplet
, the attention weight from the parent joint
to the target joint j (
) is propagated to the attention weights between the target joint j and its child joint
(
) and between the child joint
and the target joint j (
). This propagation formalizes the spatial influence of the parent joint on the entire parent-child joint chain, and the refined attention matrix elements are computed as:
The above attention propagation rule is derived from the inherent tree-like kinematic structure of the human skeleton. The spatial position of a child joint is physically determined by its parent joint and connected bone segments, so the attention correlation between target and child joints should inherit the parent-to-target attention weights. Direct superposition of attention weights will lead to gradient explosion and unstable training. Therefore, we adopt average normalization with a factor of 2 to constrain the amplitude of attention values. The selection of normalization factor is verified by ablation experiments, which proves that the factor
of 2 achieves the best trade-off between constraint effect and training stability.
This refinement operation is applied to all hierarchical-related joint triplets in the skeletal tree), yielding the hierarchical spatial attention matrix AHSDM. The normalized hierarchical spatial attention matrix is obtained by applying the Softmax function to AHSDM, and the final hierarchical spatial feature tensor is computed as:
We reshape the tensor back to to restore the temporal-frame and joint dimension, serving as the input to the subsequent HTDM module. We add a residual connection between the input f and the output fHSDM of HSDM, followed by layer normalization
to stabilize training:
3.5.3. Hierarchical temporal denoising module (HTDM).
HTDM builds on the spatial hierarchical features from HSDM to model the hierarchical temporal correlation between target joints and their child joints. Human motion is a coherent hierarchical dynamic process: the movement of a parent joint drives the correlated movement of its child joints, leading to strong temporal dependencies between hierarchical adjacent joints. Conventional temporal self-attention only captures global temporal context across frames but neglects these fine-grained hierarchical temporal correlations. HTDM addresses this by introducing a residual cross-attention branch that explicitly models the temporal correlation between target joints and their child joints, fusing it with the global temporal self-attention to refine the temporal feature representation. Fig 4 shows the architecture of HTDM.
For the residual spatial feature tensor , we extract the child joint feature tensor
by averaging the features of each target joint j and its all child joints C(j), where
denotes the number of child joints. If joint j has no descendants (j = 0),
,
,
=
. Child joint aggregation rule is as follows:
where denotes the n-th frame. This averaging operation aggregates the hierarchical temporal features of target-child joint pairs, capturing the inherent dynamic correlation between them.
HTDM adopts a dual-branch attention mechanism (self-attention branch + cross-attention branch) to model temporal dependencies. First, the temporal dimension (frame dimension) is decoupled from the joint dimension by reshaping to
, where each row represents the feature of one frame for a single joint. We construct a joint-index-based binary mask. The attention scores between different joints are masked to
, ensuring HTDM only models temporal dependencies of the same joint across all frames. The query
, key
, and value
for temporal self-attention are generated via linear projections:
where are learnable weight matrices. The temporal self-attention matrix
(capturing global temporal context) is computed as:
For the cross-attention branch used to model hierarchical temporal correlation: we reuse the query tensor from temporal self-attention. The key tensor
and value tensor
of cross-attention are obtained by linear projection from the aggregated child joint feature
. For the cross-attention branch (capturing hierarchical temporal correlation), the child joint feature tensor
is reshaped to
. We project it to generate cross-attention key
and value
via linear layers with weight matrices
and
, respectively:
The hierarchical temporal cross-attention matrix is computed using the query matrix
(shared with the self-attention branch) and the cross-attention key matrix
:
The global temporal self-attention and hierarchical temporal cross-attention are fused by element-wise addition to obtain the hierarchical temporal attention matrix AHTDM:
The fused attention matrix is normalized by the Softmax function, and the hierarchical temporal feature tensor is computed as:
3.5.4. Iterative stacking of HSDM and HTDM.
HSDM and HTDM are alternately stacked for d iterative loops to progressively refine the hierarchical spatial-temporal feature representation of joints. For each loop , the output of HTDM in the
-th loop serves as the input of HSDM in the i-th loop. This iterative refinement enables the model to capture multi-scale hierarchical spatial-temporal correlations, gradually mitigating the hierarchical error accumulation from low to high hierarchy joints in the human skeleton.
The final output feature tensor of the d-th HTDM loop is fed into a regression head (consisting of two linear layers and a ReLU activation) to predict the denoised 3D joint coordinates
, completing the denoising process of the reverse diffusion step:
where and
are the linear projection layers of the regression head.
3.5.5. Key design principles.
The design of HSTD adheres to two core principles that align with the biological and kinematic characteristics of human pose. By explicitly modeling the parent-child hierarchical dependency in both spatial and temporal dimensions, HSTD embeds the tree-like kinematic structural prior of the human skeleton into the diffusion model’s denoising process, which is critical for mitigating the error accumulation in disentangled pose estimation. Unlike conventional transformers that capture only global spatial-temporal context, HSTD decomposes the context into hierarchical adjacent context (parent-child joints) and global context, fusing them to achieve fine-grained modeling of human motion dynamics. This design enables HSTD to not only denoise the noisy bone length and direction features in the reverse diffusion process but also refine the joint feature representation by leveraging skeletal hierarchical priors, ultimately improving the accuracy of 3D human pose estimation, especially for high-hierarchy joints that are prone to severe error accumulation in existing methods.
Our framework does not reconstruct 3D poses via forward kinematics using bone length and bone direction. Disentangled variables only act as auxiliary features throughout the inference pipeline. First, initial noise is sampled independently for the bone length and bone direction feature spaces to retain anatomical priors at the beginning of denoising. Second, in each DDIM iteration, the predicted 3D pose is secondarily decomposed into and
to compute noise estimates
and
, which are further used to generate noisy features for the next iteration. Third, the final output is directly the predicted 3D joint coordinates, and disentangled features are not involved in multi-hypothesis selection. This design fully exploits the merits of disentangled representation while eliminating hierarchical error accumulation caused by kinematic chain reconstruction.
4. Loss function
The design of the loss function in HADD is tailored to the core characteristics of the disentangled diffusion framework and hierarchical skeletal modeling, aiming to achieve two key training objectives: (1) guiding the diffusion model to learn explicit human anatomical priors by disentangling 3D pose into bone length and bone direction; (2) constraining the accuracy of the denoised 3D joint coordinates regressed by the HSTD. To this end, we propose a hybrid loss function that fuses a 3D Disentanglement Loss () and a 3D Pose Loss (
). The overall loss
is a linear combination of the two sub-losses, which synergistically supervises the end-to-end training of the HADD model and ensures the consistency between disentangled bone feature learning and 3D pose regression.
4.1. 3D disentanglement loss (
)
The 3D Disentanglement Loss is the core loss for modeling explicit human anatomical priors in the forward diffusion process. Conventional diffusion-based 3D human pose estimation methods directly add noise to the 3D joint coordinates, which makes it difficult for the model to learn the inherent structural constraints of the human skeleton (e.g., fixed bone length range, reasonable bone direction variation). By contrast, HADD disentangles the 3D pose into bone length and bone direction and performs separate diffusion on these two features. The 3D Disentanglement Loss is designed to supervise the model to learn the temporal consistency and anatomical rationality of bone length and bone direction, respectively, through two sub-losses: Bone Length Loss () and Bone Direction Loss (
).
4.1.1. Bone length loss (
).
Human bone length is an approximately fixed anatomical feature in natural motion (excluding negligible physiological deformation), and its temporal variation in video sequences is extremely small. The Bone Length Loss uses the L2 Euclidean loss to constrain the consistency between the predicted bone length and the ground-truth bone length L0, which forces the model to learn the fixed bone length prior and avoid unreasonable bone length scaling in the denoising process. The mathematical definition is:
where and
represent the predicted and ground-truth length of the i-th bone in the n-th frame, respectively. The normalization factor
eliminates the impact of different input frame numbers and joint numbers on the loss scale, ensuring stable training of the model across different datasets and experimental settings.
4.1.2. Bone direction loss (
).
Bone direction characterizes the spatial orientation of the skeletal segment and is the core feature determining human pose and motion. Unlike bone length, bone direction has significant temporal variation but is constrained by human joint motion limits (e.g., the elbow joint cannot bend beyond ). Bone direction is characterized by a 3D unit vector. The L2 loss for bone direction aims to minimize the deviation between predicted unit vectors and ground-truth unit vectors, so that the predicted bone direction can fit the real labels. This model does not introduce explicit joint angle constraints, and the adopted loss cannot limit the anatomical motion range of human joints. The mathematical definition is:
where and
represent the k-th dimension (x/y/z) of the predicted and ground-truth direction vector of the i-th bone in the n-th frame, respectively.
4.1.3. Total 3D disentanglement loss.
The 3D Disentanglement Loss is the sum of the Bone Length Loss and the Bone Direction Loss, which jointly supervises the model’s learning of disentangled bone features and embeds human anatomical priors into the diffusion model’s forward and reverse processes:
This loss function ensures that the model not only preserves the fixed bone length characteristic of the human skeleton but also learns the reasonable spatial variation of bone direction, effectively mitigating the depth ambiguity problem in 2D-to-3D pose lifting and reducing the hierarchical error accumulation caused by disentanglement.
4.2. 3D pose loss (
)
The 3D Pose Loss is a direct regression loss that constrains the accuracy of the final predicted 3D joint coordinates. Although the 3D Disentanglement Loss supervises the learning of bone-level features, it cannot directly guarantee the global accuracy of the 3D pose (e.g., the cumulative error of bone direction may lead to a large deviation in the global position of the end joint). The 3D Pose Loss uses the L2 Euclidean loss to directly measure the position error between the denoised 3D pose and the ground-truth 3D pose Y0 at the joint level, which forces the HSTD to generate accurate 3D joint coordinates and compensate for the local error of disentangled bone features. The mathematical definition is:
where and
represent the k-th dimension (x/y/z) of the predicted and ground-truth 3D coordinates of the j-th joint in the n-th frame, respectively. The normalization factors 1/3, 1/N, and 1/J are used to standardize the loss scale, making it independent of the input frame number, joint number, and coordinate dimension, and ensuring that the 3D Pose Loss has a balanced contribution with the 3D Disentanglement Loss in the total loss.
Notably, the 3D Pose Loss is applied to all skeletal joints (including the root joint), which makes up for the deficiency that the 3D Disentanglement Loss only targets bone segments (). For the root joint (the only joint without a parent bone), the 3D Pose Loss is the only constraint for its 3D coordinate regression, ensuring the accuracy of the global position of the human pose in the 3D space.
4.3. Overall loss function (
)
The overall loss function of the HADD model is a linear fusion of the 3D Disentanglement Loss and the 3D Pose Loss, without introducing additional hyperparameters to weight the two sub-losses. This design is based on the consistent loss scale of the two sub-losses after standardization, which ensures their synergistic supervision of the model training without mutual interference. The mathematical definition of the overall loss is:
The overall loss function realizes the multi-level supervision of the HADD model from two perspectives: (1) The 3D Disentanglement Loss guides the model to learn human anatomical priors and disentangled diffusion features, ensuring the structural rationality of the predicted pose; (2) The 3D Pose Loss directly constrains the regression accuracy of 3D joint coordinates, ensuring the global spatial accuracy of the predicted pose.
This multi-level supervision mechanism solves the two core problems of traditional disentangled methods (hierarchical error accumulation) and traditional diffusion-based methods (inability to learn explicit anatomical priors), and enables the HSTD to fully utilize hierarchical spatial-temporal information for denoising while adhering to human skeletal structural constraints.
4.4. Computational complexity analysis
HADD’s time and space complexity are predominantly determined by three core modules: the 3D pose disentanglement strategy, the forward/reverse processes of the diffusion model, and the HSTD (including HSDM and HTDM), with key influencing hyperparameters being the number of input video frames N, human joints J, attention heads h of the Transformer, diffusion steps T, as well as the number of hypotheses H and iterations Z in the inference phase.
The core computation lies in HSDM and HTDM in HSTD. HSDM performs spatial self-attention on J joints across N frames with a complexity of , while HTDM conducts temporal self-attention and cross-attention on N frames across J joints with a complexity of
. These two modules are alternately stacked for G times, and linear operations such as feature projection and normalization are negligible for the order of complexity. Given
(a typical setting in 3D human pose estimation with video sequences) and G,
as fixed hyper-parameters independent of N and J, the simplified time complexity of the training phase is
, and the original exact complexity is
. The forward noise addition of the diffusion model and the 3D disentanglement operation are linear computations, which do not alter the polynomial order of the overall time complexity. On the basis of the training phase’s complexity, the inference phase introduces two extra hyper-parameters: H and Z. Each hypothesis and iteration requires a complete forward pass of HSTD, leading to a multiplicative factor of
for the core complexity. With
and G,
as fixed constants, the simplified time complexity of the inference phase is
, and the exact form is
. Since H and Z are finite empirical constants, the inference complexity remains a polynomial order without exponential growth.
Spatial complexity is mainly determined by the storage of feature tensors and attention weight matrices in the transformer module. The dominant storage cost comes from the core feature tensor with the shape of , while the attention weight matrices (
for spatial and
for temporal) and intermediate noise tensors of the diffusion model (
for bone length and direction) are far smaller in magnitude and thus negligible. Treating
as a fixed hyper-parameter, the simplified space complexity of the training phase is
, and the exact complexity is
. The only additional spatial overhead in inference comes from the parallel storage of feature tensors and noise data for H Gaussian hypotheses, while the serial DDIM iteration (Z) does not require parallel storage of intermediate results and thus introduces no extra space cost. With
as a fixed constant, the simplified space complexity of the inference phase is
, and the exact form is
.
5. Experiments
5.1. Datasets
To comprehensively and rigorously evaluate the performance of the proposed HADD method on monocular 3D human pose estimation, we conduct extensive experiments on two widely recognized and challenging benchmark datasets for 3D human pose estimation: Human3.6M http://vision.imar.ro/human3.6m/description.php and MPI-INF-3DHP https://vcai.mpi-inf.mpg.de/3dhp-dataset/. These datasets cover diverse human motion scenarios, different shooting conditions and various actor groups, which can fully verify the effectiveness, generalization ability and robustness of the model in handling complex human pose estimation tasks, and also provide benchmark support for sports motion kinematic analysis. The full code of this study has been open-sourced in the GitHub repository (https://github.com/ygp1234/hadd), including core source code, evaluation scripts, experimental configurations, and etc (Zenodo https://doi.org/10.5281/zenodo.21786889).
Human3.6M. Human3.6M [35] is the most classic and widely used large-scale benchmark dataset in the field of 3D human pose estimation, which provides high-quality annotated 3D human pose data collected in a controlled laboratory environment. The dataset contains a total of 3.6 million 3D human pose annotations paired with corresponding RGB images, captured from 4 different camera viewpoints with fixed shooting parameters. It involves 11 professional actors (both male and female) performing a rich set of 17 daily and human-computer interaction related action scenarios, including walking, sitting-standing transition, limb swinging and other basic motion forms, which can provide standardized baseline references for sports action kinematic modeling. Each action scenario has continuous frame sequences, which can well support the modeling of temporal information of human motion, a key requirement for video-based 3D pose estimation. In line with the experimental settings of mainstream SOTA methods [19,25] to ensure the fairness and comparability of experimental results, we adopt the standard data partition strategy for the Human3.6M dataset: the pose sequences of actors S1, S5, S6, S7, S8 are used as the training set to train the HADD model, and the pose sequences of actors S9 and S11 are used as the test set for performance evaluation. Data contains 15 action categories. A sliding window strategy with a fixed window length of 243 frames and a step size of 81 frames is used to sample continuous video sequences. No frame skipping or selective sampling is performed.
MPI-INF-3DHP. MPI-INF-3DHP [36] is a challenging 3D human pose estimation dataset that supplements the limitations of Human3.6M in terms of shooting environment and action diversity, and is more suitable for verifying the generalization ability of the model in relatively unconstrained scenarios. Unlike the fixed laboratory environment of Human3.6M, the data of MPI-INF-3DHP is collected in indoor scenes with natural light changes, and the shooting camera has slight position and angle changes, which brings more realistic challenges such as illumination variation and partial occlusions for pose estimation. The dataset involves 8 actors (4 males and 4 females) with different body shapes and heights, each performing 8 sets of diverse and complex action sequences, including walking on a treadmill, climbing stairs, jumping, dancing, and various sports-related dynamic actions with large joint motion ranges and rich temporal dynamic characteristics, which can effectively test the model’s ability to model high-amplitude sports movements and capture fine-grained temporal correlations of joints. For the data partition of MPI-INF-3DHP, we use the pose sequences of all 8 actors performing 8 action sets as the training set, and the pose sequences of 7 different and unseen action sets as the test set for model evaluation. This partition strategy that separates training and test actions can better verify the cross-action generalization ability of the HADD model, which is a key indicator for the practical application of 3D human pose estimation models in sports scenarios. In the experiment, we use the ground-truth 2D pose as the input of the model for testing (excluding the error introduced by 2D pose detection), which can independently evaluate the performance of the proposed 2D-to-3D lifting framework and avoid the interference of 2D pose estimation errors on the experimental results of the 3D pose estimation module.
5.2. Metrics
To quantitatively evaluate the accuracy of the 3D human pose estimated by HADD, we adopt two most widely used and authoritative evaluation metrics in the field of 3D human pose estimation: Mean Per Joint Position Error (MPJPE) and Procrustes Mean Per Joint Position Error (P-MPJPE). Both metrics take the Euclidean distance in millimeters (mm) as the unit, and the smaller the value of the metrics, the higher the accuracy of the estimated 3D pose. In addition, for the MPI-INF-3DHP dataset, we also introduce Percentage of Correct Keypoints (PCK) and Area Under the Curve (AUC) as supplementary evaluation metrics to comprehensively reflect the model’s performance in different error tolerance ranges. Following the experimental setting of D3DP20, all MPJPE and P-MPJPE results reported are based on J-AGG (Joint Aggregation) calculation, which aggregates the error results of all joints and all test samples to obtain the final evaluation value, ensuring the comprehensiveness and representativeness of the evaluation results.
- (1) Mean Per Joint Position Error (MPJPE): MPJPE is the most basic and direct evaluation metric for 3D human pose estimation, which measures the average Euclidean distance between the 3D coordinates of each estimated joint and the corresponding ground-truth joint over all joints and all test samples. This metric directly reflects the absolute position error of the model’s estimated 3D joints in the 3D space, and is sensitive to the global translation, scaling and rotation errors of the estimated pose relative to the ground-truth pose. The MPJPE for a single sample is calculated as:
(33)
whereand
are the 3D coordinates of the j-th joint of the m-th sample in the estimated pose and the ground-truth pose, respectively.
- (2) Procrustes Mean Per Joint Position Error (P-MPJPE): P-MPJPE is an improved evaluation metric based on MPJPE, which solves the problem that MPJPE is sensitive to the global rigid transformation (translation, rotation and uniform scaling) of the estimated pose relative to the ground-truth pose. Before calculating the joint position error, P-MPJPE first performs Procrustes analysis on the estimated 3D pose to align it with the ground-truth pose, eliminating the global translation, rotation and uniform scaling errors between the two poses. The aligned estimated pose only retains the local joint position error relative to the ground-truth pose, which can more fairly and accurately reflect the model’s ability to estimate the relative spatial structure of the human pose. The P-MPJPE for a single sample is calculated as:
(34)
whereis the 3D coordinate of the j-th joint of the m-th sample in the aligned estimated pose.
For the MPI-INF-3DHP dataset, we additionally use PCK and AUC as supplementary evaluation metrics, which are more suitable for evaluating the pose estimation performance under different error tolerance thresholds and can reflect the model’s performance in identifying valid joint positions. - (3) Percentage of Correct Keypoints (PCK): PCK measures the percentage of estimated joints whose position error is less than a given threshold relative to the ground-truth joints, which reflects the model’s ability to estimate the correct position of human joints within a certain error tolerance range. In 3D human pose estimation, the threshold is usually set as a fixed Euclidean distance (e.g., 50 mm, 100 mm) in the 3D space. The PCK for a single sample is calculated as:
(35)
whereis the preset error tolerance threshold (in mm), and
is the indicator function, which takes the value 1 if the condition in the parentheses is satisfied, and 0 otherwise. In this paper, we adopt the mainstream threshold setting in the field and use
as the standard for calculating PCK on the MPI-INF-3DHP dataset.
- (4) Area Under the Curve (AUC): AUC is the area under the PCK curve with the error tolerance threshold
as the abscissa and the PCK value as the ordinate, which can comprehensively reflect the model’s PCK performance over the entire range of error thresholds (usually from 0 mm to 200 mm). A larger AUC value indicates that the model has a higher correct rate of joint estimation under different error tolerance ranges, and the overall pose estimation performance is more stable. The AUC is calculated by numerical integration of the PCK curve:
(36)
whereis the maximum error threshold (set to 200 mm in this paper). In the actual calculation, we use the trapezoidal integration method to approximate the integral value according to the discrete PCK values at different threshold points, which is a widely used numerical calculation method in the field.
All 3D human pose data in this work are computed under the camera coordinate system. During data preprocessing, two standard operations are applied to all 3D joint coordinates: (1) Root joint alignment: the pelvis is defined as the root joint, and all coordinates are centralized by translating the pelvis to the origin of the 3D coordinate system; (2) Global scale normalization. These two operations alter the distribution of original 3D coordinates, which may cause deviations in cross-comparison with methods using different preprocessing pipelines. All comparative experiments in this paper adopt baselines with consistent preprocessing rules to ensure experimental fairness.
5.3. Baselines
To comprehensively verify the superiority of the proposed HADD, extensive comparative experiments are conducted with state-of-the-art (SOTA) methods for monocular 3D human pose estimation on the Human3.6M and MPI-INF-3DHP datasets. These comparative models are categorized into disentangled-based deterministic methods, non-disentangled-based deterministic methods and probabilistic methods according to their modeling mechanisms and pose regression strategies, and the core characteristics of each model are elaborated as follows:
- (1) Disentangled-based Deterministic Methods: This type of methods decomposes the 3D human pose estimation task into bone length and bone direction prediction based on human anatomical skeleton prior, and reconstructs 3D joint coordinates through forward kinematics, with explicit human body structural constraints integrated into the model.
- DKA [27]: A 3D human pose estimation model that performs deep kinematics analysis to explicitly predict bone lengths and directions with joint angle and symmetry constraints.
- Anatomy3D [18]: A 3D pose estimation model that incorporates human anatomical priors, decomposes poses based on bone structures and maintains bone length consistency in videos.
- Virtual Bones [37]: A model that constructs virtual bones and optimizes bone-decomposed 3D pose regression via motion projection consistency constraints.
- (2) Non-disentangled-based Deterministic Methods: This type of methods directly regresses 3D joint coordinates from 2D pose sequences without decomposing bone features, and mainly mines spatiotemporal correlation of human joints through convolutional or Transformer-based architectures to alleviate depth ambiguity in 2D-to-3D lifting.
- PoseFormer [25]: The first model to apply Transformer to 3D pose estimation for capturing spatiotemporal contextual information of joint sequences.
- P-STMO [36]: A spatial-temporal many-to-one model that mines human motion spatiotemporal features through pre-training to optimize feature extraction and fusion for 3D pose estimation.
- MixSTE19: A model that adopts a seq2seq mixed spatiotemporal encoder to fuse joint spatiotemporal features and improve the accuracy of 2D-to-3D pose conversion.
- PoseFormerV2 [26]: An upgraded version of PoseFormer, a Transformer model that explores frequency domain information to enhance the efficiency and robustness of 3D pose estimation.
- STCFormer [38]: A 3D human pose estimation model that introduces spatio-temporal criss-cross attention to achieve fine-grained interactive learning of joint sequence spatiotemporal features.
- SCT (Spectral Compression Transformer) [39]: Spectral Compression Transformer (SCT) model combines the Line Pose Graph (LPG) and a dual-stream network architecture. It compresses temporal sequences and reduces the computational cost of self-attention via Discrete Cosine Transform.
- (3) Probabilistic Methods: This type of methods models 3D human pose estimation as a probabilistic generation or denoising process, usually introducing multi-hypothesis sampling or diffusion models to capture the uncertainty of 3D pose distribution, and further improve the reliability of pose estimation through hypothesis aggregation or iterative denoising.
- MHFormer [40]: A model that generates multiple pose hypotheses based on multi-hypothesis Transformer and fuses the results to improve the robustness of 3D pose prediction.
- GFPose [33]: A probabilistic estimation model that leverages gradient fields to learn 3D human pose priors and model pose distributions with gradient field constraints.
- D3DP [20]: A diffusion-based 3D human pose estimation model that regards 2D-to-3D pose lifting as a denoising process of pose distribution and integrates multi-hypothesis aggregation strategy.
- DiffPose [34]: A highly reliable diffusion-based 3D pose model that initializes pose distribution with 2D pose heatmaps and depth distributions and constructs a Gaussian Mixture Model (GMM)-based forward diffusion process.
- DCT-DiffPose [13]: A novel framework that integrates a diffusion model with Confidence and Consistency-based Multi-Hypothesis Aggregation (CCMA) and incorporates Discrete Cosine Transform (DCT) into the self-attention mechanism to transform input data into the frequency domain.
- UDM-HPE [31]: A novel Uncertainty-Guided Diffusion Model for 3D Human Pose Estimation (UDM-HPE) uses uncertainty-aware probabilistic guidance to improve the denoising process.
5.3.1. Implementation details.
To fully elaborate the experimental setup of the HADD model, the model architecture parameters, training settings, experimental environment, data preprocessing strategies, and inference configurations are summarized as follows:
For model architecture, the disentanglement module decomposes the 3D pose of 17 human joints into 16-dimensional bone length () and bone direction (
) features with independent forward diffusion processing. The diffusion model adopts the DDPM29 framework for forward diffusion with a cosine noise schedule35 and 1000 total diffusion steps (T). We divide 17 joints into 6 hierarchies based on human skeleton kinematic tree depth. For HSTD, we deploy 128-dimensional learnable hierarchical spatial position (HSP) and temporal position (TP) embeddings, and use a MixSTE spatio-temporal Transformer19 as the backbone with 6 layers, 8 multi-head attention heads and a 512-dimensional hidden dimension, and configure the Transformer FFN with a 2048-dimensional hidden layer and GELU activation. We alternate HSDM and HTDM for 3 loops to model hierarchical spatio-temporal joint correlations, while a two-layer fully connected regression head with a 256-dimensional hidden layer and ReLU activation maps denoised bone features to 3D joint coordinate space.
The training process adopts the AdamW optimizer with an initial learning rate of , weight decay of
, betas of (0.9, 0.999) and epsilon of
, deploys a cosine annealing learning rate scheduler with 300 total epochs and 10 epochs of learning rate warm-up, sets a batch size of 8 with 4-step gradient accumulation (equivalent to an effective batch size of 32), and uses an overall loss function combining equal-weight 3D disentanglement loss
and 3D pose loss
(both L2 norm-based MSE loss). Additional training optimization strategies include mixed precision (FP16) training, gradient clipping with a norm of 1.0, pre-layer normalization for all Transformer modules, and an early stopping strategy with the model checkpoint of minimum validation loss saved for inference.
The experimental hardware environment is a single NVIDIA A100 GPU with 80 GB VRAM, and the software environment is based on Ubuntu system with Python 3.9, PyTorch deep learning framework and CUDA 11.7 computing platform for parallel acceleration, with NumPy and SciPy used for scientific computing such as Procrustes analysis in metric calculation. All baseline results reported in the manuscript are independently reproduced. We have established a unified experimental pipeline to eliminate inconsistencies in input assumptions, preprocessing, evaluation scripts, and inference budgets across all compared methods.
Data preprocessing uniformly processes Human3.6M and MPI-INF-3DHP datasets: 2D joint coordinates are extracted by OpenPose https://github.com/CMU-Perceptual-Computing-Lab/openpose detector and normalized to [0,1] (ground-truth 2D poses are directly used for MPI-INF-3DHP with only normalization), 3D joint coordinates are centered on the pelvis root joint (translated to the 3D coordinate origin) and scaled to a unified scale with random horizontal flip augmentation for training data, and continuous frame sequences are sampled by a sliding window strategy with the input sequence length N fixed at 243 for both datasets.
For inference configurations, a multi-hypothesis sampling and multi-round iteration strategy is adopted. For the basic inference setting, H is set to 5, and for the high-accuracy inference setting, H is increased to 20. Multiple hypothesis samples cover more possible pose distributions in the feature space, which effectively reduces the random sampling error of the diffusion model and improves the robustness of the prediction results. Z rounds of DDIM [30] sampling and denoising iteration are performed on the multi-hypothesis samples. For the basic inference setting, Z is set to 1, and for the high-accuracy inference setting, Z is set to 10. In each iteration, the denoised bone length and direction features from the previous round are used to generate new noisy samples, and the denoising process is repeated to gradually optimize the prediction results and reduce the estimation error of the model.
All experimental results are obtained by 5 repeated runs with different random seeds, with MPJPE and P-MPJPE calculated via the J-AGG strategy for joint and sample error aggregation, and PCK (100 mm threshold) and AUC (0–200 mm threshold range) used as supplementary metrics for MPI-INF-3DHP dataset evaluation. In all tables of this paper, the value following the “” symbol denotes the sample standard deviation (SD) of the 5 independent runs, calculated with the unbiased estimator (denominator
):
where is the sample mean of the target metric. This value reflects the dispersion of individual run results, and is not the standard error of the mean (SEM). For 95% Confidence Interval (CI), we calculate it using the t-distribution for small sample sizes (n = 5), with the formula:
where t0.025,4 = 2.776, is the sample mean, and s is the sample standard deviation.
5.4. Experimental results
Using per-task MPJPE as the core evaluation metric, we analyze the performance stability and statistical significance of the model across different human motion tasks by combining standard deviation (SD) and 95% confidence intervals (95%CI). Comparisons are conducted among three categories of methods: deterministic disentangled basis (Table 3), deterministic non-disentangled basis (Table 4), and probabilistic approaches (Table 5). A dedicated fair comparison with DiffPose is also provided (Table 6). And N, H, Z is the number of input frames, hypotheses, and iterations used in the inference stage, respectively.
The key conclusions and analyses are as follows:
- (1) Deterministic Disentangled Basis Model In the comparison with disentangled-based deterministic methods (Table 3), HADD (N = 243, H = 1, Z = 1) achieves an average MPJPE of 39.7 mm, which represents a performance improvement of 12.9%, 9.9%, and 11.4% compared to DKA (45.6 mm), Anatomy3D (44.1 mm), and VirtualBones (44.8 mm) respectively. It maintains the lowest joint position error across all action tasks, particularly in dynamic actions such as Walk and WalkT, where the error is reduced to 27.6 mm and 27.7 mm, significantly outperforming the comparative models. These results indicate that HADD’s disentanglement strategy effectively addresses the hierarchical error accumulation problem of traditional disentangled methods. By disentangling 3D poses into bone length and bone direction and integrating human anatomical priors, the accuracy of pose estimation is greatly improved. Meanwhile, its ability to model temporal information is also well-suited for the estimation requirements of dynamic actions.
- (2) Deterministic Non-Disentangled Basis Model In the comparison with non-disentangled-based deterministic methods (Table 4), HADD’s average MPJPE of 39.7 mm outperforms all comparative models including PoseFormer (44.3 mm), P-STMO (42.8 mm), MixSTE (40.9 mm), and PoseFormerV2 (45.2 mm). Even when compared to the top-performing STCFormer (40.5 mm) and SCT (40.9 mm), it achieves a slight improvement of 2.0% and 2.9% respectively. The error advantage of HADD is particularly pronounced in actions sensitive to depth information such as Greet, Phone, and Photo. For example, in the Photo action, the error is 46.7 mm, which is lower than STCFormer’s 50.5 mm and SCT’s 50.9 mm. This demonstrates that compared to traditional non-disentangled methods that directly regress 3D joint coordinates, the hierarchical spatio-temporal denoising module integrated in HADD can more fully mine the hierarchical correlation information of the human skeleton, making up for the deficiency of non-disentangled methods in underutilizing skeletal structure priors and effectively alleviating the depth ambiguity problem in the 2D-to-3D lifting process.
- (3) Probabilistic Methods In the comparison with probabilistic methods (Table 5), HADD demonstrates strong competitiveness under different inference configurations: under the basic configuration (N = 243, H = 1, Z = 1), the average MPJPE reaches 39.0 mm, outperforming MHFormer (43.0 mm) and GFPose (45.1 mm); under the conventional configuration (N = 243, H = 5, Z = 1), it achieves 39.5 mm, which is comparable to the performance of mainstream diffusion models such as D3DP, DCT-DiffPose, and UDM-HPE; under the high-precision configuration (N = 243, H = 20, Z = 10), it further reduces to 37.8 mm, slightly outperforming other probabilistic models. Notably, HADD can achieve performance equivalent to that of comparative models with high configurations under lightweight configurations with a small number of hypotheses and iterations. For diffusion model-based methods, HADD can achieve performance comparable to that of methods such as D3DP with high configurations (H = 20, Z = 10) even under lightweight inference settings (e.g., H = 5, Z = 1), reflecting the efficiency of its disentangled diffusion framework and hierarchical spatio-temporal modeling.
- (4) Separate Comparison We further performed a fair head-to-head comparison with the DiffPose series models on the Human3.6M dataset (Table 6), with strict control over input consistency to isolate the performance gain from our core framework design. Specifically, the full DiffPose model initializes the pose distribution with extra heatmaps and depth priors from an off-the-shelf 2D detector, while DiffPose-S and our HADD use only 2D joint sequences as input, aligning with the input settings of all other baselines in this work. All models are tested under the identical configuration: N = 243, H = 5, Z = 50, with MPJPE and P-MPJPE as evaluation metrics. Under the same input constraints, our HADD achieves an MPJPE of 39.2 mm and a P-MPJPE of 30.9 mm, outperforming DiffPose-S (40.1 mm MPJPE, 31.1 mm P-MPJPE) by 0.9 mm and 0.2 mm, respectively. As expected, DiffPose with additional auxiliary inputs delivers a lower error of 36.9 mm (MPJPE) and 28.7 mm (P-MPJPE), creating a notable performance gap with DiffPose-S. Without relying on any extra input information, HADD still achieves consistent performance improvement over the same-input baseline DiffPose-S, which demonstrates the effectiveness of our disentangled diffusion framework and hierarchical spatio-temporal denoising design in resolving depth ambiguity and enhancing pose estimation accuracy. The full DiffPose model uses additional heatmaps and depth priors from an off-the-shelf 2D detector, which provides extra information not available to HADD or any other baseline in our comparison. We explicitly state that this model achieves better performance due to its richer input, and we do not make SOTA claims against this setting.
- (5) Separate Comparison We further validate HADD on the MPI-INF-3DHP dataset, as shown in Table 7. Evaluations are conducted with ground-truth 2D poses as input, using three core metrics: PCK (Percentage of Correct Keypoints), AUC (Area Under the Curve), and MPJPE (Mean Per Joint Position Error) to ensure a fair comparison of 2D-to-3D lifting capability. Our HADD achieves state-of-the-art performance under the single hypothesis (H = 1, Z = 1) setting, outperforming all representative baselines (disentangle-based, transformer-based, diffusion-based SOTA methods) across all metrics. The superior performance on MPI-INF-3DHP stems from HADD’s two core designs: the disentanglement strategy in the forward diffusion process that models explicit human anatomical priors (bone length/direction) and the HSTD that effectively captures hierarchical spatial-temporal joint dependencies, mitigating error accumulation even for complex full-body movements in diverse scenarios.
5.4.1. Hierarchical joint error quantification comparison.
To quantitatively verify the capability of the proposed HADD framework to mitigate hierarchical error accumulation, we conducted comparative experiments on the Human3.6M dataset. We selected representative methods from three mainstream technical categories: disentanglement-based deterministic approaches, Transformer-based deterministic approaches, and diffusion-based probabilistic approaches. All experiments followed a unified experimental protocol with the input frame number set to 243 and a single-hypothesis inference configuration (H = 1, Z = 1). MPJPE and P-MPJPE were adopted as core evaluation metrics. Joints were divided into six hierarchies from the pelvic root node to terminal limbs in accordance with the kinematic tree depth defined in Table 1.
The hierarchical error calculation adopted the standard J-AGG pipeline consistent with main experiments: we first computed the Euclidean error of each joint per frame, then calculated average errors by joint hierarchy, across frames and over multiple repeated tests. The 95% confidence intervals calculated via the t-distribution were also reported. Detailed quantitative results are presented in Table 8, and the overall error variation trends are illustrated in Fig 5 and Fig 6.
Overall, the estimation error of all compared methods rises steadily from Hierarchy 0 to Hierarchy 5. This is an inherent characteristic of the human skeletal kinematic tree: prediction errors of parent joints propagate and amplify along the kinematic chain, peaking at terminal joints. HADD outperforms all baseline methods on both metrics across all six hierarchies, and its performance advantages become more prominent as the joint hierarchy increases. This clearly proves that the hierarchical spatio-temporal design effectively restrains error amplification across hierarchies.
For low-level joints (Hierarchy 0–1) near the root node, error propagation paths are short, leading to marginal performance gaps among different methods. When moving to middle-level joints (Hierarchy 2–3), performance differences gradually emerge. Traditional disentanglement methods suffer from rapid error growth due to error propagation caused by forward kinematic reconstruction. By explicitly modeling the spatial dependencies between parent and child joints via the hierarchical spatial denoising module, HADD effectively limits cross-layer error propagation and gains an obvious performance edge.
HADD achieves the most substantial improvements on high-level terminal joints (Hierarchy 4–5), where error accumulation is most severe. It achieves much lower MPJPE and P-MPJPE values than traditional disentanglement methods and mainstream diffusion-based models. This outcome directly demonstrates that HADD targets and resolves the error amplification issue of high-hierarchy joints, breaking the hierarchical error accumulation bottleneck of existing solutions.
HADD shows more remarkable advantages on the P-MPJPE metric. It indicates that the disentangled diffusion strategy introduces explicit human anatomical priors, enabling the predicted poses to conform better to real human skeletal structures, rather than merely fitting ground-truth absolute coordinates. Combined with the disentanglement loss constrained by fixed bone length, the model maintains reasonable bone proportions during denoising and further reduces structural deviations across hierarchies.
Distinct differences in error growth rates are observed across different technical routes. Traditional disentanglement methods have the steepest error growth curves due to severe error propagation from forward kinematics. Existing diffusion models alleviate this problem to some extent, but hierarchical error accumulation still persists without explicit modeling of skeletal hierarchies. In contrast, HADD has the gentlest error growth trend, which verifies that its hierarchical spatio-temporal denoising module effectively blocks the gradual amplification of errors along the kinematic tree and fundamentally reduces hierarchical error accumulation.
In conclusion, the hierarchical error quantification experiments provide solid quantitative evidence for the core innovations of this work. The results consistently confirm that the integration of disentangled diffusion and hierarchical spatio-temporal modeling enables HADD to effectively alleviate hierarchical error accumulation along the human skeletal kinematic tree. The framework delivers the most prominent gains for high-hierarchy terminal joints that are problematic for conventional methods, fully validating the rationality and effectiveness of the proposed approach.
6. Ablation study
To thoroughly validate the effectiveness of each core design component and strategy in the proposed HADD method, we conduct a comprehensive ablation study on the Human3.6M dataset, using 2D pose sequences extracted by the CPN detector as input. All ablation experiments follow the unified evaluation protocol of J-AGG based MPJPE and P-MPJPE, with consistent hyperparameter settings (N = 243, H = 1, Z = 1) to ensure fair and comparable results. The study focuses on three key aspects: the impact of the disentanglement strategy (input vs. output disentanglement), the contribution of each modular component in HSTD, and the effectiveness of the proposed loss function combination, aiming to quantify the performance gain brought by each design and reveal the internal working mechanism of HADD.
6.1. Impact of disentanglement strategy
To explore how disentangled representation works at different stages of the diffusion pipeline and verify the rationality of applying disentanglement only in the forward diffusion process, we conduct ablation experiments on Disentangled Input (DI) and Disentangled Output (DO). Four configurations are set up: baseline without disentanglement (w/o DI, w/o DO), disentangled output only (w/o DI, w DO), disentangled input only (w DI, w/o DO), and combined disentanglement (w DI, w DO). All experiments follow the unified settings of N = 243, H = 1 and Z = 1, with MPJPE and P-MPJPE as core evaluation metrics.
As shown in Fig 7 and Fig 8, the two disentanglement strategies produce totally different results: compared with the baseline (MPJPE: 40.23 mm, P-MPJPE: 31.56 mm), disentangled input reduces MPJPE to 39.65 mm and P-MPJPE to 31.24 mm, with relative improvements of 1.4% and 1.0% respectively. It splits the high-dimensional coupled optimization of 3D joint coordinates into two low-dimensional independent sub-tasks for bone length and bone direction, enabling the model to explicitly learn human anatomical priors such as temporal consistency of bones and kinematic constraints of bone directions. When only disentangled output is adopted, MPJPE rises to 41.72 mm and P-MPJPE rises to 33.06 mm, increasing by 3.7% and 4.7% against the baseline. This is caused by inherent hierarchical error accumulation in forward kinematics. Since joint coordinates are recursively calculated based on parent joints and bone parameters, prediction errors at low layers will propagate and amplify along the skeletal kinematic tree, resulting in large deviations of terminal joints. The model with both DI and DO achieves an MPJPE of 40.48 mm and a P-MPJPE of 32.08 mm, performing worse than the disentangled input setting. It proves that error accumulation from disentangled output offsets the benefits of disentangled input, and restricting disentanglement to the forward diffusion process is the optimal design.
We further analyze training dynamics and hierarchical errors based on Fig 9 and Fig 10. In terms of training convergence, the model with disentangled input converges faster and maintains lower training loss throughout the whole process. Its initial loss is 12 lower than the baseline, which demonstrates that disentangled representation reduces learning difficulty and accelerates gradient descent. In terms of hierarchical errors, disentangled output aggravates error propagation. Although it slightly lowers errors of Layer 1 joints, errors of joints from Layer 2 to Layer 5 keep rising, and the error of Layer 5 terminal joints increases from 65 mm to 68 mm. By contrast, disentangled input avoids forward kinematic reconstruction in the reverse process and completely eliminates such error propagation.
We repeat the ablation on the more challenging MPI-INF-3DHP dataset (Table 9), and the results are consistent. The disentangled output setting degrades overall performance, while disentangled input achieves the best results: MPJPE reaches 29.2 ± 1.7 mm, PCK is 98.5 ± 1.0% and AUC is 78.1 ± 1.9. This verifies the stable generalization ability of the proposed design across different scenarios. In conclusion, the ablation results validate the core design of HADD. We adopt disentangled bone features for separate noise injection in the forward diffusion phase to embed anatomical priors, and directly regress 3D joint coordinates in the reverse denoising phase to avoid hierarchical error accumulation. This scheme fully leverages the advantages of disentangled representation while overcoming its inherent defects.
6.2. Effect of each module
The HSTD is the core component of HADD for hierarchical spatial-temporal modeling, consisting of three key modules: Hierarchical Embedding (HE), Hierarchical Spatial Denoising Module (HSDM), and Hierarchical Temporal Denoising Module (HTDM). We sequentially add these modules to a baseline model (a diffusion-based 3D HPE model with a standard spatio-temporal transformer but no hierarchical modeling) to verify the incremental contribution of each module, with the overall performance trend shown in Fig 11 and Fig 12. To further quantify the performance gains and verify the cross-dataset generalization of the modular design, we present detailed ablation results with standard deviation and 95% confidence intervals on Human3.6M (Table 10) and MPI-INF-3DHP (Table 11) under the unified experimental setting (N = 243, H = 1, Z = 1).
Hierarchical embedding constructs spatial position embeddings for joints based on their kinematic tree depth (6 hierarchies in total), encoding both the spatial position and hierarchical information of each joint into the feature representation. On the Human3.6M dataset, adding only Hierarchical embedding brings a slight but stable performance improvement: MPJPE decreases from 40.10 ± 1.12 mm to 40.08 ± 1.10 mm, and P-MPJPE from 31.94 ± 0.99 mm to 31.83 ± 0.97 mm. This result indicates that encoding hierarchical information into the initial feature can help the model capture the inherent structural dependencies of the human skeleton, laying a foundation for subsequent hierarchical spatial-temporal modeling. The same trend is observed on the more challenging MPI-INF-3DHP dataset: after introducing HE, MPJPE drops marginally from 30.6 ± 1.8 mm to 30.5 ± 1.8 mm, while PCK and AUC remain stable with a slight upward trend. The consistent mild improvement across both datasets verifies that hierarchical position encoding provides a universal foundational benefit for skeletal feature learning.
On the basis of Hierarchical embedding, adding HSDM (which enhances the parent joint’s spatial influence on child joints via attention weight adjustment) leads to a significant performance gain. On Human3.6M, MPJPE drops by 0.40 mm to 39.68 ± 1.08 mm, and P-MPJPE drops by 0.28 mm to 31.55 ± 0.94 mm. HSDM explicitly models the spatial hierarchical dependencies between parent and child joints, guiding the model to focus on the kinematic structural constraints of the human skeleton during spatial feature learning, which effectively mitigates the spatial error accumulation of high-hierarchy joints. On MPI-INF-3DHP, the introduction of HSDM delivers a more notable improvement: MPJPE is reduced to 29.7 ± 1.7 mm, PCK rises to 98.2 ± 1.1%, and AUC increases to 77.8 ± 1.9. This more pronounced gain on MPI-INF-3DHP can be attributed to the larger motion amplitude and more complex movement patterns in this dataset, where explicit skeletal spatial constraints play a more critical role in narrowing the solution space of depth ambiguity and constraining reasonable pose distributions.
Further integrating HTDM (which captures the temporal correlation between joints and their hierarchical adjacent joints via cross-attention) further refines the model performance, forming the complete HSTD module. On Human3.6M, MPJPE is reduced by an additional 0.03 mm to 39.65 ± 1.07 mm, and P-MPJPE is reduced by 0.31 mm to 31.24 ± 0.82 mm. HTDM makes up for the deficiency of HSDM in temporal hierarchical modeling; by fusing the temporal features of the current joint and its child joints, it effectively captures the temporal consistency of human skeleton movement, and further improves the accuracy of temporal pose estimation. On MPI-INF-3DHP, adding HTDM achieves the final full-module performance of 29.2 ± 1.7 mm MPJPE, 98.5 ± 1.0% PCK and 78.1 ± 1.9 AUC. The larger improvement in structural metrics (P-MPJPE, PCK, AUC) demonstrates that HTDM further optimizes the structural rationality of the predicted pose from the temporal dimension, making the estimated pose more consistent with the inherent kinematic laws of human continuous movement.
The incremental performance improvement of the three modules is consistently observed on both datasets, confirming that the hierarchical spatial-temporal modeling of HSTD is a layered and complementary design: Hierarchical embedding provides the foundation of hierarchical feature encoding, HSDM strengthens spatial hierarchical dependencies to alleviate spatial error accumulation, and HTDM supplements temporal hierarchical correlations to enhance motion coherence. Together, they form a complete hierarchical modeling framework for human skeleton joints. The cross-dataset consistency also rules out the possibility of overfitting to a single dataset scenario, and verifies the universal effectiveness of the modular design for different motion patterns and shooting environments.
6.3. Effect of loss function
HADD adopts a combined loss function consisting of 3D pose loss () and 3D disentanglement loss (
), where
constrains the overall 3D joint coordinate prediction, and
(sum of bone length loss
and bone direction loss
) supervises the model to learn explicit human anatomical priors in the forward diffusion process. We compare the model performance under the single
setting and the combined
setting to verify the necessity of
(Table 12).
The results show that adding the 3D disentanglement loss brings a clear performance improvement: MPJPE is reduced by 0.22 mm, and P-MPJPE by a more significant 0.70 mm. This is because directly constrains the disentangled bone length and bone direction features, forcing the model to strictly follow human anatomical priors (e.g., fixed bone length, plausible joint direction) during diffusion model training, while
only constrains the final 3D joint coordinates and cannot guide the model to learn the fine-grained structural constraints of the skeleton. The larger improvement of P-MPJPE also indicates that
effectively reduces the structural deviation between the predicted pose and the ground truth, making the predicted pose more consistent with the actual human skeleton structure after Procrustes alignment. In conclusion, the 3D disentanglement loss is an essential complement to the 3D pose loss, and their combination forms a dual supervision mechanism that simultaneously constrains the overall joint coordinates and fine-grained skeletal structure, which is crucial for the model to learn accurate 3D human pose representations.
6.4. Effect of Factor 
In Equation 17, a scaling factor is used to perform mean normalization on the hierarchical attention weights. The core purpose of this design is to suppress the gradient explosion caused by the accumulation of attention weights, while ensuring the effective transmission of motion correlation between parent-child joints. It is the core hyperparameter of the HSTD module for modeling the hierarchical spatial dependencies of the human skeleton. To explore the optimal value of this hyperparameter and verify the rationality of the original design, this study conducts a complete ablation experiment on the attention scaling factor, comparing three core indicators: overall model accuracy, hierarchical joint error, and training stability under different scaling factors. All experiments strictly follow the unified control variable rules. We select Human3.6M as the main dataset, and complete HADD model (disentangled diffusion + HSTD module), only the attention scaling factor is modified, while all other structures and parameters remain completely unchanged. Lightweight single hypothesis H = 1, Z = 1. Based on the original scaling factor of 2, representative gradient normalization coefficients are selected to set up 6 control groups, e.g., 1.0, 1.5, 2.0, 2.5, 3.0 and 4.0. All experiments adopt 5 independent repeated runs (with different random seeds), and the results are presented in the format of mean ± standard deviation [95% confidence interval], as shown in Table 13.
From the MPJPE and P-MPJPE results in Table 13, the core conclusions can be drawn:
- (1) The optimal value is
=2.0: The scaling factor of 2 set in the original paper achieves the lowest MPJPE (39.7 mm) and P-MPJPE (30.9 mm) on Human3.6M, which are reduced by 5.5 mm and 4.2 mm respectively compared with the non-normalized
=1.0, with a significant accuracy improvement;
- (2) Too small factor (
2.0): Unstable gradients lead to distortion of the model’s pose structure, and the overall error continues to rise as the factor decreases, reaching the peak when
=1.0;
- (3) Too large factor (
2.0): The spatial attention correlation between parent-child joints is weakened, and HSDM cannot effectively model the hierarchical dependencies of the human skeleton tree. The overall error continues to rise as the factor increases, and when
=4.0, the error is close to the level of non-normalization.
6.5. Robustness experiments
To verify the anti-interference capability of HADD under common input degradation scenarios in real-world applications, we conducted systematic robustness tests on two benchmark datasets: Human3.6M (controlled laboratory environment) and MPI-INF-3DHP (unconstrained real-world scenario). The experiments strictly followed the same data partitioning, input frame length (N = 243), and preprocessing strategies as the main experiments. A unified lightweight single-hypothesis inference configuration (H = 1, Z = 1) was adopted to exclude performance gains from multi-hypothesis sampling and iterative denoising, purely evaluating the inherent anti-perturbation capability of the model architecture itself.
We designed three types of perturbation scenarios covering typical failure modes of real-world 2D detectors: Random Joint Occlusion (simulating partial limb occlusion caused by obstacles), Gaussian Noise in 2D Input (simulating electronic noise during image acquisition and transmission), and Detector Confidence-Based Keypoint Dropout (simulating detection failures of real-world detectors on low-quality regions). All experiments took the undisturbed scenario (0% occlusion / =0 noise) as the baseline, used the J-AGG (Joint Aggregation) strategy to calculate Mean Per Joint Position Error (MPJPE) as the core evaluation metric. All results are the average of 5 runs with different random seeds.
Random joint occlusion simulates the situation where obstacles randomly occlude different parts of the human body in the scene. In the experiment, we randomly set the 2D joint coordinates to zero at ratios of 10%, 20%, 30%, and 40% to gradually increase occlusion intensity, and compared the performance degradation trends of HADD with mainstream SOTA methods. The detailed quantitative results are shown in Table 14, and the MPJPE change curves on the two datasets are shown in Fig 13 and Fig 14.
The experimental results show that the MPJPE of all methods shows an upward trend with the increase of occlusion ratio, but the performance degradation speed of HADD is significantly slower than that of all baseline methods. On the Human3.6M dataset, when the occlusion ratio reaches 40%, the MPJPE of HADD is 59.1 mm, which is 3.8 mm lower than that of the second-best performing UDM-HPE (62.9 mm) and 19.1 mm lower than that of the traditional disentangled method Anatomy3D (78.2 mm). On the more challenging MPI-INF-3DHP dataset, the MPJPE of HADD is only 49.3 mm at 40% occlusion, which is 2.7 mm lower than UDM-HPE (52.0 mm) and 86.9 mm lower than Anatomy3D (136.2 mm). Even in the mild occlusion (10%) scenario, HADD maintains a significant advantage: the MPJPE is 42.9 mm on Human3.6M and 31.8 mm on MPI-INF-3DHP, both superior to contemporary diffusion models D3DP and UDM-HPE.
This advantage stems from the synergistic effect of HADD’s hierarchical spatio-temporal modeling and anatomical priors: when some joints are occluded, the model can reasonably infer the positions of occluded joints by using information from unoccluded parent/child joints in the skeletal hierarchical structure, combined with anatomical constraints such as bone length invariance, thus effectively alleviating information loss caused by occlusion.
Gaussian noise simulates 2D joint coordinate jitter caused by image sensor noise, compression distortion, and transmission interference. In the experiment, we added zero-mean Gaussian noise to the normalized 2D joint coordinates, with noise intensity gradually increasing from 0.01 to 0.09, covering noise interference scenarios from mild to severe. The detailed quantitative results are shown in Table 15, and the MPJPE change curves are shown in Fig 15 and Fig 16.
The results show that HADD exhibits the strongest noise suppression capability. On the Human3.6M dataset, when the noise intensity reaches =0.09, the MPJPE of HADD is 62.1 mm, which is 3.8 mm lower than UDM-HPE (65.9 mm) and 9.1 mm lower than MixSTE (71.2 mm). On the MPI-INF-3DHP dataset, the MPJPE of HADD is 55.7 mm at
=0.09, which is 4.8 mm lower than UDM-HPE (60.5 mm) and 82.8 mm lower than Anatomy3D (138.5 mm). Even in the mild noise (
=0.01) scenario, the performance of HADD is superior to all baseline methods, and its advantage continues to expand with the increase of noise intensity.
This result verifies the synergistic effect of the inherent denoising property of diffusion models and the HSTD module: diffusion models can recover clean pose distributions from noisy inputs, while the HSTD module further filters noise interference by modeling hierarchical spatio-temporal correlations between joints. Meanwhile, the bone length and bone direction priors constrained by the disentanglement loss ensure that the predicted pose always conforms to human anatomical laws, avoiding unreasonable limb deformations caused by noise.
Random coordinate zeroing is an idealized perturbation method and cannot fully reflect the failure modes of real-world 2D detectors. In practical applications, low-confidence outputs from detectors usually correspond to occluded, blurred, or truncated regions rather than randomly distributed joints. For this reason, we supplemented the keypoint dropout experiment based on CPN detector confidence, which is currently the standard realistic perturbation protocol for evaluating occlusion robustness in the 3D HPE field.
In the experiment, we used the original output confidence scores of the CPN detector on the two datasets, dropped 10%, 20%, 30%, and 40% of keypoints in ascending order of confidence scores (set their coordinates to zero), and retained the same single-hypothesis inference configuration as the random occlusion experiment. The detailed quantitative results are shown in Table 16.
The experimental results show that the performance degradation of all methods under confidence-based dropout is milder than that under random occlusion, because the dropped keypoints are those with inherently poor detection quality rather than high-quality reliable joints. Even so, HADD still maintains a significant performance advantage. On the Human3.6M dataset, the MPJPE of HADD is 49.8 mm at 30% keypoint dropout, which is 2.0 mm lower than UDM-HPE (51.8 mm); at 40% dropout, it is 56.3 mm, which is 3.3 mm lower than UDM-HPE (59.6 mm). On the MPI-INF-3DHP dataset, the MPJPE of HADD is 39.2 mm at 30% dropout, which is 2.1 mm lower than UDM-HPE (41.3 mm); at 40% dropout, it is 46.5 mm, which is 2.6 mm lower than UDM-HPE (49.1 mm).
The three types of perturbation experiments designed in this study (random zeroing, Gaussian noise, confidence dropout) constitute a rigorous stress test system, verifying the inherent anti-interference ability of HADD under input noise and partial occlusion. Among them, the keypoint dropout experiment based on detector confidence is closer to the failure mode of real 2D detectors, indicating that HADD has better robustness under such common degradation scenarios. However, it should be clearly stated that these experiments are still conducted in controlled dataset environments and do not cover challenges such as complex lighting changes, multi-person interactions, and severe full-body occlusion in unconstrained in-the-wild scenarios. Therefore, this study does not claim that HADD has general real-world robustness, and its generalization ability in unconstrained in-the-wild scenarios still needs further verification.
6.6. Visual analytics
To intuitively verify the effectiveness of HADD in 3D human pose estimation, we conducted a visual comparison between HADD and mainstream SOTA methods on typical action sequences of the Human3.6M test set S11 in Fig 17. Both 3D spatial views and 2D projection views of the 3D poses are presented, with comparisons focusing on three core aspects: the spatial accuracy of joint positions, the continuity of limb movements, and the estimation precision of high-hierarchy joints (hands and head). In the experiment, all methods adopted a unified input configuration (243-frame 2D joint sequences) and a single-hypothesis inference setting (H = 1, Z = 1) to ensure the fairness of the visual results; the predicted 3D joint coordinates were uniformly centered on the pelvis root joint and normalized in scale to eliminate the interference of global position and scale differences on visual comparisons.
All original images from the Human3.6M dataset are removed. This figure is independently generated from raw 3D joint coordinates, and only human skeletons are presented for illustration.
The predicted results of HADD have no obvious joint overlap or position confusion in the 2D projection view, while the comparison methods all have varying degrees of depth estimation ambiguity. This further verifies that HADD effectively alleviates the depth ambiguity problem in monocular 3D human pose estimation through the disentangled diffusion strategy and hierarchical spatio-temporal modeling, realizing more physically reasonable and visually realistic 3D pose prediction. Overall, the visualization results are highly consistent with the quantitative experimental conclusions, intuitively demonstrating the comprehensive advantages of HADD in the detailed fitting of 3D human poses, movement continuity, and high-hierarchy joint estimation. The predicted 3D poses not only comply with the physical constraints of human anatomical structures but also match the dynamic laws of real human movements.
6.7. Computational efficiency
To comprehensively assess the practical applicability of the proposed HADD method, we conduct a detailed computational efficiency analysis on the Human3.6M dataset, alongside a comparison with state-of-the-art 3D human pose estimation methods (including disentangle-based, non-disentangle-based, and diffusion probabilistic approaches). We measure the core efficiency metrics of all compared methods under a unified experimental environment (single NVIDIA A100 80GB GPU, CUDA 11.7, PyTorch 2.0.1, mixed-precision FP16 training and inference). For the training phase, all models were uniformly trained for 300 epochs with an effective batch size of 32 (batch size 8 + 4-step gradient accumulation), and all diffusion models adopted a cosine noise schedule with 1000 total diffusion steps. All experiments are performed on a unified hardware platform. We evaluate four core efficiency metrics: FLOPs (Floating-Point Operations) (G), inference time per frame (ms), parameter counts and training time, as shown in Table 17. Only the pure computational time for end-to-end model training is measured, excluding data loading, logging, and model checkpoint writing time. For the inference phase, a typical single-video deployment configuration with a batch size of 1 is adopted. All diffusion-based methods in Table 17 use the lightweight single-hypothesis single-step inference setting (number of hypotheses H = 1, DDIM sampling steps Z = 1). All metrics are measured exclusively for the 2D-to-3D pose lifting module, excluding 2D pose detection, data preprocessing, and postprocessing steps, ensuring exactly the same statistical scope as all baseline methods.
While HADD is slower and has higher computational complexity than several deterministic baselines (MixSTE, STCFormer, Anatomy3D, Virtual Bones), it outperforms all diffusion/probabilistic baselines in both inference speed and FLOPs, while delivering superior or comparable pose estimation accuracy. This makes HADD the most efficient diffusion-based method among the compared approaches, with a single-frame inference latency that meets the basic real-time requirements of many downstream applications. In summary, HADD does not achieve absolute optimality across all dimensions, but it precisely addresses the core pain point of diffusion-based methods: “high accuracy but difficult deployment.” Compared with deterministic methods, it achieves a significant accuracy improvement at an acceptable efficiency cost. Compared with other diffusion-based methods, it maintains state-of-the-art (SOTA) accuracy while significantly reducing computational overhead.
6.8. Discussion
Experimental results and ablation studies fully validate the effectiveness and superiority of our proposed HADD model for monocular 3D human pose estimation. Its core innovation lies in the deep integration of disentanglement strategy and diffusion model, which addresses two critical bottlenecks of existing mainstream methods: the lack of explicit human anatomical prior constraints in diffusion-based approaches, and severe hierarchical error accumulation in disentanglement-based methods. In the forward diffusion process, we disentangle 3D human pose into bone length and bone direction, and apply Gaussian noise to the two features separately instead of the whole 3D joint coordinates. This design enables the model to independently learn the temporal consistency of bone length and spatial variation rules of bone direction, efficiently embedding human anatomical priors and alleviating depth ambiguity in 2D-to-3D mapping. Benefiting from the improved accuracy of high-hierarchy terminal joints and the robustness to dynamic motions, the framework can provide reliable quantitative data support for competitive sports technique diagnosis, sports rehabilitation gait assessment, and youth physical fitness testing.
Unlike traditional disentanglement methods that directly regress bone parameters and reconstruct poses via forward kinematics, we confine disentanglement to the forward diffusion stage and take joint coordinates as the final optimization target in the reverse process, completely avoiding hierarchical error propagation along the skeletal kinematic tree. In the reverse diffusion process, our Hierarchical Spatial and Temporal Denoising (HSTD) module explicitly models the hierarchical correlation of parent-child joints via two sub-modules: HSDM strengthens the spatial attention weight of parent joints to child joints to characterize hierarchical spatial dependency, while HTDM fuses global temporal context and hierarchical temporal correlation between target and child joints to capture the linkage motion law of joints. The alternate iteration of the two modules achieves fine-grained modeling of hierarchical spatio-temporal information and effectively mitigates error accumulation in high-hierarchy joints. Our hybrid loss function combining 3D disentanglement loss and 3D pose loss realizes dual constraints at the bone and joint levels. The disentanglement loss guides the model to learn anatomical priors for structural rationality, while the pose loss directly constrains joint coordinate regression for global spatial accuracy. Ablation results confirm that this dual supervision brings significant accuracy improvement, especially for the P-MPJPE metric. This dual-level supervision mechanism at bone and joint dimensions aligns well with the analytical logic of sports biomechanics, and can directly output structured kinematic parameters for movement technique analysis.
HADD also achieves a favorable balance between accuracy and computational efficiency. It outperforms or matches mainstream diffusion-based probabilistic methods with fewer hypotheses and iterations in inference, with reasonable single-frame latency and FLOPs to meet the real-time requirements of downstream applications such as VR/AR, digital human animation, intelligent video surveillance, as well as real-time action feedback in sports training and on-site competition motion capture. Meanwhile, the model only takes 2D joint sequences as input without additional auxiliary information, showing strong deployment compatibility. Although HADD shows superior anti-interference ability under three types of controlled perturbations, its performance still degrades significantly under extreme occlusion scenarios. In addition, this study did not evaluate the robustness of HADD on real in-the-wild datasets (such as 3DPW, Human3.6M in-the-wild), nor did it test its adaptability to outputs from different 2D detectors. For competitive sports scenarios, challenges such as high-speed motion blur, multi-person limb occlusion and complex field backgrounds also remain to be further adapted. These are key issues that need to be addressed in future work.
This work has several limitations verified by experimental results. First, the model degrades obviously under heavy occlusion. When the occlusion/keypoint dropout ratio reaches 40%, its MPJPE rises to 56.3 mm on Human3.6M and 49.3 mm on MPI-INF-3DHP, indicating limited compensation ability for severe occlusion. Second, all experiments adopt single-person datasets, so the model cannot adapt to multi-person interaction and mutual occlusion scenarios, and is difficult to directly apply to posture analysis in multi-person competitive sports. Third, tested only on laboratory datasets, the model shows poor resistance to complex real-world interference, as reflected by obvious error growth under high Gaussian noise. Fourth, this framework only focuses on 3D joint estimation and cannot support human mesh and surface reconstruction. In future work, we will add occlusion-aware modules, extend the model for multi-person tasks, adopt in-the-wild datasets to improve generalization, explore joint pose and mesh co-estimation, and carry out lightweight adaptation research for typical sports scenarios to support real-time posture analysis on training grounds and portable devices.
7. Conclusion
In this paper, we propose HADD to solve the core bottlenecks of existing methods. By disentangling 3D pose into bone length and bone direction for dimension-wise noise addition in the forward diffusion process, we embed human anatomical priors while avoiding error accumulation of traditional disentanglement methods. In the reverse process, our HSTD module models the hierarchical spatio-temporal dependency of parent-child joints to alleviate error amplification in high-hierarchy joints. Combined with a hybrid loss function, the model achieves dual optimization of structural prior learning and hierarchical spatio-temporal modeling, and provides technical support for sports motion analysis and rehabilitation effect assessment. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets show that HADD achieves consistent performance improvement over existing SOTA methods: it reduces average MPJPE by up to 12.9% over traditional disentanglement-based methods, optimizes average MPJPE by 2.0% over non-disentangled deterministic methods, and outperforms mainstream diffusion-based probabilistic methods under lightweight inference configurations, with significant advantages in high-hierarchy joint estimation and dynamic action sequences, suitable for fine-grained capture of dynamic athletic movements. For future work, we will focus on different directions: improving occlusion robustness via multi-modal information fusion; enhancing generalization ability in unconstrained real-world scenarios with large-scale datasets and self-supervised learning; expanding the model to realize joint 3D pose and mesh estimation; performing lightweight optimization for edge devices and real-time applications, and exploring its targeted adaptation for sports training scenarios. In addition, models incorporating additional auxiliary information (such as heatmaps and depth priors) can achieve further performance improvements, which is an important direction for future work.
References
- 1.
Trojak M, Jurgas A, Stanuch M, Kownacki L, Skalski A. Mixed Reality 6D Object Pose Estimation using Deep Learning for Visual Markerless Surgical Navigation. In: IEEE Conference on Virtual Reality and 3D User Interfaces, VR 2025 - Abstracts and Workshops, Saint Malo, France, March 8-12, 2025. IEEE; 2025. pp. 947–52. https://doi.org/10.1109/VRW66409.2025.00193
- 2. Gao P. Key technologies of human–computer interaction for immersive somatosensory interactive games using VR technology. Soft Comput. 2022;26(20):10947–56.
- 3. Benbelkheir Y, Lerga A, Ardaiz O. A virtual reality direct-manipulation tool for posing and animation of digital human bodies: An evaluation of creativity support. Multimodal Technol Interact. 2024;8(7):60.
- 4.
Okano M, Kanai K, Katto J. Performance Evaluations of IEEE 802.11ad and Human Pose Detection towards Intelligent Video Surveillance System. In: IEEE International Conference on Consumer Electronics - Taiwan, ICCE-TW 2019, Yilan, Taiwan, May 20-22, 2019. IEEE; 2019. pp. 1–2. https://doi.org/10.1109/ICCE-TW46550.2019.8991819
- 5.
Mwali S, Kao C, Chi C, Lin Y. Using OpenPose for Enhancing CPR Training: A Data-Driven Approach. In: IEEE Symposium on Computers and Communications, ISCC 2025, Bologna, Italy, July 2-5, 2025. IEEE; 2025. pp. 1–8. https://doi.org/10.1109/ISCC65549.2025.11325798
- 6. Huang J, Hashim H, Norman H, Zaini MH, Zhang X. Automatic detection of teacher behavior in classroom videos using AlphaPose and Faster R-CNN algorithms. PeerJ Comput Sci. 2025;11:e2933. pmid:40567803
- 7. Dong K, Zhou Y, Riou K, Yun X, Sun Y, Subrin K, et al. Spatial–temporal–geometric graph convolutional network for 3-D human pose estimation from multiview video. IEEE Trans Instrum Meas. 2025;74:1–13.
- 8. Pan X, Li G, Zhang N, Li J. 3D human pose estimation based on a Hybrid approach of Transformer and GCN-Former. J Visual Commun Image Represent. 2026;115:104696.
- 9. Bian S, Wang J, You Y, Yu Z, Sun Y, Wu W. Enhancing human pose estimation accuracy with pyramid fusion vision transformers. Visual Computer. 2026;42(2):143.
- 10. Kappan MM, Sandoval EB, Meijering E, Cruz F. A survey on deep learning for 2D and 3D human pose estimation. Artif Intell Rev. 2026;59(1):32.
- 11. Lee H, Ryu J. Toward efficient generalization in 3D human pose estimation via a canonical domain approach. IEEE Access. 2025;13:79118–36.
- 12. She L, Liang H, Sun H, Chen Y. 3D hand pose estimation based on hand denoising diffusion probabilistic robustness model. Int J Patt Recogn Artif Intell. 2025;40(02).
- 13. Zhong L, Chen F, Zheng B, Feng R, Zhang L, Wan J. DCT-DiffPose: A lightweight diffusion model with multi-hypothesis for 3D human pose estimation. IEEE Access. 2025;13:73319–31.
- 14. Shi L, Wu S, Yang S, Qiu W, Qiang D, Zhao J. Text‐guided diffusion with spectral convolution for 3D human pose estimation. Comput Graph Forum. 2025;44(7).
- 15. Asebe Teka N, Gedamu Alemu K, Assefa M, Akmel F, Zhou Z, Wu W, et al. AMT-Net: Adversarial motion transfer network with disentangled shape and pose for realistic image animation. IEEE Access. 2025;13:92712–29.
- 16. Jiang M, Li J, Lu Y, Zhang F. D-NeuRA: customizable dynamic neural rendering for human avatars with disentangled pose and appearance. J King Saud Univ Comput Inf Sci. 2025;37(8).
- 17.
Liu Y, Jiang H, Ouyang X. Demo: FreePose: Real-Time View-Invariant 3D Human Pose Estimation via Motion-View Disentanglement. In: Proceedings of the 31st Annual International Conference on Mobile Computing and Networking, ACM MOBICOM 2025, Hong Kong, November 4-8, 2025. ACM; 2025. pp. 1189–91. https://doi.org/10.1145/3680207.3765586
- 18. Chen T, Fang C, Shen X, Zhu Y, Chen Z, Luo J. Anatomy-aware 3D human pose estimation with bone-based pose decomposition. IEEE Trans Circuits Syst Video Technol. 2022;32(1):198–209.
- 19.
Zhang J, Tu Z, Yang J, Chen Y, Yuan J. MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022. IEEE; 2022. pp. 13222–32. https://doi.org/10.1109/CVPR52688.2022.01288
- 20.
Shan W, Liu Z, Zhang X, Wang Z, Han K, Wang S, et al. Diffusion-Based 3D Human Pose Estimation with Multi-Hypothesis Aggregation. In: IEEE/CVF International Conference on Computer Vision, ICCV 2023, Paris, France, October 1-6, 2023. IEEE; 2023. pp. 14715–25. https://doi.org/10.1109/ICCV51070.2023.01356
- 21. Pramono A, Chang I-C, Dewi Puspasari B. PRAMGCN-Net: 3D human pose estimation with a parameterized routing adjacency modulation graph convolutional network. IEEE Access. 2025;13:159050–63.
- 22. Meng X, Liu Z. Yoga pose recognition using dual structure convolutional neural network. PeerJ Comput Sci. 2025;11:e2907. pmid:40567706
- 23. Yan X, Xie J, Liu M, Li H, Gao H. Hierarchical local temporal network for 2D-to-3D human pose estimation. IEEE Internet Things J. 2025;12(1):869–80.
- 24. Wei F, Xu G, Wu Q, Qin P, Pan L, Zhao Y, et al. Whole-body 2D human pose estimation based on human keypoints distribution constraint and adaptive Gaussian factor. Complex Intell Syst. 2025;11(10).
- 25.
Zheng C, Zhu S, Mendieta M, Yang T, Chen C, Ding Z. 3D Human Pose Estimation with Spatial and Temporal Transformers. In: 2021 IEEE/CVF International Conference on Computer Vision, ICCV 2021, Montreal, QC, Canada, October 10-17, 2021. IEEE; 2021. pp. 11636–45. https://doi.org/10.1109/ICCV48922.2021.01145
- 26.
Zhao Q, Zheng C, Liu M, Wang P, Chen C. PoseFormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. IEEE; 2023. pp. 8877–86. https://doi.org/10.1109/CVPR52729.2023.00857
- 27.
Xu J, Yu Z, Ni B, Yang J, Yang X, Zhang W. Deep Kinematics Analysis for Monocular 3D Human Pose Estimation. In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, WA, USA, June 13-19, 2020. Computer Vision Foundation. IEEE; 2020. pp. 896–905. https://doi.org/10.1109/CVPR42600.2020.00098
- 28.
Wang G, Zeng H, Wang Z, Liu Z, Wang H. Motion Projection Consistency Based 3D Human Pose Estimation with Virtual Bones from Monocular Videos. CoRR. 2021. http://arxiv.org/abs/2106.14706
- 29.
Ho J, Jain A, Abbeel P. Denoising Diffusion Probabilistic Models. In: Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual; 2020.
- 30.
Song J, Meng C, Ermon S. Denoising Diffusion Implicit Models. In: 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReview.net; 2021.
- 31. Liu Z, Wang Y. Uncertainty-guided diffusion model for 3D human pose estimation. Neurocomputing. 2025;641:130306.
- 32. Li J, Bai Z, Kong D, Chen D, Li Q, Yin B. 3d human pose estimation based on conditional dual-branch diffusion. Multim Syst. 2025;31(1):5.
- 33.
Ci H, Wu M, Zhu W, Ma X, Dong H, Zhong F, et al. GFPose: Learning 3D Human Pose Prior with Gradient Fields. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. IEEE; 2023. pp. 4800–10. https://doi.org/10.1109/CVPR52729.2023.00465
- 34.
Gong J, Foo LG, Fan Z, Ke Q, Rahmani H, Liu J. DiffPose: Toward More Reliable 3D Pose Estimation. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. IEEE; 2023. pp. 13041–51. https://doi.org/10.1109/CVPR52729.2023.01253
- 35. Ionescu C, Papava D, Olaru V, Sminchisescu C. Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments. IEEE Trans Pattern Anal Mach Intell. 2014;36(7):1325–39. pmid:26353306
- 36.
Shan W, Liu Z, Zhang X, Wang S, Ma S, Gao W. P-STMO: Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation. In: Computer Vision - ECCV 2022 - 17th European Conference, Tel Aviv, Israel, October 23-27, 2022, Proceedings, Part V. vol. 13665 of Lecture Notes in Computer Science. Springer; 2022. pp. 461–78. https://doi.org/10.1007/978-3-031-20065-6_27
- 37. Wang G, Zeng H, Wang Z, Liu Z, Wang H. Motion projection consistency-based 3-D human pose estimation with virtual bones from monocular videos. IEEE Trans Cogn Dev Syst. 2023;15(2):784–93.
- 38.
Tang Z, Qiu Z, Hao Y, Hong R, Yao T. 3D Human Pose Estimation with Spatio-Temporal Criss-Cross Attention. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023. IEEE; 2023. pp. 4790–9. https://doi.org/10.1109/CVPR52729.2023.00464
- 39. Zheng Z, Yang L, Zhu H, Ye M. Spectral compression transformer with line pose graph for monocular 3D human pose estimation. Pattern Recognit. 2026;172:112507.
- 40.
Li W, Liu H, Tang H, Wang P, Gool LV. MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 18-24, 2022. IEEE; 2022. pp. 13137–46. https://doi.org/10.1109/CVPR52688.2022.01280