Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

TrackStudio: An integrated toolkit for markerless tracking

  • Hristo Dimitrov ,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Validation, Visualization, Writing – original draft, Writing – review & editing

    hristo.dimitrov@mrc-cbu.cam.ac.uk

    Affiliation MRC Cognition and Brain Sciences Unit, University of Cambridge, Cambridge, United Kingdom

    ⨯
  • Viktorija Pavalkyte,

    Roles Conceptualization, Data curation, Project administration, Writing – review & editing

    Affiliation MRC Cognition and Brain Sciences Unit, University of Cambridge, Cambridge, United Kingdom

    ⨯
  • Giulia Dominijanni,

    Roles Conceptualization, Writing – review & editing

    Affiliations MRC Cognition and Brain Sciences Unit, University of Cambridge, Cambridge, United Kingdom, Center for Neuroprosthetics and School of Engineering, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland

    ⨯
  • Tamar R. Makin

    Roles Funding acquisition, Resources, Supervision, Writing – review & editing

    Affiliation MRC Cognition and Brain Sciences Unit, University of Cambridge, Cambridge, United Kingdom

    ⨯

Abstract

Markerless motion tracking has advanced rapidly in the past 10 years and currently offers powerful opportunities for behavioural, clinical, and biomechanical research. While several specialised toolkits provide high performance for specific tasks, using existing tools still requires substantial technical expertise. There remains a gap in accessible, integrated solutions that deliver sufficient tracking for non-experts across diverse settings. TrackStudio was developed to address this gap by combining established open-source tools into a single, modular, GUI-based pipeline that works out of the box. It provides video recording preprocessing, recording synchronisation, automatic 2D and 3D pose estimation, and visualisation without requiring any programming skills. We supply a user guide with practical advice for video acquisition, camera calibration, video synchronisation, and experimental setup, alongside documentation of common pitfalls and how to avoid them. To validate the toolkit, we tested its performance across three environments using either low-cost webcams or high-resolution cameras, including challenging conditions for body position, lighting, space, and obstructions. Across 76 participants, average inter-frame correlations exceeded 0.98 and average triangulation errors remained low (<13.6 mm for hand tracking), demonstrating stable and consistent tracking. We further show that the same pipeline can be extended beyond hand tracking to other body and face regions. TrackStudio provides a practical, accessible route into markerless tracking for researchers or laypeople who need reliable performance without specialist expertise.

Introduction

Tracking human movements has been critical to the vast majority of research on motor control and neuromotor disorders. For example, motion capture has been pivotal in defining the fundamental biomechanics of human gait and movement coordination [1] as well as characterising fine motor impairments in Parkinson’s disease [2]. Nevertheless, motion tracking has often required specialised equipment and expertise to collect annotated movement data. Since the advent of optoelectronic systems that record reflective markers placed on body landmarks, optical tracking has become the gold standard for movement tracking due to its high precision [3–5]. However, the use of optical tracking requires specialised equipment and tends to be restricted to dedicated environments (research labs, production studios, etc.). Optical tracking entails further technical limitations, such as (relative) high cost and required technical skills for an effective setup. Furthermore, the use of markers represents a fundamental limitation as they require constant visibility and precise placement, both of which will affect the system’s accuracy and increase setup complexity. Other alternatives, such as inertial measurement units (IMUs) [6,7] or magnetic tracking [8] have been developed to address some of these issues, e.g., reducing cost and requirements for marker visibility. However, these approaches have not gained broader adoption, as their sensitivity to environmental interference (metals and magnets), tendency for error drift, and need for specialised expertise continue to pose barriers.

Markerless tracking (MLT), which is the use of plain video footage to identify and track key landmarks on the human body, offers an alternative to marker-based tracking [3,4,9]. Current progress in computer vision and machine learning [10,11] allows for full automation of the process without the need for manual video annotation, simplifying the processing and expanding its usage to people outside the field of computer vision. The process of MLT (illustrated in Fig 1) begins in two dimensions (2D) with an automated detection of specific body parts in individual images or videos (e.g., the position of a wrist, elbow, or fingertip). These landmarks are identified by computer vision models, which effectively act as “virtual markers” which are placed consistently across video frames in a process called 2D Annotation. Due to readily available large, annotated image datasets, and because the task is constrained to image coordinates, several pre-trained, off-the-shelf tools now exist and make automated landmark detection more accessible. For instance, MediaPipe Hand Landmarker [12], provides robust hand landmark detection across images, video, and live camera streams, while OpenPose [11] can detect body, face, and hand key points in real time, even in multi-person settings. These 2D tools eliminate the need for manual marking (such as in optical tracking) and make pose estimation available without the need for creating new models.

thumbnail
Fig 1.

A) 2D Hand Annotation based on Mediapipe Hand Landmarks. B) Example of a camera calibration process, required to obtain 3D information. C) Exemplary results of 3D data annotation. The displayed video frame has the hand landmarks labelled onto the original video. In the middle, there is a projection of the 3D acquired data of that frame. D) Extracting hand features such as angles and volumes.

https://doi.org/10.1371/journal.pone.0358647.g001

For many applications, however, 2D information is not enough. While 2D tracking can determine where a joint appears in an image or video, it cannot capture depth, orientation, or the spatial relationships between body parts. This data might be critical for analyses of movement kinematics, such as calculating joint angles, measuring stride length, or extracting other biomechanically meaningful features. Obtaining this information requires reconstructing three-dimensional (3D) poses from multiple camera views. Conceptually, this relies on 3D triangulation – the process of locating a point in space by combining its 2D positions from different camera viewpoints. For triangulation to correctly relate 2D pixel coordinates to 3D, the cameras themselves must be precisely calibrated – that is, their position, orientation, and lens distortions must be known. To obtain these parameters, typically, a reference checkerboard with known dimensions is recorded (illustrated in Fig 1) and then processed by a camera-calibration algorithm. Additional challenges for triangulation include ensuring that the same landmark is consistently identified across views (multi-view correspondence) and reducing noise in the reconstructed trajectories. Toolkits such as Anipose [13] and Pose3D [14] offer tools for 3D triangulation and camera calibration as an open-source solution. Anipose also provides video overlays of resulting virtual markers and can integrate with another existing tracker software called DeepLabCut [15], to enable tracking of external objects and certain animals.

All these tools (MediaPipe, OpenPose, Anipose, DeepLabCut) eliminate the manual annotation process, simplify camera calibration, and improve consistency and reproducibility. Additionally, recent advancements in MLT make the process simpler and more accurate each year. For example, the Radical Live service [16] removes the need for high-end computers due to its cloud service. New algorithms, such as ATHENA [17] use novel triangulation methods and mechanisms to prevent hand-switching errors (accidentally swapping which hand is being tracked), which bring performance closer to optical tracking.

Despite the barriers of specialised equipment and machine learning accuracy being lifted by various recent packages and toolkits [18–20], MLT is still not fully accessible for general use. The necessity for users to have coding knowledge, the ability to diagnose installation issues, and the ability to navigate through inconsistent standards, presents a technical barrier for wider MLT adoption. For comparison, Nexus (Vicon, Oxford, UK), a popular software for optical tracking, provides an intuitive graphical user interface (GUI) with ready-to-use features without the need for any coding or installation troubleshooting. Here we aim to close this gap by presenting TrackStudio – a toolkit that provides an accessible software tool for people wishing to employ MLT without requiring an extensive technical background. This community open-source resource also enables experts to expand the toolkit through integrating different methods for annotating 2D, 3D, or computing camera calibration for a variety of videos.

The TrackStudio toolkit (see S2 TrackStudio GitHub Files in S2 File) is a collection of custom-written code and open-source libraries, aimed at simplifying and optimising MLT. It is designed to provide easy access and adequate performance in generalised settings, which could be easily utilised by non-experts. In addition to MLT, the TrackStudio toolkit includes modules for video trimming, and visualisation of preprocessing steps, as well as 2D/3D video labelling to support accessible quality inspection. All these features are integrated into a Python GUI, allowing users to access and customise the different tools and their parameters without having to interact with or write any code. The only inputs the TrackStudio requires are video recording(s) using a simple web camera(s), and a recording of a calibration board if 3D MLT is required (instructions provided in the Appendix). To facilitate evaluation and demonstration of the toolkit, example videos are available on the TrackStudio GitHub repository (https://github.com/dimitrov-hristo/TrackStudio/tree/master/examples/videos), allowing users to explore its full functionality.

Materials and methods

Toolkit overview

The TrackStudio toolkit, including the utilities for video trimming, visualisation, configuration file editing, and tool integration was developed by the lead author (H. Dimitrov). The pose estimation models used for 2D/3D annotation, video labelling, and calibration were sourced from established open-source libraries.

2D video annotation is performed using Google’s MediaPipe framework [12]. This tool was chosen due to its robust and lightweight implementation, as well as the stable performance in a variety of settings (lights, background, occlusions, etc.) [21–23]. However, the standalone 3D model has limitations regarding depth estimation [23,24]. Therefore, Anipose’s approach [13] is utilised for 3D triangulation, video labelling, and integration capabilities with DeepLabCut (used for animal models and inclusion of custom markers [15]). By leveraging the strengths of these open-source libraries, we aim to extract optimal performance in a variety of conditions, while maintaining modularity for future expansion in the fast-moving MLT field.

One of the biggest challenges for using tools such as MediaPipe [12], OpenPose [11], FreeMoCap [25] or Anipose [13] is their maintenance and the continuously changing behaviour of the packages they rely on. This results in installation issues – e.g., incompatibility between package versions within and across different tools as well as issues due to operating systems and silicon. Thus, the approach we take is to package and upload the environments required for the pose estimation models and utilities, which prevents all of these issues. The disadvantage of this approach is that the user is limited to a particular version of these open-source tools. However, this approach guarantees functionality and maintains installation simplicity, especially since these tools are not updated frequently, but the packages that they use are, which can break the system’s utility.

Another important consideration in the MLT process is the acquisition of high-quality, synchronised recordings, as the MLT performance is ultimately dependent on the quality of the video data. Factors like occlusions (either self or external objects or barriers), light conditions, movement speed, lens and recording distortions, clothing, or calibration quality can significantly affect the performance of any markerless (and in many cases optical) tracking. At the same time, we aim to provide the easiest and least demanding setup, which can be used in diverse environments. Thus, we have created a brief manual (see S1 Appendix – TrackStudio Manual) aimed at illustrating how to install TrackStudio, record quality videos, perform MLT, and address challenges pertaining to common pitfalls of MLT. In addition, we provide experimental validation of the toolkit across 3 different setups, each introducing its own MLT challenges, as well as test videos on TrackStudio’s page (https://github.com/dimitrov-hristo/TrackStudio/tree/master/examples/videos).

Experimental validation

Validation environments

To test the TrackStudio toolkit in realistic and challenging settings, we recorded 76 participants (recruited in the period 29/03/2022–30/05/2025) in 3 different setups (seated, supine, and mixed), interacting with various objects (see Table 1). Ethical approval for the data collections was obtained from the Cambridge Psychology Research Ethics Committee for the data used seated and mixed protocols and from the University College London Research Ethics Committee (UCL REC) for the data used in the seated protocol. All participants provided written informed consent before taking part. These setups vary in number and type of cameras, tracked body parts, space, light conditions, background, body and object positions, and tasks. This is aimed to provide an evaluation of the performance of the toolkit in variable conditions and occlusion scenarios. The first two environments (seated and supine) involve recordings completed with simple web cameras (Logitech Brio, 1080p, 60 Hz) and evaluating the capabilities of a low-cost setup. The third environment involves high-resolution RGB cameras (FLIR Blackfly S BFS-U3-23S3C-C). It aims to evaluate the capabilities of the toolkit for tracking multiple body parts across different body positions (sitting and standing), in a confined space, with multiple tasks, obstructions, and multi-day recordings.

The seated environment tested the hand tracking performance across different tasks. The setup consisted of a table surrounded by three cameras, with participants sitting on a chair and performing object manipulation tasks (Fig 2A). The data from this setup represents 45 people (27 female, mean age = 25.2 ± 4.47). Each participant performed a set of three tasks involving challenging manipulations of everyday objects [26]. Furthermore, participants also completed the same tasks wearing an augmentation device (The Third Thumb, Dani Clode Designs) on the tracked hand used for object manipulation. This hand-worn device introduces further occlusions to the MLT setup. Each task was repeated multiple times, yielding 10 minutes of video per task and a total of approximately 30 minutes of automatically labelled, tracked videos per participant and condition (with and without wearing an augmentation device). The three cameras were synchronised by the timestamps of the recording computer.

thumbnail
Fig 2. Validation environments, illustrating the sitting setup with 3 web cameras (A), the lying down setup with 4 web cameras (B), and the mixed setup with 5 high-resolution cameras (C).

For the lying down setup, web cameras are circled in red and LED light in green for clarity.

https://doi.org/10.1371/journal.pone.0358647.g002

The supine environment tested hand tracking performance in a challenging body position in a constrained, low-light environment. The supine setup took place in a mock MRI room and consisted of four web cameras, with participants lying on the MRI bed being partially positioned inside the bore (only their head and neck). A height-adjustable table with an integrated LED light was mounted on the participants’ torso, serving as both a base for object placement and the designated start position of their hands. The LED light was used for camera synchronisation (Fig 2B). Participants interacted with 10 different daily-life objects in two distinct ways – performing actions away from the body and towards the body. This setup produced data from 26 participants (14 female, mean age = 54.7 ± 12.8), comprising a total of 10 minutes of auto-labelled tracked video data per participant.

The mixed environment tested hand, arms and face tracking across multiple days and with position changes (sitting and standing) within a session. The setup used five high-resolution RGB machine-vision cameras distributed around the testing space. A red LED light was mounted in the upper corner of each camera to ensure video synchronisation. Participants performed seven different object manipulation tasks while sitting at a table and repeated a subset of three of these tasks while standing, with the table height adjusted accordingly (Fig 2C). Here too, the hand augmentation device was worn for the majority (five) of the tasks during sitting and for all of the tasks during standing, imposing an additional occlusion challenge for tracking the biological fingers. This setup was tested on 5 participants (3 female, mean age = 25.7 ± 3.91). Each task was repeated multiple times, producing approximately 10 minutes of video per task. Hence, there was a total of 70 minutes of automatically labelled, tracked video in the sitting condition and 30 minutes for the standing condition per participant per session. In total, there were 5 sessions in 5 consecutive days.

Preliminary workflow usability evaluation

To assess the accessibility of the software for non-expert users, we conducted a brief structured usability evaluation with 10 participants reporting beginner-to-intermediate coding experience and no prior experience with the software. Before the session, participants were asked to download the software files and read the accompanying appendix manual. During the session, participants independently installed and configured the software using the example dataset, then completed the full processing workflow: automatic video trimming, camera calibration, 2D and 3D annotation, and 2D and 3D video labelling.

Task completion times were recorded throughout. To account for differences in computer speed, participant task time was calculated as total elapsed time minus processing time, allowing estimation of user-active time independent of hardware-dependent computation. Overall time from installation to completion of the final task is reported in the main text, with task-level timings provided in S1 Table.

Pipeline performance metrics

To assess the performance of the pipeline, we have focused on quality metrics for MLT (for more details on ground truth comparison, see Tony Hii et al., 2023 [22]). The quality metrics are: 1) inter-frame Pearson’s cross-correlation; 2) movement smoothness; 3) 3D error approximation. The inter-frame Pearson’s cross-correlation is at the level of the marker’s 3D vector magnitude and is performed between each 50 ms of data, discarding the change of direction. We use 50-ms intervals instead of consecutive frames and ignore correlations during changes of direction to avoid artificially high (minimal motion between frames) or low (natural movement reversal) values that do not reflect true tracking performance. The inter-frame cross-correlation per trial is computed by taking the median across all hand markers. Movement smoothness was quantified using the log dimensionless jerk (LDJ) metric. It represents the time- and amplitude-normalised integral of the squared jerk (the third derivative of position), providing a dimensionless measure of motion smoothness [27,28]. The LDJ per trial is computed by taking the median across all hand markers. As a guideline for interpreting LDJ, previous research indicates that simple actions/object manipulations result in LDJ in the range of 7–10 [29,30] and more complex movements/actions in the range of 11–14 [31]. Lastly, the 3D error approximation is representative of the amount of error incurred during the 3D triangulation. It is computed by calculating the error between the marker predicted by the 2D annotation and the 3D marker reprojected back into 2D space. To calculate this error in meaningful units (instead of camera pixels), we have converted pixels to mm by using the distance between the camera and the tracked body-part and the camera lens’ parameters. The error per trial is computed by taking the median across all hand markers. To contextualise acceptable magnitudes, state-of-the-art multiview triangulation methods typically achieve 13–34 mm mean per-joint position error on human benchmarks [32,33]. To capture further the number of large errors, we have calculated the percentage of errors that are above 10 mm (for hand and face) and above 30 mm (for elbows and shoulders), calculated per video frame and expressed as a percentage of the total number of frames. All of the MLT and metrics computation was performed on a standard laptop (Dell G5, Intel Core i5-8300H, Nvidia GeForce GTX 1060, 16GB RAM).

As TrackStudio combines Google’s MediaPipe for 2D annotations and Anipose for camera calibration and 3D reconstruction, prior research already exists on their accuracy compared to ground truth [13,21–24,34–38]. These studies provide a strong technical basis for the expected performance of the underlying methods. However, end-to-end tracking accuracy may still vary with camera configuration, calibration, movement characteristics, occlusion and environmental conditions. We therefore conducted a complementary target-referenced positional-accuracy assessment within the mixed environment recording configuration detailed above.

Participants performed a reaching task in which they sequentially touched eight fixed labelled targets (12 mm 20 mm) with their fingertips. The targets were distributed across four pylons, with each pylon containing one target positioned at a higher and one at a lower height. For each expected target contact, we calculated the distance between the reconstructed fingertip position and the independently specified location of the relevant physical target.

A target contact was classified as successful when a 5-mm-radius sphere centred on the reconstructed fingertip intersected the relevant target region. We report the target-acquisition rate and, for unsuccessful contacts only, the median fingertip-sphere-to-target gap in millimetres. This benchmark provides an external, target-referenced assessment of positional accuracy in the task and recording environment used in the present study.

Statistical analysis

All statistical analyses were conducted in JASP (v0.18.3) [39] using both frequentist and Bayesian approaches. For each setup, the three metrics (inter-frame cross-correlation, LDJ, and 3D error approximation) were reduced to a single median per subject per condition/task. Correlation coefficients were Fisher r-to-z transformed before inferential analyses. Descriptive statistics and figures are reported in the original correlation metric.

For the seated and supine setups, conditions were compared using paired t-tests (i.e., within-participant comparisons tested across participants) with accompanying Bayesian t-tests. For the Bayesian t-tests, the corresponding Bayes Factor (BF10), defined as the relative support for the alternative hypothesis, was reported. We used the threshold of BF10 < 1/3 as positive evidence in support of the null, consistent with previous research [40,41]. The Cauchy prior width was set at the conventional default of 0.707 [40,41]. Furthermore, agreement between the two fixed conditions (augmentation device worn vs not worn or away vs towards) was assessed using the two-way mixed-effects, absolute-agreement, single-measure Interclass Correlation Coefficient (ICC(3,1)), with ICC < 0.5 = poor, 0.5–0.75 = moderate, 0.75–0.9 = good, and > 0.9 = excellent reliability.

For the mixed setup, for each body part (hand, face, arms), data were examined for consistency across days and between augmentation device conditions (device worn vs not worn) when sitting using both frequentist and Bayesian two-way repeated-measures ANOVAs (factors: Day Device). If neither analysis indicated reliable effects or interactions, defined as non-significant frequentist results and BF10 < 1, data were considered stable and averaged across days and task types. This criterion was used solely to confirm measurement stability, as limited-sample Bayesian analyses often produce inconclusive Bayes Factors even when true differences are minimal [40,42,43]. Descriptive statistics (mean ± SD) of the aggregated data are reported thereafter.

Statistical assumptions were assessed via Q–Q plots and Shapiro–Wilk tests applied to paired differences (for t-tests) or residuals (for ANOVA). Sphericity was evaluated using Mauchly’s test, with Greenhouse-Geisser correction applied when violations were detected. Where distributional assumptions were materially violated, Wilcoxon signed-rank tests were applied, and z-statistic and Rank-Biserial Correlation were reported instead of t-score and Cohen’s d.

Results

To assess TrackStudio’s performance with low-cost webcams under different conditions, we examined inter-frame cross-correlation, movement smoothness (LDJ), and re-projection error (absolute and percentage >10 mm) in the seated (Fig 3) and supine (Fig 4) setups. Cross-correlations were extremely high in both configurations (seated: r = 0.98 ± 0.004; supine: r = 0.999 ± 0.0001), indicating smooth transitions and minimal marker jumps, e.g., due to occlusions. LDJ values followed previously reported trends from other studies [29–31]: simple movements yielded lower LDJs (7–10) [29,30] and more complex actions produced higher values (11–14) [31]. The supine setup, involving a single object manipulation movement (low complexity), showed an average LDJ of 8.36 ± 0.33, while the seated setup, which required continuous performance of daily manipulation tasks (higher complexity), registered 10.55 ± 0.36.

thumbnail
Fig 3. Seated setup analysis plots, illustrating inter-frame cross-correlation, log dimensionless jerk, and 3D error, with colours signifying different conditions.

https://doi.org/10.1371/journal.pone.0358647.g003

thumbnail
Fig 4. Supine setup analysis plots, illustrating inter-frame cross-correlation, log dimensionless jerk, and 3D error, with colours signifying different conditions.

https://doi.org/10.1371/journal.pone.0358647.g004

The 3D error (incurred during triangulation) remained small across participants (seated: 13.6 mm ± 10 mm; supine: 8.95 mm ± 4.1 mm), with very few large errors (seated: 0.58% ± 0.17%; supine: 0.55% ± 0.16%). These errors are consistent with state-of-the-art results on human benchmarks [32,33]. Furthermore, there was no significant effect on error from either the added obstruction in the seated setup (p = 0.37, z = 0.91, Rank-Biserial = 0.156, BF10 = 0.23, ICC = 0.67) or the task type in the supine setup (p = 0.094, z = –1.689, Rank-Biserial = –0.379, BF10 = 0.636, ICC = 0.95). Likewise, inter-frame cross-correlations did not differ between the two seated conditions (p = 0.08, t = –1.783, d = –0.266, BF10 = 0.69, ICC = 0.44) or the two supine tasks (p = 0.248, t = –1.183, d = –0.232, BF10 = 0.388, ICC = 0.56). As expected, LDJ differed significantly within both setups, reflecting the differing movement complexities (seated: p = 0.006, t = 2.886, d = 0.43, BF10 = 6.01, ICC = 0.49; supine: p < 0.001, t = –4.151, d = –0.814, BF10 = 90.3, ICC = 0.75).

To explore the benefits of higher-resolution cameras, we evaluated a mixed setup involving sitting and standing tasks, tracking the hand, face, and arms in a confined space over five consecutive days. Given the limited number of participants, these analyses were considered exploratory. No significant day-to-day or device-wearing differences were detected (0.93 > p > 0.15, 0.33 < BF10 < 0.848, see S2 Table for full details), suggesting no strong evidence for systematic changes across the recorded sessions. Therefore, Fig 5 shows participant-wise averages across days. Cross-correlations were again extremely high across all segments (0.999 ± 0.0001). Hand movements showed LDJ values (9.75 ± 0.62) similar to those in the seated webcam setup, reflecting comparable task structure. The 3D reprojection error for the hand was the lowest among all setups (4.41 mm ± 0.53 mm), with almost no large errors (0.0048% ± 0.0016%). Errors for face (9.52 mm ± 3 mm) and arms (17.6 mm ± 5.65 mm) were higher, likely reflecting the increased difficulty of tracking facial features and the larger anatomical variability of elbow and shoulder joint centres. Nevertheless, all reported errors remain within the range of values reported by current state-of-the-art approaches evaluated on human benchmarks [32,33].

thumbnail
Fig 5. Mixed setup analysis plots, illustrating inter-frame cross-correlation, log dimensionless jerk, and 3D error, with colours signifying different body parts and x-axis denoting different conditions.

https://doi.org/10.1371/journal.pone.0358647.g005

In the mixed-camera setup, we further assessed target-referenced positional accuracy during the separate target-reaching task. Across 730 reaches across the 5 participants and 5 days, the 5-mm-radius sphere centred on the index fingertip intersected the relevant target region in 91% of observations. For the remaining observations, the median gap between the fingertip sphere and the target boundary was 10.3 mm with interquartile range of 14.7 mm (raw data and summary results reported in S1 Validation Task Data).

Finally, all participants who took part in our preliminary assessment of the workflow usability completed successfully all the tasks from the first try. The average user-active completion time from installation to 3D video labelling was 04:44 minutes (range: 02:09 minutes-09:33 minutes), excluding computer processing time (for full details see S1 Table).

Discussion

To move MLT beyond expert use, we aimed to shift the focus from demonstrations of technical accuracy alone towards expectations of usability, repeatability, and adaptability. The TrackStudio toolkit presented here demonstrates that a streamlined, pre-configured workflow can deliver stable performance across different tasks and environments without the need for technical expertise. It is not intended to improve or replace established pose-estimation or triangulation algorithms such as MediaPipe, Anipose, or DeepLabCut. Rather, its contribution is practical and workflow-based: it integrates these tools within an accessible graphical interface, with a simplified installation process and constrained analysis workflow that guides users from video preparation to pose estimation and visualisation. This is important because, in practice, successful MLT use depends not only on algorithmic performance, but also on the user’s ability to assemble compatible tools, manage software environments, define configuration files, prepare input videos appropriately, and transfer data between different stages of the workflow. By reducing this dependency on technical expertise, TrackStudio enables non-technical users to obtain reproducible outputs that are comparable to those achievable through more technically demanding implementations.

This positioning is important when comparing TrackStudio with other commonly used markerless tracking workflows. Tools such as ATHENA [17], DeepLabCut [15], Bonsai [44], OpenCap [18], and FreeMoCap [25] have substantially expanded access to markerless tracking, each with clear strengths. DeepLabCut/Anipose and Bonsai-based workflows offer high flexibility and are particularly valuable for custom key points, multi-camera reconstruction, or real-time experimental control, but they still require technical skills and choices around installation, annotation, configuration, and/or pipeline assembly. ATHENA provides an important step toward highly accurate no-annotation, GUI-based hand tracking, but it is currently focused on hand kinematics and requires short command-line inputs for installation and launching. OpenCap and FreeMoCap similarly reduce barriers to full-body motion capture yet are primarily organised around their own capture and processing workflows and require technical skills for installation and usage. TrackStudio fills a different gap: it provides a pre-configured, GUI-based route from video preparation to tracking and visualisation, while remaining modular enough to incorporate outputs from tools such as DeepLabCut when specialist tracking is required.

Further strength of our pipeline lies not in outperforming task-specific implementations under ideal conditions, but in delivering dependable results across imperfect ones. Although direct comparison is limited, our target-referenced analysis within the mixed-camera setup, demonstrated broadly comparable results with prior markerless hand-tracking studies (reporting fingertip errors of approximately 11–21 mm) [45,46]. However, in many research contexts, the practical barrier is not whether a model can achieve a millimetre precision, but whether the pipeline can be deployed with limited setup time and budget, non-specialist staff, and variable recording environments. The current findings suggest that TrackStudio occupies this middle ground effectively: it enables reproducible tracking without requiring control over hardware, room layout, or extensive expertise.

More broadly, this work highlights the importance of reducing not only algorithmic complexity but also procedural friction. Many existing pipelines fail to translate beyond expert laboratories because the barrier lies in configuration, installation, or uncertainty over how to align components. By packaging a predefined environment with integrated documentation, informed default settings, and processes for recovery from common pitfalls, TrackStudio offers practical accessibility to support a broader user base, including those who may face challenges when working with command line interfaces.

At the same time, the toolkit does not preclude expert use and upgrades to its utility. Its modular organisation allows substitution of models for 2D/3D tracking, calibration methods, and extension to additional body regions or multi-person tracking. The preliminary tests of face and arm tracking illustrate this adaptability, even if performance varies with anatomical complexity and marker definition. These differences are expected and highlight where task-specific refinement may be needed if sub-centimetre precision is critical.

There are, of course, boundaries to what a general-purpose workflow can resolve. The present approach and validation were limited to single-person recordings in relatively structured settings. Out-of-scope scenarios, including multi-person contexts, close physical interaction, or crowded scenes, introduce additional challenges that go beyond the single-person tracking problem evaluated here. In these settings, tracking errors may arise from identity switching across frames, incorrect assignment of landmarks to individuals, inter-person occlusion, or ambiguity when overlapping body parts are detected. As TrackStudio was designed and validated for single-person recordings, its applicability to multi-person contexts should therefore be tested separately and may require additional identity-tracking, camera coverage, or manual correction procedures. In addition, highly dynamic full-body movements, severe or unpredictable occlusions, or applications requiring clinical-grade kinematic accuracy may still benefit from tailored optimisation, additional cameras, or marker-based systems.

The current evaluation also highlighted practical limits related to recording space and hardware. All environments involved a single individual within an appropriately sized capture volume. When the recording space was not appropriately sized, as in the mixed setup, higher-end cameras were required to improve the field of view and shutter speed. Increasing the number of cameras may further improve coverage, but also increases demands on video capture, data transfer, synchronisation, and visibility of synchronisation cues (hence we used an LED per camera in the mixed setup). Nonetheless, the aim of the present work was not to eliminate the need for specialised solutions, but to make MLT realistically accessible in scenarios where it is currently under-used.

In the near future, we plan to expand the TrackStudio toolkit to different operating systems (Linux-based and MacOS-based), add advanced feature extraction tools, expand the methods for video recording synchronisation, and include multi-person tracking approaches. With this, we hope to expand further the utility and usability of the toolkit as well as improve the camera synchronisation process.

Conclusion

By offering a pre-configured, GUI-based system, that supports diverse recording conditions without bespoke engineering, TrackStudio lowers the barrier to MLT adoption and facilitates experimentation rather than discouragement at the setup stage. In doing so, it also opens possibilities for educational activities, pilot work, and exploratory recording where the cost of failure is typically too high to justify time investment. TrackStudio does not represent an optimal system, but rather a practical one: a toolkit that lowers the barrier to entry while remaining adaptable.

Supporting information

S1 Appendix. TrackStudio Manual.

Installation and usage of TrackStudio alongside advice for markerless tracking setups.

https://doi.org/10.1371/journal.pone.0358647.s001

(PDF)

S1 Table. Completion times for each stage of the software workflow in 10 novice users.

Total elapsed time, computer processing time, and participant task time (total elapsed time minus processing time) are shown for installation, configuration, video processing, annotation, and labelling tasks.

https://doi.org/10.1371/journal.pone.0358647.s002

(XLSX)

S2 Table. Full statistical results for comparisons of tracking performance across testing days and device-wearing conditions.

https://doi.org/10.1371/journal.pone.0358647.s003

(XLSX)

S1 Validation Task Data. Raw data and summary results of the target-referenced positional-accuracy assessment.

https://doi.org/10.1371/journal.pone.0358647.s004

(ZIP)

S2 File. TrackStudio GitHub Files.

TrackStudio code files provided here in addition to being made available at a public GitHub repository.

https://doi.org/10.1371/journal.pone.0358647.s005

(ZIP)

Acknowledgments

We would like to thank Lucy Dowdall, Payton Kang, and Ema Jugovič for providing additional data for validation and Jonathan A. Michaels for contributing with guidance and ideas for setup improvements.

References

  1. 1. Winter DA. Biomechanics and motor control of human movement. John Wiley & Sons. 2009.
  2. 2. Roemmich RT, Field AM, Elrod JM, Stegemöller EL, Okun MS, Hass CJ. Interlimb coordination is impaired during walking in persons with Parkinson’s disease. Clin Biomech (Bristol). 2013;28(1):93–7. pmid:23062816
  3. 3. Scataglini S, Abts E, Van Bocxlaer C, Van den Bussche M, Meletani S, Truijen S. Accuracy, Validity, and Reliability of Markerless Camera-Based 3D Motion Capture Systems versus Marker-Based 3D Motion Capture Systems in Gait Analysis: A Systematic Review and Meta-Analysis. Sensors. 2024;24(11):3686.
  4. 4. Kanko RM, Laende EK, Davis EM, Selbie WS, Deluzio KJ. Concurrent assessment of gait kinematics using marker-based and markerless motion capture. J Biomech. 2021;127:110665. pmid:34380101
  5. 5. van der Kruk E, Reijne MM. Accuracy of human motion capture systems for sport applications; state-of-the-art review. Eur J Sport Sci. 2018;18(6):806–19. pmid:29741985
  6. 6. Filippeschi A, Schmitz N, Miezal M, Bleser G, Ruffaldi E, Stricker D. Survey of Motion Tracking Methods Based on Inertial Sensors: A Focus on Upper Limb Human Motion. Sensors (Basel). 2017;17(6):1257. pmid:28587178
  7. 7. García-de-Villa S, Casillas-Pérez D, Jiménez-Martín A, García-Domínguez JJ. Inertial Sensors for Human Motion Analysis: A Comprehensive Review. IEEE Trans Instrum Meas. 2023;72:1–39.
  8. 8. Franz AM, Haidegger T, Birkfellner W, Cleary K, Peters TM, Maier-Hein L. Electromagnetic tracking in medicine--a review of technology, validation, and applications. IEEE Trans Med Imaging. 2014;33(8):1702–25. pmid:24816547
  9. 9. Mathis A, Schneider S, Lauer J, Mathis MW. A Primer on Motion Capture with Deep Learning: Principles, Pitfalls, and Perspectives. Neuron. 2020;108(1):44–65. pmid:33058765
  10. 10. Wang J, Sun K, Cheng T, Jiang B, Deng C, Zhao Y, et al. Deep High-Resolution Representation Learning for Visual Recognition. IEEE Trans Pattern Anal Mach Intell. 2021;43(10):3349–64. pmid:32248092
  11. 11. Cao Z, Hidalgo G, Simon T, Wei S-E, Sheikh Y. OpenPose: Realtime Multi-Person 2D Pose Estimation Using Part Affinity Fields. IEEE Trans Pattern Anal Mach Intell. 2021;43(1):172–86. pmid:31331883
  12. 12. Zhang F, Bazarevsky V, Vakunov A, Tkachenka A, Sung G, Chang CL. MediaPipe Hands: On-device Real-time Hand Tracking. CoRR. 2020.
  13. 13. Karashchuk P, Rupp KL, Dickinson ES, Walling-Bell S, Sanders E, Azim E, et al. Anipose: A toolkit for robust markerless 3D pose estimation. Cell Rep. 2021;36(13):109730. pmid:34592148
  14. 14. Sheshadri S, Dann B, Hueser T, Scherberger H. 3D reconstruction toolbox for behavior tracked with multiple cameras. JOSS. 2020;5(45):1849.
  15. 15. Mathis A, Mamidanna P, Cury KM, Abe T, Murthy VN, Mathis MW, et al. DeepLabCut: markerless pose estimation of user-defined body parts with deep learning. Nat Neurosci. 2018;21(9):1281–9. pmid:30127430
  16. 16. Motion R. Real-time AI motion capture in your browser. https://radicalmotion.com/live 2025. Accessed 2025 October 15.
  17. 17. Mulla DM, Costantino M, Freud E, Michaels JA. ATHENA: automatically tracking hands expertly with no annotations. J Neurophysiol. 2025;134(6):2003–12. pmid:41269685
  18. 18. Uhlrich SD, Falisse A, Kidziński Ł, Muccini J, Ko M, Chaudhari AS, et al. OpenCap: Human movement dynamics from smartphone videos. PLoS Comput Biol. 2023;19(10):e1011462. pmid:37856442
  19. 19. TensorFlow. MoveNet: Ultra Fast and Accurate Pose Detection Model. https://www.tensorflow.org/hub/tutorials/movenet Accessed 2025 October 15.
  20. 20. Contributors M. OpenMMLab Pose Estimation Toolbox and Benchmark; 2020. Available from: https://github.com/open-mmlab/mmpose
  21. 21. Bittner M, Yang WT, Zhang X, Seth A, van Gemert J, van der Helm FCT. Towards Single Camera Human 3D-Kinematics. Sensors. 2023;23(1).
  22. 22. Tony Hii CS, Beng Gan K, You HW, Zainal N, Ibrahim NM, Azmin S, et al. Frontal Plane Gait Analysis using Pose Estimation Models. In: 2023 IEEE 2nd National Biomedical Engineering Conference (NBEC), 2023. 1–6. https://doi.org/10.1109/nbec58134.2023.10352623
  23. 23. Coster MD, Rushe E, Holmes R, Ventresque A, Dambre J. Towards the extraction of robust sign embeddings for low resource sign language recognition. arXiv. 2023. http://arxiv.org/abs/2306.17558
  24. 24. Amprimo G, Ferraris C, Masi G, Pettiti G, Priano L. GMH-D: Combining Google MediaPipe and RGB-Depth Cameras for Hand Motor Skills Remote Assessment. In: 2022 IEEE International Conference on Digital Health (ICDH), 2022. 132–41. https://doi.org/10.1109/icdh55609.2022.00029
  25. 25. Queen P, Cherian A, Wirth T, Idehen E, Matthis JS. freemocap. https://github.com/freemocap/freemocap 2024.
  26. 26. Dowdall L, Molina-Sanchez M, Dominijanni G, da Silva E, Pavalkyte V, Jugovic E. Developing a sensory representation of an artificial body part. bioRxiv. 2025.
  27. 27. Balasubramanian S, Melendez-Calderon A, Burdet E. A robust and sensitive metric for quantifying movement smoothness. IEEE Trans Biomed Eng. 2012;59(8):2126–36. pmid:22180502
  28. 28. Flash T, Hogan N. The coordination of arm movements: an experimentally confirmed mathematical model. J Neurosci. 1985;5(7):1688–703. pmid:4020415
  29. 29. Bayle N, Lempereur M, Hutin E, Motavasseli D, Remy-Neris O, Gracies JM, et al. Comparison of various smoothness metrics for upper limb movements in middle-aged healthy subjects. Sensors. 2023;23(3):1158.
  30. 30. Engdahl SM, Gates DH. Reliability of upper limb movement quality metrics during everyday tasks. Gait Posture. 2019;71:253–60. pmid:31096132
  31. 31. Gulde P, Hermsdörfer J. Smoothness Metrics in Complex Movement Tasks. Frontiers in Neurology. 2018;9.
  32. 32. Bartol K, Bojanić D, Petković T, Pribanić T. Generalizable human pose triangulation. 2022. https://arxiv.org/abs/2110.00280
  33. 33. Iskakov K, Burkov E, Lempitsky V, Malkov Y. Learnable Triangulation of Human Pose. In: International Conference on Computer Vision (ICCV); 2019.
  34. 34. Hii CST, Gan KB, Zainal N, Ibrahim NM, Azmin S, Desa SHM, et al. Automated Gait Analysis Based on a Marker-Free Pose Estimation Model. Sensors. 2023;23(14):6489.
  35. 35. Amprimo G, Masi G, Olmo G, Ferraris C. Deep Learning for hand tracking in Parkinson’s Disease video-based assessment: Current and future perspectives. Artif Intell Med. 2024;154:102914. pmid:38909431
  36. 36. Sprague AH, Vogel C, Williams M, Wolf E, Kamper D. Evaluation of commercial camera-based solutions for tracking hand kinematics. Sensors. 2025;25(18):5716.
  37. 37. Maggioni V, Azevedo-Coste C, Durand S, Bailly F. Optimisation and Comparison of Markerless and Marker-Based Motion Capture Methods for Hand and Finger Movement Analysis. Sensors (Basel). 2025;25(4):1079. pmid:40006308
  38. 38. Uribe M, Damiani A, Macellari N, Gonzalez I, Vega E, Grigsby EM, et al. Automatic Pipeline for Kinematic Tracking of Clinically Relevant Upper-Limb Motor Tasks in Humans and Macaque. Annu Int Conf IEEE Eng Med Biol Soc. 2025;2025:1–5. pmid:41336093
  39. 39. Team J. JASP (Version 0.18.3)[Computer software]. 2025.
  40. 40. Dienes Z. Using Bayes to get the most out of non-significant results. Frontiers in Psychology. 2014;5.
  41. 41. Wetzels R, Matzke D, Lee MD, Rouder JN, Iverson GJ, Wagenmakers E-J. Statistical Evidence in Experimental Psychology: An Empirical Comparison Using 855 t Tests. Perspect Psychol Sci. 2011;6(3):291–8. pmid:26168519
  42. 42. van Doorn J, Aust F, Haaf JM, Stefan AM, Wagenmakers E-J. Bayes Factors for Mixed Models: Perspective on Responses. Comput Brain Behav. 2023;6(1):127–39. pmid:36879767
  43. 43. Keysers C, Gazzola V, Wagenmakers E-J. Using Bayes factor hypothesis testing in neuroscience to establish evidence of absence. Nat Neurosci. 2020;23(7):788–99. pmid:32601411
  44. 44. Gonzalez M, Gradwell MA, Thackray JK, Temkar KK, Patel KR, Abraira VE. Using DeepLabCut-Live to probe state dependent neural circuits of behavior with closed-loop optogenetic stimulation. J Neurosci Methods. 2025;422:110495. pmid:40436321
  45. 45. Abdlkarim D, Di Luca M, Aves P, Maaroufi M, Yeo S-H, Miall RC, et al. A methodological framework to assess the accuracy of virtual reality hand-tracking systems: A case study with the Meta Quest 2. Behav Res Methods. 2024;56(2):1052–63. pmid:36781700
  46. 46. Reimer D, Podkosova I, Scherzer D, Kaufmann H. Evaluation and improvement of HMD-based and RGB-based hand tracking solutions in VR. Frontiers in Virtual Reality. 2023;4.