Figures
Abstract
Labor trafficking increasingly intersects with online environments, yet there are limited computational detection approaches. To examine existing methods and their potential applications, this article reviews literature that applies or advocates analytics and machine learning (ML) for detecting labor trafficking (LT). It synthesizes methodologies across two contexts: (a) internet-facilitated LT (IF-LT), where digital platforms enable recruitment and coordination of trafficking, and (b) online-accessible LT (OA-LT), where offline signals of trafficking can be retrieved from online sources. Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines, six bibliographic databases were searched between January 2002 and September 2025 and complemented with additional journals, backward and snowball searches, and expert recommendations. Of 2,151 records, we assessed 51 articles for full-text and 21 were mapped and qualitatively synthesized (14 LT-focused and 7 methodologically transferable). IF-LT research clusters around online recruitment, recruitment fraud, illicit massage businesses (IMBs), while OA-LT examines supply chains, child labor, violence indicators on social media, modern slavery disclosures and high-risk sectors such as the fishing industry. Common approaches include case studies with explainable analytics, interpretable classifiers, topic modeling and advanced natural language processing (NLP) methods like sentiment analysis, transformers and large language models; positive-unlabeled learning, probabilistic modeling, and spatial analysis. Despite attention to online LT, the review reveals a gap in quantitative applications of state-of-the-art ML to online content for LT detection. Interdisciplinary and transferable techniques developed for detecting online sex trafficking (ST) offer an adaptable pathway towards automation and law enforcement applications. Beyond methodological gaps, the review emphasizes responsible and context-aware AI, including weighted indicator modeling, remedies for class imbalance, multimodal data integration, label-efficient learning with expert annotations, and clear diagnostic reporting to support accountable deployment.
Citation: Pradhan R, Shehory O (2026) Online labor trafficking and machine learning: A scoping review and research agenda. PLoS One 21(9): e0358425. https://doi.org/10.1371/journal.pone.0358425
Editor: Jan Vrba, Akademia Jagiellońska w Toruniu: Akademia Jagiellonska, CZECHIA
Received: March 12, 2026; Accepted: September 1, 2026; Published: September 24, 2026
Copyright: © 2026 Pradhan, Shehory. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All relevant data are within the manuscript and its Supporting Information files.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.
Introduction
Labor trafficking (LT) is a worldwide violation of human rights that occurs both domestically and internationally. It involves various actors, such as criminal networks, businesses, employers, governments, consumers, civil society organizations, recruitment agencies, and vulnerable populations. The Protocol to Prevent, Suppress and Punish Trafficking in Persons, supplementing the United Nations Convention against Transnational Organized Crime (hereafter the Palermo Protocol), defines trafficking in persons as the recruitment, transportation, transfer, harboring or receipt of persons, achieved through the threat or use of force, coercion, abduction, fraud, deception, the abuse of power or of a position of vulnerability, or the giving or receiving of payments to obtain the consent of a person having control over another, for the purpose of exploitation [1]. Exploitation is defined to include, at a minimum, the exploitation of the prostitution of others or other forms of sexual exploitation, forced labor or services, slavery or practices similar to slavery, servitude, and the removal of organs. By convention, the trafficking in persons for sexual exploitation is termed sex trafficking (ST) and forced labor and services as labor trafficking (LT). Despite its prevalence, LT remains unevenly visible in research, policy, and technological interventions, particularly when compared with ST. This imbalance reflects not only differences in empirical accessibility, but also deeper asymmetries in how exploitation is conceptualized, operationalized, and detected, especially in digital environments.
This phenomenon is primarily observed in the unskilled labor sector, which is subject to less regulatory scrutiny and often experiences seasonal fluctuations in labor demand [2,3]. The Action-Means-Purpose (AMP) model illustrates the trafficking process and the route taken by a trafficker; at the minimum, the presence of one element from each column (Fig 1) points to a potential trafficking instance [4].
Produced by the authors based on the AMP model described by Polaris [4].
Recognizing the prevalence of LT, the United Nations (UN), International Labour Organization (ILO) and various institutions have produced comprehensive indicator lists of human trafficking (HT). Volodko et al. [5] operationalized these indicators for the identification of trafficking signals in job advertisements targeted at migrant job-seekers. The Global Report published by the UN emphasized the growing use of internet-based platforms for illicit activities, particularly in the context of recruitment, noting that almost half of the victims were affected by online recruitment among the set of court cases analyzed [6].
In the field of analytics and operations research, there is a significant focus on ST, despite LT being estimated to account for more than 60% of all trafficking instances. However, LT only constitutes 12% of the studies in a prior review [7]. This disparity is particularly salient in the application of artificial intelligence (AI) and machine learning (ML). One of the prominent computational applications addressing forced labor is satellite detection of brick kilns [8], whereas the methodological maturity, labeled datasets, and benchmark tasks are far more developed for ST than for LT [9,10].
Research that employs ML techniques specifically for online LT detection is even more limited [11,12]. This highlights the disproportionate representation of LT compared to ST, as indicators of sexual exploitation in online contexts are more explicit, while LT indicators are typically indirect, ambiguous and context dependent. Also, the broader HT literature often equates HT with ST, rather than distinguishing between trafficking types and their distinct exploitative purposes [13–16]. As a result, AI-based approaches risk inheriting and amplifying these conceptual biases. This, in turn, shapes which forms of exploitation become computationally legible and which remain underrepresented.
In response to this gap, this review consolidates existing efforts that use analytics and ML to detect LT across two contexts: internet-facilitated LT (IF-LT), where recruitment or coordination is done via digital platforms such as online recruitment, recruitment fraud, and illicit massage businesses (IMBs); and online-accessible LT (OA-LT), where relevant signals can be retrieved online although operations occur offline, such as in supply chains, child labor, violence indicators on social media, modern slavery disclosures, and high-risk sectors such as the fishing industry.
Importantly, the application of AI in these contexts is not merely a technical exercise, but a socio-technical intervention with potential consequences for workers, employers, and enforcement institutions. Misclassification, surveillance, and uneven data availability raise questions of accountability, transparency, and appropriate human oversight – issues that are central to contemporary debates on AI and society.
While prior reviews have examined human trafficking broadly, often focusing on sex trafficking, there is still no consolidated synthesis of machine learning applications specifically addressing labor trafficking across both internet-facilitated and online-accessible contexts. As a result, existing work remains fragmented, making it difficult to compare methods and identify common challenges in applying machine learning to labor trafficking detection.
To address this gap, this review maps and synthesizes existing methodologies and the context in which they are applied, and examines their implications for future AI-assisted labor trafficking detection. The review is guided by the following research questions:
- Mapping: What is the state of labor trafficking research that applies or advocates the use of machine learning or analytics across IF-LT and OA-LT contexts?
- Synthesis: What can be learned from the existing scientific evidence and countermeasures to improve online labor trafficking detection through responsible and context-aware AI methods?
Methods
In developing the review method, we were guided by the Arksey and O’Malley [17] and Joanna Briggs Institute (JBI) guidance [18], which include (1) identifying the research question, (2) identifying relevant literature, (3) screening relevant studies based on the inclusion criteria, (4) data charting, and (5) discussion of results. The reporting is guided by the Preferred Reporting Items for Systematic Reviews and Meta‑Analyses extension for Scoping Reviews (PRISMA-ScR) [19] ensuring documentation and reproducibility.
Prior to the review process, a protocol to guide the search, eligibility criteria and data-charting procedures was developed by the authors’ team to ensure methodological transparency because scoping reviews are not eligible for PROSPERO [20] registration. The Protocol summary is available in S2 File. In order to identify relevant literature we created a search and selection strategy which is described below:
Search strategy
We used the JBI Population-Concept-Context (PCC) framework (Table A in S1 File) to consolidate and operationalize the strategy for the search. The population details the trafficking phenomenon related to labor trafficking, modern slavery, and trafficking networks. The concept of interest is the analytical and machine learning techniques applied to trafficking detection or analysis of trafficking-related activities. The context encompasses the digital or online-accessible environments in which trafficking signals can be observed.
The search was systematically carried out for literature published between January 2002 and September 2025 with boolean, block-based search strings to ensure replicability (Table B in S1 File) across six databases namely SpringerLink, IEEE Xplore, Taylor & Francis, ACM Digital Library, ScienceDirect and Scopus. The detailed database search strategies and retrieved records are available in Table C in S1 File. The starting date (2002) was chosen to capture two decades of development in this field following the adoption of the Palermo protocol, which catalyzed attention and reporting on trafficking, leading to subsequent research. These databases contain the majority of the relevant literature, providing international and interdisciplinary coverage, where studies related to human trafficking are typically published. Secondly, we conducted manual searches across ArXiv and ResearchGate for conference articles to capture emerging literature prior to their formal publication. Third, we ran backward searches on reviews and snowball searches for mapped articles. We also included ‘labour trafficking’ in addition to ‘labor trafficking’ as specific industries used the term interchangeably in various academic and grey areas like webpages, policy proceedings and media etc. Search terms were selected to broadly reflect the HT literature, informed by the first author’s domain expertise, and finalized collaboratively by the author team.
Eligibility criteria
To achieve the aims of this review, we focused on the articles addressing LT in online contexts. The study includes the literature that apply or explicitly advocate ML or analytics, rather than discussing LT conceptually or theoretically. Online LT in this review refers to both IF-LT and OA-LT data that can be derived from the internet such as recruitment boards, job reviews, web content, open geodata or remotely sensed. The included studies had to be in English, published as peer-reviewed journal articles or conference papers, preprints, or doctoral dissertations. Grey literature in the included set comprises the preprints and dissertations. Additional grey literature, namely agency and institutional publications, supports definitional, contextual and indicator-framework claims and was not screened against these criteria. As the LT-specific ML literature is sparse, we also conducted targeted searches of online ST literature to extract transferable methodologies (as demonstrated, e.g., by [21], and cited by ILO [22]). These studies are also included in the data charting but reported in a separate column as methodological articles. Studies related to IMBs, which are commonly known for their association to sexually oriented services and may operate as a hybrid of LT and ST, were included under the LT category because of the involvement of these environments in deception and coercion in the recruitment and exploitation of victims. This is based on the Polaris findings from trafficking hotline massage business cases, in which survivors reported some form of labor trafficking in all cases.
Study selection
Records retrieved from database searches were imported into Rayyan systematic review software [23] for a two-stage screening (S2 File). This includes title and abstract screening followed by full-text eligibility assessment. Studies were evaluated against the inclusion criteria. Screening decisions were discussed between the authors for agreement on inclusion.
Data charting
A standardized data charting form (Table 1) was produced to extract key study characteristics such as research type, the data source, country represented in the study, techniques and methods, key findings and a separate column showing the category whether the articles directly refer to LT or methodological transferability.
Results
Study selection process
The study selection process is illustrated in the PRISMA flowchart (Fig 2). The full database yields were imported into the Rayyan systematic review software tool. Those that could not be imported were separately managed in Excel with standard bibliographic fields such as authors, title, year, source, abstract, digital object identifier. Following the database-specific search and export-stage filtering, 2,151 records were retained for deduplication and screening. After deduplication using the title string and author names, 2,119 unique articles remained and went through title and abstract screening. While transitioning to full-text assessment, we also identified 11 articles via snowball sampling, citation tracking, and expert recommendations, including some important references. In total, 51 studies were retained and proceeded to full-text eligibility check based on the inclusion criteria. At the full-text eligibility stage, studies were excluded if they were unrelated to labor trafficking contexts, did not involve data sources, did not apply analytics or machine learning or did not suggest a data-driven way forward. Because the LT-based data-driven literature is limited, seven additional studies were included for their methodological transferability (MT), taken from sex trafficking contexts which could inform labor trafficking detection. After full-text screening, we conducted final agreement on eligibility and methodological adequacy. This resulted in the final set of 21 articles: 14 LT-focused and 7 MT studies.
Mapping of included studies
Table 1 presents the mapping list of the 21 included studies as a data extraction table that records the article name and the characteristics such as the research type, the data source, country represented in the study, techniques and methods, key findings and a separate column showing the category if the articles directly refer to LT or methodological transferability.
Literature addressing IF-LT and OA-LT is concentrated in a few geographies, namely Lithuania, the Netherlands, the United States, the United Kingdom, and Brazil. Most datasets are country-specific and based on labor information. Studies draw on diverse online sources, including web-based recruitment data for unskilled labor, social media data such as Facebook and Twitter, Geographic Information System (GIS) data, business reviews, publicly accessible websites (e.g., Google Places), online news, official labor statistics, synthetic data, corporate disclosures, case articles, and interviews published online. There are also analyses of data scraped from websites that flagged trafficking risk. The majority of research employing automation techniques was conducted between 2018 and 2025, suggesting a growing interest in automation-based analytical approaches. This trend underscores both the growing need and the expanding scope for scalable methods, as well as the increasing relevance of research addressing the online dimensions of LT. Much of this research concentrates on online recruitment, particularly in job advertisements and forced labor in industries such as fisheries. A significant focus is placed on the illicit massage industry which is a hybrid form of LT alongside ST. Literature also mentions child labor and exploitation in supplier locations. The categorization of methodologies for detecting LT includes content analysis and exploratory studies, as well as predictive modeling, classification tasks and spatial analysis. Traditional ML techniques such as Naive Bayes, Logistic Regression, Support Vector Machines (SVM), and Random Forest models with decision trees are commonly used for classification, clustering and inferential modeling, with logistic regression also being combined with topic modeling in certain cases. Probabilistic approaches like Bayesian Networks and positive-unlabeled (PU) learning with ML classifiers are utilized for handling uncertainty. Natural Language Processing (NLP) and text analysis play a key role, employing advanced techniques like transformer models BERT, CNN, and LSTM, along with general deep learning models and large language models (LLMs). Sentiment analysis (e.g., VADER) further supports the detection of exploitation.
Overall, the evidence base is limited, with only a handful of analytical and ML intervention based studies. Qualitative designs in LT detection research such as field studies, policy analyses and interpretative studies outweighed quantitative ones. Limited data accessibility, with ST often treated as synonymous with HT, often obscures the studies that capture specific aspects of LT.
Discussion
This section first clarifies the relations between the concepts: forced labor, trafficking and slavery. It then discusses the result set theme-wise and concludes the review’s implication into future research.
Conceptual distinction in concepts related to labor trafficking
Researchers argue that LT is best understood as part of a continuum of exploitation lying between decent labor and forced labor [24–26] causing harm to workers and the reputation of industries. The literature reviewed in this article mentions the terms forced labor, human trafficking (also called “trafficking in persons”), slavery and modern slavery. As set out in the Introduction, HT constitutes a single offense under the Palermo Protocol, within which ST and LT are distinguished by the purpose element rather than by act or means [1]. Forced labor is work exacted under the menace of any penalty and not offered voluntarily, the two elements being cumulative [27]. Slavery denotes the status or condition of a person over whom powers attaching to the right of ownership are exercised. Modern slavery, by contrast, is an umbrella term used in advocacy and vary across jurisdictions whose scope has been seen in the literature [25,26]. These concepts overlap substantially in practice (Fig 3). Trafficking as a mechanism, with progression, can lead to forced labor and, in the extreme, can escalate to a state of slavery. Forced labor and slavery are defined by the condition or status of the person, whereas trafficking is defined by the process leading into that condition. Hybrid cases, of which IMBs are the clearest instance, involve labor exploitation and sexual exploitation of the same workforce and therefore satisfy more than one enumerated purpose simultaneously.
(Produced by the authors based on the definitions from ILO, “What is forced labour?”).
Methodological implications
The themes across the literature pertaining to LT applying or advocating ML techniques include job advertisements for labor recruitment leading to trafficking and job advertisements for scam, IMBs, modern slavery risks in supply chains, child labor, violence on social media, Modern Slavery Act disclosure analysis and articles annotation; and one of the high-risk sectors such as the fishing industry alongside methodologically transferable insights from escort ads in ST detection literature. These themes and their subproblems are represented in the taxonomy diagram (Fig 4).
The taxonomy situates the data sources and their analytical contexts across the reviewed studies for LT, the methodologically transferable ST domain, and the hybrid IMBs context. Apart from the data charting that maps the literature reviewed here, the discussion offers a synthesis chart of transferability and design proposition (Table 2) for developing systems to detect online LT complementing the narrative that follows.
With globalization [28] and the large-scale proliferation of affordable internet facilities, the transnational HT market has expanded beyond borders, within the illicit and licit sectors [13]. In terms of the AMP model (Fig 1), exploitation commonly uses deception at recruitment [3], such as promising wages higher than the market norms [29]. Recognizing this, across the recruitment domain, considerable research starting with an early foundational work [9], shaped global policies and technological interventions in this domain. The literature suggests that online job ads contain actionable signals such as violations of the national wage-and-hour standards, vague job role descriptions, off platform redirects, benefits such as accommodation, transport to work, transfer to destination country and help with settling although not necessarily indicate trafficking but can be present to mask the exploitative conditions [5,30]. Upon analysis it was observed that indicators usually co-occur rather than being present in isolation and only 10 out of 59 indicators laid out by the UNODC [31] could be operationalized, highlighting the challenges of using indicators developed for offline use in the online sphere. Community-specific platforms such as for the Chinese-speaking immigrants to the US refer to cultural and linguistic patterns in job advertisements. This is essential in mitigating bias in online HT detection [32] making the model more interpretable with preemptive measures [33]. Related work on recruitment fraud [34] shows overlapping tactics such as mass recruitment, offer of attractive packages, and typical writing style. These works suggest exploring topic modeling to extract thematic information and leveraging LLMs for analysis and reporting, as traditional ML models require extensive training data, LLMs may be helpful to outperform them in classification tasks [35]. While these studies demonstrate the potential of ML for recruitment-stage labor trafficking detection, they are largely limited to single platforms, specific linguistic, cultural context, and indicators characterizing a particular set of industries [5,30,33,34]. The most persisting gap is the absence of ground truth validation against confirmed cases and limited reporting of evaluation metrics. At the dataset level, some advertisements with too short texts were removed for reliable topic extraction [30], potentially biasing the dataset and results. In addition, heavily imbalanced datasets in this domain require resampling techniques while mitigating the bias.
IMBs exemplify a distinct hybrid combination and facilitation of both LT and ST [36]. These businesses advertise massage or therapeutic services disguised as legitimate businesses, but exploit workers through low wages and unlawful working hours, confiscating passports, sometimes coupling these services with sexual exploitation. Interpretable models with ensemble methods and features enriched with review text, licensing status, business attributes such as opening hours, recruiter metadata and geo-demographic indicators can surface risk while preserving transparency for decision-makers [37,38]. Spatial correlates such as proximity to airports and ports, neighborhood type, and local cost structures helped forecast relocation after raids and shutdowns [39]. Active learning pipelines for labeling NLP tasks in the absence of labeled data and alternative sources such as the Google Places API extend coverage [40,41]. In parallel, triangulation frameworks from computational criminology show how linguistic markers derived from embeddings, context-aware lexicon and gender guesser can be combined with geospatial layers to produce more robust, multi-dimensional risk assessments [42]. More recently, graph based ML also has been explored to model relationships across the complex information in advertisements and business entities to improve detection [43]. The IMB studies are predominantly US-based and trained on data from specific states such as Texas or Florida limiting both geographic generalizability and reproducibility of findings. The interpretability of models specifically with risk scores and decision trees trades off the accuracy because they use simpler models in order to achieve the transparency. Such a methodological gap can be addressed with explainable AI techniques.
The recruitment into massage industries for labor exploitation that occurs through deceptive online job advertisements, places IMBs within the internet-facilitated category (IF-LT) [36]. However, the detection methodologies applied by researchers rely predominantly on post-hoc online-accessible traces such as customer reviews, business location data and operating hours [37–41,43], which are characteristic of OA-LT detection contexts. The IF-LT and OA-LT categories therefore capture two distinct analytical aspects. One, through which trafficking is facilitated and the other, digital traces available for detection. This explains how to design the computational detection in either case. Detection systems targeting IMBs may therefore benefit from integrating both recruitment-stage signals alongside post-hoc consumer signals for a broader detection approach.
Within the supply chain contexts, important multi-source monitoring of information comes from social media, news, articles, audit data on child labor and other broader modern slavery risks. These are based on the demand-side evidence rather than recruitment or source-side detection which should be more preventive [44,45]. In this scenario, the UK Modern Slavery Act is a groundbreaking legislative approach to combating HT within corporate supply chains by requiring companies to disclose their anti-trafficking efforts with its “reporting mandate”. Analysis of these statements emphasize high-level policies rather than just the operational measures to detection [46,47]. Although modern slavery monitors the disclosure, it cannot verify the actual compliance that limits their utility in mitigating exploitation. Predictive modeling with lasso regularization on highly imbalanced data from socioeconomic, demographic, and rescue operations to identify modern slavery risk in Brazil can be further extended to multiclass models in urban and rural settings. Such data can be integrated to the information accessed online which are highly imbalanced to draw better inference on risk.
Recent work provides evidence for the applicability of LLMs to assist human reviewers with the system providing sentence level decisions, token level rationales and calibrated probabilities in assessing model’s confidence in its predictions. The AIMSCheck system [35], shows how LLMs can move beyond the binary classification towards rational clarity, guiding human reviewers through lengthy modern slavery compliance documents, and supports more informed decision-making. It is also seen that Fine-tuned models outperformed zero-shot and few-shot approaches in the compliance classification task [35,48], suggesting that models should be context-aware for reliable performance. Together, these developments imply accountable AI deployment in labor trafficking mitigation.
Often the lack of ground truth constrains fully supervised systems. To address this, RaFoLa [49] offers a rationale-annotated corpus for detecting indicators of forced labor drawn from news articles from traffick analysis hub, the Business and Human Rights resource center, and ILO newsroom for annotation which is rationale-oriented and derived mainly from the indicators list from the ILO. While this corpus reflects the demand-side evidence which is the public reporting of the incident, rather than the preventive recruitment-stage detection, it can be extended by developing sentence-level justification, improving label efficiency and supporting multi-label classification of co-occurring indicators in the online recruitment context.
Sector-specific analysis of a prominent hotspot of forced labor and trafficking is on the fishing industry which notes that vessels involved with forced labor exhibit systematically different behavior from other vessels. Combined together with vessel telemetry, technical characteristics, PU signs of forced labor could be predicted [50]. While this analysis does not include economic, social, or political causes of forced labor, data integration with other labor dynamics such as point of entry of worker, recruitment characteristics through recruitment agencies both online or offline recruitment, living conditions of workers, point of exit for the worker along with the restrictions imposed in leaving the workplace such as debt bondage or unemployment [51] would allow a comprehensive assessment of forced labor at sea.
Additionally, methodologies from ST domain, such as the semi-supervised model conducted by [21] offers adaptability for LT detection as recommended by ILO [22]. It is based on language patterns, words-of-interest, and victim characteristics. Similarly [52], develops a methodology for interpreting language in customer-to-customer (C2C) marketplace ads with pseudo-labels requiring minimal human intervention for training and evaluating NLP models. Lightweight, efficient text embeddings such as the “bag-of-tricks” [53] approach with heuristic relabeling have also shown advantages for large-scale screening compared with heavier paragraph embedding baselines in similar domains [54] and can also be translated for LT scenarios. It is a good example of a system that can operate in limited label settings by careful combination with proven ML techniques. Collectively, these studies suggest a practical direction as outlined in the way forward section.
Across the studies, methodological robustness varies considerably. The recruitment-stage studies present foundational but exploratory work with relatively smaller datasets and evaluation reporting. These studies showcase a methodological progression from indicator counting [5] to unsupervised ML [30] to computational analysis [37] and supervised classification [34]. Trafficking detection studies across IMBs are more developed with applied [37,40,41], interpretable [38] ML, spatial analysis [39], qualitative triangulation [42] and graph ML [43]. These report quantitative metrics across larger datasets containing reviews and advertisements. For example, studies on IMBs employ structured review datasets combined with census and geospatial data. This enables the reporting of precision–recall trade-offs and risk scores, in contrast to recruitment-stage studies where evaluation is often limited or absent. Supply chain studies [44] mostly rely on demand-side evidence such as rescue operation data [46], proprietary modern slavery statements [47], with LLM approaches as the most advanced design [35]. The fishing industry study addresses a distinct high-risk sector but remains constrained by label scarcity [50]. As an example, the fisheries study applies positive-unlabeled learning to vessel telemetry data to address the label scarcity. This is distinct from IMBs or supply chain studies where proxy labels from customer reviews and corporate disclosures provide a structure for model training. The geographic and methodological coverage across these studies are diverse but limits cross-country comparison due to the different application areas. Methodologically transferable studies [21,52,54] offer adaptable approaches, though validation on labor-specific datasets is needed. Together these patterns point to the need for standardized evaluation frameworks and shared benchmark datasets to advance methodological maturity in this field.
Structural constraints
The methodological limitations that constrain the field's maturity are best understood not as incidental gaps that a better model would resolve, but as structural constraints around data, methodological fragmentation, and context-bound transferability, which reinforce one another. The most fundamental is the data problem, which is the absence of a stable computational target. Because trafficking is hidden and legally contested, researchers rarely have confirmed cases and instead rely on proxies, but are not ground truth instances itself, so a model's reported accuracy reflects agreement with the proxy rather than detection of trafficking itself. Constitutive elements of offense such as coercion, debt bondage and document retention are realized inside the employment relationship, after recruitment, so recruitment-stage data cannot contain them, and deception requires the advertisement to appear lawful. The weak discriminative signal is therefore a property of the offense rather than a deficiency of the data. This also partially restricts transfer of trained signal from sex trafficking, where the advertisement is the point of sale and the artifact analyzed carries the transaction itself; in labor recruitment it describes conditions that do not yet exist, which is why the strongest available signals are economic irregularities within lawful documents rather than evidence of an offense. Fragmentation follows whichever records an institution happens to generate. Because these records are created for different purposes, they encode different phenomena, whether consumer perceptions, corporate compliance reports, enforcement outcomes, or behavioral anomalies. Consequently, datasets capture different constructs rather than different observations of the same phenomenon. Even when studies operate within the same domain, their findings cannot be readily synthesized because they measure fundamentally different aspects of exploitation. Hence, successive studies cannot readily build on one another. Together, these structural constraints determine what a detection system can and cannot see, and in turn who is scrutinized and where intervention is directed.
Socio-technical implications and governance considerations
The application of AI and machine learning to online LT detection should be understood not just as a technical enhancement of investigative capacity, but as a socio-technical intervention that reshapes how exploitation is identified and acted upon [55]. Given the ambiguous nature of the online trafficking indicators, models risk identifying proxies for exploitation rather than the actual instances. Algorithmic systems inevitably reflect the assumptions embedded in their data sources, labels, and modeling choices, thus influencing which forms of labor exploitation become visible to institutions and which remain obscured.
A central implication concerns computational visibility. As this review indicates AI-based approaches to LT rely heavily on indirect and context-aware indicators, often derived from fragmented or unevenly available online data. For example, studies focusing on IMBs [37–39] rely only on the customer reviews and make exploitation visible only when digital traces exist. These constraints risk privileging certain regions, sectors (e.g., construction and domestic work), or forms of exploitation that are more observable online. Consequently, AI systems may reinforce existing blind spots in LT research and enforcement, rather than correct them.
Another implication relates to risk and accountability [56]. Misclassification in LT detection has non-trivial consequences: false positives may stigmatize legitimate businesses or workers, while false negatives may leave exploitative practices unaddressed. Unlike many commercial AI applications, errors in this domain intersect directly with human rights and legal processes. These stakes emphasize the need for transparency in model design, clear reporting of uncertainty, and mechanisms for review by human decision-makers.
The reviewed literature also raises concerns regarding surveillance and power asymmetries. Many AI-driven approaches depend on large-scale monitoring of online content, supply-chain disclosures, or social media activity including vessel telemetry [50], corporate disclosure analysis [47], and social media scraping [45]. While such data may be publicly accessible, their aggregation and algorithmic analysis can expose already vulnerable populations such as migrant or informal workers, while affording limited visibility into the practices of more powerful actors. This asymmetry highlights the importance of carefully delimiting the scope of automated monitoring and embedding AI tools within governance frameworks that protect individual rights.
Finally, these findings point to the necessity of human-in-the-loop and institutional safeguards. AI systems for LT detection should be designed to support, rather than replace, expert judgment by investigators, regulators, and civil society organizations. Incorporation of domain expertise during model development, validation, and deployment can help mitigate bias, improve contextual sensitivity, and enhance trust. Moreover, decisions about which tasks to automate and which to leave to human discretion, should be informed by legal, ethical, and organizational considerations, not solely by technical feasibility. AIMSCheck [35] exemplifies this through sentence-level justifications without replacing human judgment.
Together, these socio-technical considerations suggest that progress in AI-enabled LT detection depends as much on governance, accountability, and institutional design as on advances in data and algorithms. Responsible deployment requires viewing AI not as an autonomous solution, but as a component within a broader ecosystem of social, legal, and policy interventions. As a scoping review, this study does not evaluate or benchmark the empirical performance of individual models, but instead synthesizes methodological approaches, data sources, and conceptual assumptions across the existing literature.
Way forward
Building on the theme-based synthesis in the preceding section, this review identifies methodologies and the transferability for advancing research that can detect LT at scale. Importantly, the objective of such systems is not to ascertain individual cases of LT, but to flag suspicious instances for closer human examination. As the discussion notes, indicators play an important role in shaping detection systems; they should be modeled not as simple binary tags but as a combination of weak and strong indicators. This can improve discriminatory power by weighting features on predictive strength rather than treating them as binary.
Beyond the existing literature, future work should focus on the integration of heterogeneous data and the systematic combination of automation with expert input. One possible approach is the use of active learning and expert labeling of borderline ads coupled with feature importance. This can refine indicator weights efficiently while penalizing noisy indicators, and can be further supported by weak supervision frameworks to auto-generate probabilistic labels in label-scarce settings. Wage and hour information, which is an empirically strong indicator, should be integrated with contextual features to generalize the ability of a model. Lexicons are valuable for precision and embeddings for recall, combining them in ensembles can balance performance when expanded with domain-specific vocabulary to strengthen coverage. Likewise, unstructured corpora benefit from topic discovery so that the extracted features can be used alongside the supervised classifiers for potential exploitation detection.
As the majority of ads are non-exploitative, two-stage models are appropriate to handle zero-inflated data distributions. In scenarios where features are numerous relative to confirmed positive cases, regularization and class imbalance should be treated with a combination of data-level methods (random under-sampling and over-sampling) and synthetic variants such as Synthetic Minority Over-sampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN) and algorithm-level approaches (e.g., cost-sensitive loss functions).
While detection tools can aid law enforcement, their deployment may also introduce risks, such as the over targeting of migrant workers or excessive surveillance of vulnerable populations. Such tools and systems should serve as a decision-support rather than autonomous detection. Accordingly, the evaluation of such tools should extend beyond technical performance metrics to include potential harms and unintended social consequences for affected populations [57]. Taken together, the mapping and synthesis of the literature highlights the need to integrate additional data sources, develop adaptable lexicons, operationalize the knowledge for decision making, reinforce active learning, and empirically test the indicators alongside reporting practices. In particular, transparent reporting of precision, recall, specificity, sensitivity and calibration together with false-positive consequences is essential to support responsible application in real-world settings [55].
Moving forward, our forthcoming research aims to translate some of these approaches into a unified, context-aware pipeline that bridges data, methodology and detection systems for online LT detection, while maintaining explicit attention to accountability, interpretability, and human oversight in deployment contexts.
Limitations
This scoping review is subject to limitations and sources of potential selection biases across its scope, evidence-base and source-composition. Firstly, the search process was restricted to English-language publications, potentially excluding relevant studies in other languages with significant data-driven research in non-English speaking regions. Secondly, despite searching six databases, the review may have omitted some relevant qualitative or conceptual studies across grey literature, and reports addressing labor trafficking through other methodological approaches. Thirdly, consistent with scoping review guidelines, this review does not perform a formal quality appraisal or comparative evaluation of model performance across the identified data sources.
Additionally, across the studies there is a heterogeneity in reporting of the study application success. While some provide evaluation metrics others provide qualitative findings, making relative comparison of effectiveness difficult across methods. Accordingly, the synthesis is exploratory in nature, mapping the methodologies in the existing research rather than ranking them. Finally, the relatively limited volume of data-driven research in this domain also influenced the representation and scope of the reviewed literature. For example, the scope of the regions mentioned in the included studies particularly United States, United Kingdom, Lithuania, Netherlands, Australia, Canada, Britain and Brazil with dominance of US-based studies, shows a geographic skew limiting the data representations and hence context-based applicability of findings across geographies and contexts.
Beyond the scope, a constraint in this field of study is the accessibility of datasets across the included studies, as they considerably vary in terms of accessibility. Some studies draw on fully publicly accessible sources such as Yelp reviews [37], job advertisements [5,30], open government data and news articles [49], where independent replication is possible. Some datasets are owned by organizations, such as the compiled Modern Slavery Act statement databases [47], and may be partially accessible upon request from the authors. Others are restricted to law enforcement or government agencies, including rescue operations data [46], and cannot be accessed by independent researchers. Studies examining IMBs scraped data from RubMaps platform have migrated domains following law enforcement pressure [37–41,43] effectively inaccessible for replication from the same source. A similar constraint applies to methodologically transferable studies where platforms such as Backpage.com have been shut down [21] and certain datasets remain law enforcement restricted [54]. This inaccessibility constrains reproducibility of the studies and limits improvement of the models. This can be mitigated by the creation of ethically curated, anonymized data accessible upon request to researchers.
Finally, two features of the evidence base introduce considerations for interpretation of the findings. The inclusion of grey literature such as preprints and dissertations broadens coverage and offers valuable methodological insight, but this work has not undergone peer review, so its findings are more indicative than confirmatory. Similarly, the sex-trafficking studies were included in the discussion for their methodological transferability rather than their subject matter; even though they inform a useful range of tested methods, the conclusions that rest on them remain conditional on the transferability limits discussed earlier. Both choices reflect the sparse LT-specific literature and were necessary for adequate coverage to further the research in this domain. They widen the evidence base at some cost to homogeneity and should be weighed accordingly.
Supporting information
S1 File. PCC Framework, Search query, database specific search strategy, query details.
https://doi.org/10.1371/journal.pone.0358425.s001
(DOCX)
References
- 1.
United Nations. Protocol to prevent, suppress and punish trafficking in persons, especially women and children, supplementing the United Nations Convention against Transnational Organized Crime [Internet]. New York: United Nations; 2000 [cited 2026 Apr 17]. Available from: https://www.ohchr.org/en/instruments-mechanisms/instruments/protocol-prevent-suppress-and-punish-trafficking-persons
- 2.
Andrees B. Forced labour and human trafficking: A handbook for labour inspectors [Internet]. Geneva: International Labour Office; 2008 [cited 2026 Apr 17]. Available from: https://www.ilo.org/sites/default/files/wcmsp5/groups/public/@ed_norm/@declaration/documents/publication/wcms_097835.pdf
- 3. Europol. European Union serious and organised crime threat assessment (SOCTA) 2017: Crime in the age of technology [Internet]. The Hague: Europol; 2017 [cited 2026 Apr 17]. Available from: https://www.europol.europa.eu/cms/sites/default/files/documents/report_socta2017_1.pdf
- 4. Polaris Project. Understanding human trafficking [Internet]. Washington (DC): Polaris; 2019 Oct 16 [cited 2026 Jul 16]. Available from: https://polarisproject.org/understanding-human-trafficking/
- 5. Volodko A, Cockbain E, Kleinberg B. “Spotting the signs” of trafficking recruitment online: exploring the characteristics of advertisements targeted at migrant job-seekers. Trends Organ Crim. 2019;23(1):7–35.
- 6. United Nations Office on Drugs and Crime. Global report on trafficking in persons 2020 [Internet]. Vienna: UNODC; 2021 [cited 2026 Apr 17]. Available from: https://www.unodc.org/documents/data-and-analysis/tip/2021/GLOTiP_2020_15jan_web.pdf
- 7. Dimas GL, Konrad RA, Lee Maass K, Trapp AC. Operations research and analytics to combat human trafficking: A systematic review of academic literature. PLoS One. 2022;17(8):e0273708. pmid:36037198
- 8. Boyd DS, Jackson B, Wardlaw J, Foody GM, Marsh S, Bales K. Slavery from Space: Demonstrating the role for satellite remote sensing to inform evidence-based action related to UN SDG number 8. ISPRS Journal of Photogrammetry and Remote Sensing. 2018;142:380–8.
- 9. Latonero M. Human trafficking online: The role of social networking sites and online classifieds [Internet]. SSRN; 2011 [cited 2026 Jul 16]. Available from: https://papers.ssrn.com/sol3/papers.cfm?Abstract_id=2045851
- 10. Latonero M, Wex B, Dank M. Technology and labor trafficking in a network society: General overview, emerging innovations, and Philippines case study [Internet]. SSRN; 2015 [cited 2026 Jul 16]. Available from: https://papers.ssrn.com/sol3/papers.cfm?Abstract_id=2574676
- 11. Sweileh WM. Research trends on human trafficking: a bibliometric analysis using Scopus database. Global Health. 2018;14(1):106. pmid:30409223
- 12.
Xian LY, Logeswaran R. Human Trafficking Through Data Sharing and Analytics. In: 2022 IEEE International Conference on Distributed Computing and Electrical Circuits and Electronics (ICDCECE), 2022. 1–4. https://doi.org/10.1109/icdcece53908.2022.9793090
- 13. Cockbain E, Bowers K, Vernon L. Using Law Enforcement Data in Trafficking Research. The Palgrave International Handbook of Human Trafficking. Springer International Publishing. 2019. p. 1709–32.
- 14. Strauss K. Sorting victims from workers: Forced labour, trafficking, and the process of jurisdiction. Prog Hum Geogr. 2017;41: 140–58.
- 15. Efrat A. Global efforts against human trafficking: The misguided conflation of sex, labor, and organ trafficking. Int Stud Perspect. 2016;17: 34–54.
- 16. Gezinski LB, Gonzalez-Pons KM. Sex trafficking and technology: A systematic review of recruitment and exploitation. J Hum Traffick. 2024;10: 497–511.
- 17. Arksey H, O’Malley L. Scoping studies: Towards a methodological framework. Int J Soc Res Methodol. 2005;8:19–32.
- 18. Peters MDJ, Godfrey C, McInerney P, Munn Z, Tricco AC, Khalil H. Chapter 11: Scoping reviews. In: Aromataris E, Munn Z. JBI manual for evidence synthesis. Adelaide (Australia): Joanna Briggs Institute; 2020.
- 19. Tricco AC, Lillie E, Zarin W, O’Brien KK, Colquhoun H, Levac D, et al. PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and Explanation. Ann Intern Med. 2018;169(7):467–73. pmid:30178033
- 20. Booth A, Clarke M, Dooley G, Ghersi D, Moher D, Petticrew M, et al. The nuts and bolts of PROSPERO: an international prospective register of systematic reviews. Syst Rev. 2012;1:2. pmid:22587842
- 21. Alvari H, Shakarian P, Snyder JEK. Semi-supervised learning for detecting human trafficking. Secur Inform. 2017;6:1.
- 22. International Labour Organization. Use of digital technology in the recruitment of migrant workers [Internet]. Geneva: ILO; 2021 [cited 2026 Apr 17]. Available from: https://www.ilo.org/publications/use-digital-technology-recruitment-migrant-workers
- 23. Ouzzani M, Hammady H, Fedorowicz Z, Elmagarmid A. Rayyan-a web and mobile app for systematic reviews. Syst Rev. 2016;5(1):210. pmid:27919275
- 24.
Laczko F, Gozdziak EM. Data and research on human trafficking: a global survey. Geneva: International Organization for Migration; 2005.
- 25. O'Connell Davidson J. Modern slavery: The margins of freedom. London: Palgrave Macmillan; 2015.
- 26. Weitzer R. Human trafficking and contemporary slavery. Annu Rev Sociol. 2015;41: 223–42.
- 27.
International Labour Organization. What is forced labour? [Internet]. Geneva: ILO; [cited 2026 Apr 17]. Available from: https://www.ilo.org/topics/forced-labour-modern-slavery-and-trafficking-persons/what-forced-labour
- 28. Aronowitz AA, Veldhuizen ME. The human trafficking–organized crime nexus. The Routledge Handbook of Transnational Organized Crime. London: Routledge. 2021. p. 232–52.
- 29.
Janušauskienė D. Lithuanian migrants as victims of human trafficking for forced labour and labour exploitation abroad. Exploitation of migrant workers in Finland, Sweden, Estonia and Lithuania: uncovering the links between recruitment, irregular employment practices and labour trafficking. Helsinki (Finland): HEUNI; 2013. pp. 305–52. (HEUNI Report Series; No. 75).
- 30. Cascavilla G, Catolino G, Palomba F, Andreou AS, Tamburri DA, Van Den Heuvel W-J. Unsupervised labor intelligence systems: A detection approach and its evaluation: A case study in the Netherlands. In: Barzen J, Leymann F, Dustdar S, editors. Service-oriented computing. Cham: Springer International Publishing; 2022. pp. 79–98.
- 31.
United Nations Office on Drugs and Crime. Human trafficking indicators [Internet]. Vienna: UNODC; [date unknown] [cited 2026 Apr 17]. Available from: https://www.unodc.org/pdf/HT_indicators_E_LOWRES.pdf
- 32. Hundman K, Gowda T, Kejriwal M, Boecking B. Always lurking: Understanding and mitigating bias in online human trafficking detection. Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society. New Orleans LA USA: ACM; 2018. pp. 137–43.
- 33.
Zhou S, Peng J, Ferrara E. Tracing the unseen: uncovering human trafficking patterns in job listings. In: Proceedings of the International Workshop on Data for the Wellbeing of the Most Vulnerable at the 18th International AAAI Conference on Web and Social Media (ICWSM 2024) [Internet]; 2024 Jun 3; Buffalo (NY). Palo Alto (CA): AAAI Press; 2024 [cited 2026 Jul 26]. Available from: https://workshop-proceedings.icwsm.org/pdf/2024_27.pdf
- 34. Mohd Hanif AH, Maarop N, Kamaruddin N, Samy GN. Machine learning approach in predicting fraudulent job advertisement. Int J Acad Res Bus Soc Sci. 2024;14: 1182–93.
- 35. Bora AE, Arodi A, Zhang D, Bannister J, Bronzi M, Fansi Tchango A, et al. AIMSCheck: Leveraging LLMs for AI-assisted review of modern slavery statements across jurisdictions. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Vienna, Austria: Association for Computational Linguistics; 2025. pp. 9109–35.
- 36. Polaris Project. Massage business trafficking quick facts [Internet]. Washington (DC): Polaris Project; 2018 [cited 2026 Apr 17]. Available from: https://polarisproject.org/wp-content/uploads/2021/03/Massage-Business-Trafficking-Quick-Facts.pdf
- 37. Li R, Tobey M, Mayorga ME, Caltagirone S, Özaltın OY. Detecting human trafficking: Automated classification of online customer reviews of massage businesses. Manuf Serv Oper Manag. 2023;25: 1051–65.
- 38. Tobey M, Li R, Özaltın OY, Mayorga ME, Caltagirone S. Interpretable models for the automated detection of human trafficking in illicit massage businesses. IISE Transactions. 2022;56(3):311–24.
- 39. White A, Guikema S, Carr B. Why are you here? Modeling illicit massage business location characteristics with machine learning. J Hum Traffick. 2024;10: 20–40.
- 40. Ouyang R. Machine learning for tangible effects: natural language processing for uncovering the illicit massage industry & computer vision for tactile sensing [dissertation]. Cambridge (MA): Harvard University; 2023 [cited 2026 Apr 17]. Available from: https://dash.harvard.edu/entities/publication/044ba436-5786-4fec-9549-ecc1c41775d1
- 41. Tobey M. Data science and machine learning methodologies for detecting human trafficking risk in the illicit massage industry [dissertation]. Raleigh (NC): North Carolina State University; 2023 [cited 2026 Apr 17]. Available from: https://repository.lib.ncsu.edu/items/ec7c9b8f-ef95-4443-ac0e-215c8a366459
- 42. de Vries I, Radford J. Identifying online risk markers of hard-to-observe crimes through semi-inductive triangulation: The case of human trafficking in the United States. Br J Criminol. 2022;62:639–58.
- 43.
Garg V, Özaltın OY, Mayorga ME, Bosisto S. Detecting illicit massage businesses by leveraging graph machine learning. In: Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence (IJCAI-25); 2024. p. 9647–55. https://doi.org/10.24963/ijcai.2025/1072
- 44. Thöni A, Taudes A, Tjoa AM. An information system for assessing the likelihood of child labor in supplier locations leveraging Bayesian networks and text mining. Inf Syst E-Bus Manage. 2018;16(2):443–76.
- 45. Nallakaruppan MK, Srivastava G, Gadekallu TR, Reddy PK, Krishnan S, Połap D. Child tracking and prediction of violence on children in social media using natural language processing and machine learning. In: Rutkowski L, Scherer R, Korytkowski M, Pedrycz W, Tadeusiewicz R, Zurada JM. Artificial intelligence and soft computing. Cham: Springer Nature Switzerland; 2023. pp. 560–9.
- 46.
da Silva Santos M, Ladeira M, Van Erven GCG, Luiz da Silva G. Machine Learning Models to Identify the Risk of Modern Slavery in Brazilian Cities. In: 2019 18th IEEE International Conference On Machine Learning And Applications (ICMLA), 2019. 740–6. https://doi.org/10.1109/icmla.2019.00132
- 47. Nersessian D, Pachamanova D. Human trafficking in the global supply chain: Using machine learning to understand corporate disclosures under the UK Modern Slavery Act. Harv Hum Rights J [Internet]. 2022 [cited 2026 Apr 17];35: 1–46. Available from: https://journals.law.harvard.edu/hrj/wp-content/uploads/sites/83/2022/05/35HHRJ1-Nersessian.pdf
- 48. Chae Y, Davidson T. Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning. Sociological Methods & Research. 2025;55(2):501–67.
- 49.
Mendez Guzman E, Schlegel V, Batista-Navarro R. RaFoLa: A Rationale-Annotated Corpus for Detecting Indicators of Forced Labour. In: Proceedings of the Language Resources and Evaluation Conference, 2022. 3610–25. https://doi.org/10.63317/4i2g7vc3hjbg
- 50. Joo R, McDonald G, Miller N, Kroodsma D, Farthing C, Belhabib D, et al. Towards a responsible machine learning approach to identify forced labor in fisheries. arXiv:2302.10987 [Preprint]. 2023 [cited 2026 Apr 17]. Available from: https://arxiv.org/abs/2302.10987
- 51. Stringer C, Whittaker DH, Simmons G. New Zealand’s turbulent waters: the use of forced labour in the fishing industry. Glob Netw. 2016;16: 3–24.
- 52. Perez AR, Rivas P. Combatting human trafficking in the cyberspace: A natural language processing-based methodology to analyze the language in online advertisements. arXiv:2311.13118 [Preprint]. 2023 [cited 2026 Apr 17]. Available from: https://arxiv.org/abs/2311.13118
- 53.
Joulin A, Grave E, Bojanowski P, Mikolov T. Bag of Tricks for Efficient Text Classification. In: Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, 2017. 427–31. https://doi.org/10.18653/v1/e17-2068
- 54. Kejriwal M, Ding J, Shao R, Kumar A, Szekely P. FlagIt: A system for minimally supervised human trafficking indicator mining. arXiv:1712.03086 [Preprint]. 2017 [cited 2026 Apr 17]. Available from: https://arxiv.org/abs/1712.03086
- 55. Giommoni L. Why we cannot identify human trafficking from online advertisements. J Hum Traffick. 2024;1-15. Epub 2024 Dec 8.
- 56. Nair P, Lefebvre G, Garrel S, Molamohammadi M, Rabbany R. Ask before you build: Rethinking AI-for-Good in human trafficking interventions. arXiv:2506.22512 [Preprint]. 2025 [cited 2026 Jul 16]. Available from: https://arxiv.org/abs/2506.22512
- 57. Deeb-Swihart J, Endert A, Bruckman A. Ethical Tensions in Applications of AI for Addressing Human Trafficking: A Human Rights Perspective. Proc ACM Hum-Comput Interact. 2022;6(CSCW2):1–29.