Figures
Abstract
Hikikomori, or prolonged social withdrawal, is an issue of global relevance. The HRI-15 is a brief tool for its assessment. This study enhances its utility by adding clinical thresholds to identify risk and person-centered clinical profiles, and by developing a machine-learning(ML)–based computerized adaptive test (CAT) to enable rapid screening. Data from a national survey of Italian adolescents (N = 8,755) were used to conduct ROC analysis and latent profile analysis (LPA). The findings indicated that a score of ≥ 42 achieved optimal classification performance, and four profiles with distinct meanings and systematic differences in anxiety, depression, impulsivity, and risk behaviors were identified. A multivariate conditional inference tree was estimated to develop an ML-based adaptive version of the instrument. The CAT reduced administered items by 53% (7.03/15), accurately reproducing full-length scores (r = .77–.997) and the corresponding classification; alignment was assessed with latent transition analysis (entropy = .96). The procedure was integrated into an application supporting administration, scoring, and reporting (score, risk, profiles, graphs). Combining a cross-validated cutoff and clinically useful profiles with CAT and its automated administration and reporting application strengthens triage and personalization, enabling large-scale, multidomain screening and earlier intervention for hikikomori risk.
Citation: Colledani D, Anselmi P, Monacis L, Genetti B, Fassinato D, Gomez Perez LJ, et al. (2026) Interpretable machine learning for hikikomori screening: The adaptive HRI-15. PLoS One 21(8): e0355595. https://doi.org/10.1371/journal.pone.0355595
Editor: Alberto Greco, Universita degli Studi di Pisa, ITALY
Received: March 3, 2026; Accepted: July 23, 2026; Published: August 20, 2026
Copyright: © 2026 Colledani et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: The data, analysis code, and link to the Shiny application are publicly available on OSF: https://osf.io/3857s/overview?view_only=78a57773986941068b6b094b225487b3.
Funding: This study received financial support from Department for Antidrug Policies [collaboration agreement CA09, 2021]. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The authors have declared that no competing interests exist.
Introduction
In recent years, the phenomenon of prolonged social withdrawal, known as hikikomori, has become a major challenge for mental health, education, and healthcare systems [1]. Hikikomori refers to a persistent pattern in which individuals substantially limit social participation and relationships outside the immediate family, remaining confined to their room or home for a long time [2–5]. In its broader sense, the term denotes a significant and problematic form of prolonged social withdrawal. Although hikikomori has been discussed in contemporary nosological debates and is mentioned in the DSM-5/DSM-5-TR [6] in relation to cultural concepts of distress, it is not listed as a discrete DSM-5/DSM-5-TR disorder. Rather, it is better understood as a clinically relevant and culturally inflected syndrome/condition whose nosological status remains under debate. Proposed research criteria typically emphasize marked home-based social withdrawal lasting at least six months, together with functional impairment or distress.
Although originally described in Japan in connection with school absenteeism (futoko) under intense educational demands [3,4,7–9], the phenomenon now exhibits a global spread across different age groups, with reported cases from Asian and Western settings, including the USA and EU [7,10–19]. This trajectory accelerated during the COVID-19 pandemic, which imposed extended periods of isolation and shifted social interaction and work into digital environments [20–22].
Despite the absence of unified global epidemiological estimates, available research suggests that hikikomori affects about 1.2% of the world population [17]. In Japan, studies have reported prevalence between 0.9% and 2.4%, with some studies indicating 22% in specific samples of young people [2,5,15,23,24] Estimates of similar magnitude emerge in other Asian contexts: Hong Kong 5.0–9.10% [25,26], mainland China 1.90–6.57% [27–29], and Korea around 1.4% [30]. In Western countries, estimates range from 2.03% to 16.13% in the EU [31,32] and are about 2.65% in the United States [7]. In general, prevalence figures vary widely across studies due to differences in cultural context, target population, and measurement tools; for example, a national survey in Oman found rates close to 44%, markedly surpassing estimates from other countries [33]. In Western countries, although systematic surveillance remains limited, converging evidence points to a substantial problem, especially among young people. Consistent with this, a national study on Italian adolescents showed that social withdrawal tendencies are already present at age 13, underscoring the need for targeted early detection [34]. The literature further indicates that, even though onset typically occurs in adolescence [20,24,35,36], once established, withdrawal often persists over the years, with detrimental consequences for personal development, academic/occupational functioning, and psychological well-being [16,19,24,37–41]. Taken together, these findings underscore the need for early identification and large-scale screening to limit chronicity and enable timely interventions.
To address this need, psychometric research has developed specialized screening instruments. Among the most widely used self-report measures are the NEET-Hikikomori Risk [42], the Hikikomori Questionnaire [43], and the Hikikomori Risk Inventory [44]. The HRI-24, developed jointly by Japanese and Italian researchers, shows robust psychometric properties and cross-cultural invariance, supporting its use as a promising instrument for international research on hikikomori. Moreover, its five-factor structure (anthropophobia, agoraphobia, lethargy, paranoia, and depressive mood) allows for the evaluation of distinct facets of the hikikomori condition. This is noteworthy because it may enable the creation of specific risk profiles, which would be clinically valuable as they could have unique antecedents and outcomes and inform personalized preventive and therapeutic pathways. Another strength of the instrument comes from the availability of a brief version, which facilitates its use in broad-spectrum screening and epidemiological studies. The brief form (HRI-15) has demonstrated adequate psychometric properties, including age- and gender-invariance, as well as ROC-derived cut-off scores showing sufficient diagnostic accuracy [11].
Although the instrument is valid and relatively time-efficient, it can still be challenging to use when multiple psychopathological conditions must be screened at once. In clinical practice, research surveys, and epidemiological studies, comprehensive assessments routinely require broad-spectrum screening across several health domains [45,46]. These procedures are essential for targeting interventions and describing populations and phenomena under study. However, administering extensive questionnaires increases testing time and respondent burden, raising the risk of fatigue, incomplete responses and reduced data quality [46–48]. These demands can be especially challenging for vulnerable individuals and require significant resources from service providers [45,49].
Computerized Adaptive Testing (CAT) offers a useful strategy to mitigate these issues. By tailoring item administration to prior responses, CAT markedly reduces the number of items while preserving score precision [45,50,51]. In recent years, integration with machine learning (ML) has made CAT procedures particularly efficient, accurate, and straightforward to implement. ML-based CAT can reduce administration burden without compromising validity, maintain high classification accuracy, and deliver reliable clinical categorizations, including in longitudinal settings [52–54].
This study aims to introduce an ML-based computerized adaptive version (ML-based CAT) of the HRI-15 that shortens administration and reporting time while improving classification reliability and clinical utility. The procedure automatically returns dimensional scores, flags scores that exceed the ROC-based cut-off for hikikomori risk, and assigns respondents to previously defined risk profiles (e.g., via latent profile analysis). By providing faster, more accurate, and practice-ready outputs, the ML-based CAT has the potential to enhance the accuracy and efficiency of clinical, scientific, and epidemiological assessments and screening procedures.
Aims and analytic overview
To meet the proposed objectives, the study advances through two principal stages: first, the definition of detailed diagnostic criteria and clinical risk profiles; second, the development of an ML-based CAT version of the instrument with an automated administration and reporting application.
As mentioned above, the definition of specific risk profiles based on HRI dimensions is expected to be particularly valuable, as it has the potential to identify specific risk groups, allowing for the analysis of their distinctive characteristics and associated factors. This, in turn, should lead to the development of more personalized and targeted interventions, fully leveraging the tool’s potential. In parallel, the development of an ML-based CAT version of the tool has the potential to reduce respondent boredom and fatigue thereby increasing enrollment rates and survey adherence, facilitating longitudinal monitoring, simultaneously reducing time, costs, and scoring errors.
Diagnostic categorization
To be diagnostically useful, test scores must be translated into explicit decision criteria. In this regard, this study aims to make two primary contributions: the establishment of a cut-off score for identifying hikikomori risk using receiver operating characteristic (ROC) analysis; and the delineation of latent profiles derived from the HRI-15 dimensions through latent profile analysis (LPA).
Establishing a cut-off score for classifying at-risk individuals is critical to enhancing the practical value of any screening or diagnostic instrument. Operationally, ROC analysis provides a rigorous procedure for determining, among all possible test scores, the threshold that most effectively differentiates individuals with and without a specific condition [55]. ROC analyses require two inputs: observed scores on the target measure and an external criterion used as a pragmatic indicator of whether cases meet the target condition. For each possible score, sensitivity (true positive rate) and specificity (true negative rate) are estimated, and the optimal cut-off score is selected as the score that maximizes both [55]. In this study, consistent with the literature indicating a general factor underlying the HRI-15, the total scale score was used as the target measure to establish the optimal cut-off score for the scale [11]. The external criterion was a self-report item assessing how often participants locked themselves in their rooms during the past six months. Responses were recorded on a frequency scale and dichotomized as not at risk (“Never”, “Only once”, “1–2 times per month”) or at risk (“Every week (but not every day)”, “Every day”). This item was considered appropriate for calibrating a practical screening threshold because it directly reflects a core behavioral manifestation of prolonged social withdrawal. Previous work on the HRI-15, using a similar procedure, identified a cut-off score (≥ 37) with encouraging performance (i.e., AUC = 0.80; sensitivity = 0.78; specificity = 0.71; accuracy = 0.72; [11]); however, those analyses relied on a binary external criterion (“Have you ever experienced the tendency to lock yourself in your room for several months, never going out, not even to eat meals or to entertain social relations?”; yes/no). The present study advances this evidence by using a larger sample and a more detailed, frequency-based criterion, thereby enabling a more behaviorally specific cut-off calibration and a more detailed assessment of classification performance.
While establishing precise cut-off scores is crucial for clinical practice, these values are confined to reflecting only the total score. Although suitable for preliminary screening, this approach overlooks potentially meaningful profile differences, thereby limiting the development of individualized interventions and the advancement of understanding of the phenomenon. To address this limitation, the present study proposes a novel contribution through LPA. LPA is a person-centered, model-based approach that partitions a population into latent groups characterized by within-class homogeneity and between-class heterogeneity on a set of indicators [56,57]. It is a rigorous method that yields clinically meaningful information, which in turn can facilitate the development of more individualized prevention and treatment approaches [58,59]. Specifically, in this work, LPA was used to detect latent profiles based on the five HRI-15 dimensions. The emerging profiles were also examined against a set of psychosocial indicators (e.g., anxiety, depression, impulsivity, substance use, and related functional outcomes) to assess their clinical utility. Notably, this profile-based characterization has not yet been undertaken for the HRI; however, it has the potential to substantially expand the instrument’s utility by fully leveraging its value for screening and interventions.
CAT development
In clinical and research contexts, assessment batteries often incorporate multiple instruments to gain a nuanced understanding of individuals and populations. While this approach provides valuable insights, it also increases the workload for respondents, professionals, and organizations. The use of brief scales can be an effective strategy for managing this issue. However, when many constructs need to be evaluated, the overall assessment length can still be quite extensive, even if each construct is measured with a brief scale. CAT is an innovative and effective method for improving measurement efficiency without compromising accuracy [50,51]. Computerized adaptive tests (CATs) personalize the testing experience by presenting respondents with a reduced number of items selected dynamically for each person based on item properties and the responses progressively provided by participants during the testing process. Consequently, unlike fixed-form tests, CATs do not administer the entire set of items to each respondent, but focus on a restricted set of items that are most informative for each respondent at a given time [60,61]. Most CATs originate in the framework of item response theory (IRT) and have demonstrated notable efficiency in in various fields, including education, mental health and personality assessment, as they facilitate the precise measurement of disorder severity using a markedly limited number of items [62–66].
More recent approaches to CAT development draw on ML algorithms based on recursive binary partitioning, such as Classification and Regression Trees (CART). Compared to IRT-based procedures, these methods are generally easier to implement because they do not require calibration of large item banks and adherence to strict distribution or dimensionality assumptions (e.g., unidimensionality). Furthermore, their straightforward integration into administration systems and provision of transparent, clinically relevant decision rules enhance their utility and appeal in psychological and psychodiagnostic settings [34,46,52–54,67,68]. Early work in this area focused on procedures that assigned respondents to discrete categories (e.g., diagnosed vs. not diagnosed; at risk vs. not at risk). This framework is known as computerized adaptive diagnosis (CAD; [46]). Recently, however, CART-based approaches have been extended to estimate continuous outcomes, such as test scores, disorder severity, and trait levels. This enables a more nuanced quantitative assessment that goes beyond simple categorization, thus expanding their range of applications [53,67,69].
In general, CART algorithms operate by developing a tree structure that can be conceptualized as a flowchart defined through a recursive, top-down, divide-and-conquer process that generates a sequence of if-else rules aimed at creating classifications (decision trees, DT) or continuous predictions (regression trees, RT; model trees, MT; [70,71]). The tree comprises a root (the variable from which prediction/classification begins), internal nodes (additional variables considered along the path), branches, and leaves. Branches are constituted by a sequence of nodes, connected through specific decision rules (e.g., a numerical threshold of ≤ 1 vs. > 1 for a 0–5 variable) and leaves correspond to the termination of a branch, representing the final prediction of the defined process (class or score).
As binary partitioning methods, CART algorithms split cases according to the rules coded in the tree structure. The development of the algorithm starts with the selection of a root variable that includes the entire dataset, as well as an associated splitting rule that divides the sample into subsets that are as homogeneous as possible with respect to the outcome (classification or score). The procedure then continues recursively within each subset until specific stopping criteria are met (e.g., maximum depth, minimum node size, or no further gain in homogeneity; [71]). In the context of categorical outcomes, the objective is to maximize node purity, a concept that is typically quantified via information gain and entropy reductions [71–74]. In the context of continuous outcomes, the implementation of splits is based on the objective of minimizing within-node variance. This is frequently operationalized as a reduction in standard deviation [71].
Model definition typically follows a two-step approach. First, the tree is developed using a dataset (training dataset) that contains predictors and outcomes for all cases. Then, its predictive performance is evaluated using an independent dataset (testing dataset) with the same structure, but comprising cases that were not used in the development stage [75,76]. This cross-validation approach provides an estimate of generalizability and practical utility of the learned tree for predicting unseen data.
An emerging body of research indicates that tree-based adaptive procedures are effective in psychometric and diagnostic applications [46,53,67–69,77]. Conceptually, once a tree structure has been developed, it can be used to guide the adaptive administration of tests. Respondents are presented only with items located along the branch to which they are progressively assigned based on their responses [46,52,69,77,78]. This architecture offers several practical advantages. First, the adaptive process employs a precompiled scheme based on the established branching structure, thus allowing for rapid deployment with minimal runtime computation [68]. This contrasts with many IRT-based implementations, which require more complex procedures [60,61]. Second, this approach is less dependent on stringent theoretical assumptions, facilitating its application across clinical, school, and epidemiological contexts [46,54].
Evidence indicates that ML-based CATs perform well for both categorical classification and continuous score estimation, yielding valid and time-efficient assessments that closely align with full-length measures in cross-sectional and longitudinal designs [53,54,69,77–80]. Moreover, agreement with external variables is typically indistinguishable from that obtained with full-scale administrations [53,67].
Building on these foundations, the present study aims to demonstrate that an ML-based CAT can improve psychometric screening and diagnostic assessment by producing clinically meaningful outputs (accurate score estimation, risk-profile categorization, and precise cutoff-based classification) while simultaneously reducing respondent burden and facilitating test administration. This efficiency has the potential to enable large-scale, multi-domain screening, making administration more feasible in real-world settings without compromising accuracy or interpretability. In this context, the HRI-15 is an appropriate test case because it assesses distinct dimensions relevant to hikikomori, allowing the identification of clinically useful risk profiles that can substantially complement the essential cutoff-based classification that remains fundamental for screening. The manuscript details the development of an ML-based CAT and introduces a dedicated administration, scoring, and reporting application that returns dimensional scores in real time, flags cutoff exceedances, assigns LPA-based profiles, and generates a concise, practitioner-oriented report suitable for clinical, educational, and epidemiological contexts.
Method
Participants
Data were drawn from a large, nationally coordinated, school-based survey of Italian adolescents aged 11–17 years. The sampling plan was designed to produce a representative student cohort across the country and employed probabilistic selection with stratification by age band (11–13; 14–17), school type, and geographical area. To achieve at least 4,000 valid questionnaires per age group—and assuming a response rate of 15–20%—676 schools across all Italian regions were invited. Final participation rates were 22.7% for middle schools (ranging 20.0–26.2% across areas) and 22.6% for high schools (range 18.5–26.8%). All regions contributed data, ensuring broad national coverage.
In total, 10,181 questionnaires were submitted. Following data-quality checks, 1,426 records (14.0%) were excluded due to ineligibility (age outside target range), incomplete responses, or duplicate submissions caused by technical issues during online entry. The final analytic dataset thus included 8,755 valid cases.
Written informed consent was obtained from students and their legal guardians prior to participation. Data collection and processing complied with national privacy regulations, guaranteeing respondent anonymity. The study received ethical approval from the National Ethics Committee of the Italian National Institute of Health (Protocol PRE BIO CE 0010655, March 22, 2022) and adhered to the principles of the 1964 Declaration of Helsinki and its subsequent amendments.
The final sample comprised 8,755 students (M age = 14.03 years, SD = 1.98). Of them, 3,623 were middle-school students (41.4%; M age = 11.99, SD = 0.81) and 5,132 were high-school students (58.6%; M age = 15.48, SD = 1.08). Within the high-school subsample, 7.7% attended fine-arts secondary schools, 31.8% vocational institutes, 29.4% technical institutes, and 31.1% other types of high schools. The gender distribution was approximately balanced (4,187 females, 47.8%; 4,291 males, 49.0%), with 277 students (3.2%) not disclosing gender. Most participants reported Italian nationality (N = 7,535; 86.1%).
Measures
All measures were collected using closed-ended items through a computer-assisted web interviewing platform during school hours. Trained supervisors were present to oversee the process. The battery comprised: demographic indicators (age, gender, region), lifestyle and behavioral indicators (e.g., participation in social-media challenges, alcohol use, cannabinoids use), and standardized scales assessing hikikomori risk and related psychosocial constructs.
Lifestyle and behavioral indicators.
One item assessed the tendency to self-isolate over the past six months (“Over the last six months, have you isolated yourself in your room for extended periods of time without ever going out—not even to eat meals or engage in social interactions?”). The response options were: 1 = Never, 2 = Only once, 3 = One to two times per month, 4 = Every week (but not every day), and 5 = Every day. For the analyses, the responses were dichotomized as not-at-risk (answers from 1 to 3) and at-risk (answers from 4 to 5).
Daily use of social media and video games over the previous week was assessed using two separate items. Participants were asked to quantify their usage by answering the following questions: “In the past week, how many hours per day did you use social media?” and “In the past week, how many hours per day did you play video games?” Responses were recorded on a five-point scale: 1 = Less than 1 hour, 2 = 1–2 hours, 3 = 3–6 hours, 4 = 7–8 hours, 5 = 9 hours or more.
An additional set of dichotomous (yes/no) items was administered to assess various risk behaviors. The following variables were included in the study: participation in social media challenges during the past six months (“In the past six months, have you taken part in social media challenges?”); history of alcohol consumption (“Have you ever drunk alcohol?”); history of sending erotic/sexual messages (i.e., sexting, “Have you ever sent erotic/sexual messages, videos, or personal photos via smartphone, e-mail, or webcam?”); and history of cannabis use (“Have you ever used cannabis?”; for ethical reasons the last two questions were presented only to the high school students group).
As a protective behavior, the tendency to seek assistance from parents or caregivers was measured using a yes/no item (“When you have problems, do you talk with your parents/caregivers?”).
Hikikomori risk.
Hikikomori risk was assessed with the Hikikomori Risk Inventory–15 (HRI-15; [11]). The scale is the short form of the HRI-24 [44] and comprises 15 items rated on a five-point Likert scale (1 = Strongly disagree; 5 = Strongly agree), referring to the past six months. The instrument indexes five dimensions: anthropophobia, agoraphobia, paranoia, lethargy, and depressive mood. These dimensions were derived through factor-analytic refinement from items tapping hikikomori symptomatology and related features as theorized by [8]. Anthropophobia, reflects fear of people and social contact (e.g., “I feel severely anxious when I meet people outside”); agoraphobia denotes avoidance of settings where assistance might be difficult to obtain during panic or intense anxiety (e.g., “I avoid enclosed public spaces such as shops, cinemas, theaters, or banks”); lethargy captures low energy and behavioral disengagement (e.g., “I sleep many hours a day because I often feel weak”); paranoia reflects suspiciousness and mistrust (e.g., “I am afraid of being deceived by others”); and depressive mood indexes sadness, anhedonia, and discomfort (e.g., “I feel a sense of inner emptiness”). Despite the inherent multidimensionality of the structure of the scale, existing evidence suggests that a robust general factor underlies the unidimensional interpretation of the total score [11]. The short form demonstrated satisfactory psychometric properties, with an area under the curve (AUC) of .80, sensitivity of .78, specificity of .71, and accuracy of .72. It also exhibited invariance by gender and age group (middle school versus high school). In this sample internal consistency was α = .92 for the total scale score, while it was .84, .83, .77, .82, and .79 for anthropophobia, agoraphobia, paranoia, lethargy, and depressive mood, respectively.
Related psychosocial constructs.
Social anxiety and depressive symptoms were assessed using the Italian adaptations of the DSM-5 Severity Measures for Social Anxiety [81] and Depression [82]. The social anxiety scale includes 10 items rated from 0 (“never”) to 4 (“all the time”) referring to the past seven days (e.g., “In the past seven days, I spent a lot of time thinking about what to say or how to behave in social situations”). The depression scale comprises 9 items rated from 0 (“not at all”) to 3 (“nearly every day”) referring to the past week (e.g., “In the past seven days, how often did you feel down, depressed, irritable, or hopeless?”). Higher total scores indicate greater symptom severity; internal consistency was α = .93 for social anxiety and α = .89 for depression.
Impulsivity was included as a divergent construct because, in the HRI-24 validation, sensation seeking (e.g., BSSS) was employed to assess divergent validity; sensation seeking is closely related to impulsivity, and both reflect approach-oriented, externalizing tendencies that stand in contrast to social withdrawal, the core of hikikomori [44,83–86]. Trait impulsivity was measured with the Italian BIS-15 (Barratt Impulsiveness Scale–15; [87]), a brief form of the BIS-11 [88]. Although the BIS-15 has not been specifically validated in Italian adolescents, the full version has been adapted for this population [89]. Items are rated on a four-point scale (e.g., “I have self-control”). Following recent recommendations [90], BIS-15 scores were treated as a continuous score; higher scores indicate greater impulsivity (α = .77).
Analysis strategy
ROC analysis.
ROC analyses were conducted on the HRI-15 total score using an external criterion based on the item: “Over the past six months, have you stayed in your room for extended periods without going out—not even for meals or social interactions?”. Responses were dichotomized as positive (“Every week, but not every day” or “Every day”) versus negative (“Never” through “One to two times per month”). The dataset was randomly partitioned into a training set (approximately 67% of the total sample) and a test set (remaining 33%). The training set (n = 5,866; 48.9% male; M age = 14.00, SD = 1.99) was used to estimate the ROC curve, compute the area under the curve (AUC) with its 95% confidence intervals, and determine the optimal cut-off via Youden’s index (Sensitivity + Specificity – 1; higher values indicate better combined sensitivity and specificity). The selected threshold was then applied to the independent test data set (n = 2,889; 49.3% male; M age = 14.10, SD = 1.96), to compute sensitivity, specificity, and overall accuracy with their 95% confidence intervals. ROC analyses were performed in the statistical environment R [91] using the pROC package [92]. To examine the robustness of the threshold to alternative operationalizations of the external criterion, sensitivity analyses were conducted using different dichotomizations of the room-confinement item. Specifically, the primary coding defined positivity as responses 4–5 (“Every week (but not every day)” / “Every day”), whereas a stricter coding treated only the response “Every day” as positive. This comparison was used to assess whether the resulting cut-off remained stable under a narrower definition of social withdrawal. Finally, supplementary calibration analyses were conducted by estimating logistic regression models in the training dataset and applying them to the testing dataset to obtain predicted probabilities. Calibration was evaluated by comparing predicted risk probabilities with observed outcomes in the testing dataset. The Brier score was used as a summary index of the overall discrepancy between predicted probabilities and observed outcomes, with lower values indicating better probabilistic accuracy [93]. Calibration intercept and calibration slope were also examined: the ideal values are 0 for the intercept and 1 for the slope; negative or positive intercept values indicate systematic overestimation or underestimation of risk, respectively, whereas slopes below or above 1 indicate overly extreme or overly moderate predicted probabilities, respectively [94].
Latent Profile Analysis (LPA).
LPA was employed to detect latent profiles based on the five HRI-15 scales. LPA is a person-centered approach that classifies individuals into groups defined by comparable response patterns on a set of indicators, providing a holistic view of individual differences that can be especially informative in applied research and clinical practice [95–99]. To date, no latent-profile solution for the HRI-15 has been reported, making this analysis a novel contribution. The analysis was therefore conducted on the full sample to maximize precision and obtain realistic parameter estimates. In particular, a sequence of one- through five-class models was estimated and compared, using standardized scores on the five HRI-15 scales as indicators. To identify the preferred solution, multiple statistical and substantive criteria were considered jointly, including Akaike’s Information Criterion (AIC), the Bayesian Information Criterion (BIC), the sample-size–adjusted BIC (SABIC), and entropy. Entropy is a measure of classification accuracy ranging from 0 to 1. Higher values indicate a clearer separation between classes. Conversely, for AIC, BIC, and SABIC greater accuracy is suggested by lower values [57]. Two comparative tests were also considered for model evaluation: the Vuong–Lo–Mendell–Rubin likelihood ratio test (VLMR; [57]) and the Lo–Mendell–Rubin likelihood ratio test (LMR; [100]). These tests are used to compare a model with C latent classes against a model with C–1 classes. Significant p-values indicate that the model with C classes better fits the data than a more parsimonious one (i.e., the model with one fewer class). In contrast, non-significant p-values suggest retaining the more parsimonious model [57,100]. Finally, class sizes and the interpretability of the resulting profiles were taken into account in selecting the solution to be retained. As mixture-model fit indices often fail to converge on a single optimal solution, no single statistic was considered to be conclusive. Instead, model selection was based on an overall balance of relative model fit, classification quality, parsimony, class size, and interpretability [57,93,97]. Analyses were conducted using Mplus 7.4 [56], with the robust maximum likelihood estimator (MLR; [101]) and mixture specification. Between-class differences on a set of psychosocial indicators were examined using ANOVA and χ2 statistics, with eta-squared (η2) and Cramer’s V as effect size measures.
Machine-Learning Computerized Adaptive Test (CAT).
CAT development: A real-data simulation approach with cross-validation was used to develop the CAT. Specifically, a multivariate conditional-inference tree (ctree) was estimated on the training dataset (67%; random split). In the model the 15 items of the HRI-15 [11,44] were used as predictors, while the HRI-15 total score (sum) and the five standardized (M = 0, SD = 1) scores on its subscales (anthropophobia, agoraphobia, paranoia, lethargy, depressive mood) were the joint outcome variables. The ctree algorithm, as with other regression trees, employs recursive binary partitioning [102]. However, in this framework, in contrast to conventional approaches, both variable and splitting rule selection are determined by permutation tests. This allows for the management of multiple outcomes within a single tree and reduces selection bias [103–105]. At each node, the algorithm tests a global null hypothesis of independence between the predictors and the outcome variable(s); if the null hypothesis is not rejected, the node becomes a leaf. If the null hypothesis is rejected, the predictor with the smallest p-value (adjusted for multiplicity) is selected for splitting; and for that predictor, all possible binary partitions are evaluated, choosing the split that maximizes the test statistic across branches. For a single continuous outcome, the split statistic can be interpreted as a standardized difference between the mean outcome values in the two child nodes defined by a candidate split. For ordered outcomes, the comparison is instead based on a linear-rank statistic computed from the outcome scores/ranks. For multivariate outcomes, the statistic combines the differences between the two child nodes across the outcome vector into a single quadratic-form test statistic [11,44]. At each split, specific node-size constraints have to be met. Specifically, the constraints pertain to the minimum number of observations required in the nodes (or the sum of the case weights) to attempt a split (i.e., the minimum split or “minsplit”) and the minimum number of observations that must be included in each leaf (i.e., the minimum bucket or “minbucket”). The thresholds for these parameters are set to balance tree parsimony against predictive accuracy (higher thresholds favor simplicity; lower thresholds allow for a closer fit, at the risk of overfitting). At the end of the branches growing process, leaves provide piecewise-constant predictions representing the mean of the outcome variable(s) in a specific subset (which is equivalent to the intercept of an intercept-only model for that subset). The branches growth stops when no covariate shows sufficient association with the outcome(s) at the prespecified significance threshold (i.e., “mincriterion”). This threshold is the p-value that determines whether a node should be split (it is typically set as 1 − α; e.g., a mincriterion of .95 corresponds to α .05, meaning that p ≤ 0.05 is required for a split) or when node/leaf sample-size constraints (i.e., minsplit/minbucket) are not met; therefore, no pruning or smoothing is required. In the present study, candidate tuning configurations were examined over a broad grid of plausible values for mincriterion (.10, .20, .50, .70, .90), minsplit (10, 20, 50), and minbucket (10, 20, 50). Candidate configurations were first compared by 5-fold cross-validation within the training subset. They were then refitted on the full training subset and evaluated on the corresponding independent test subset to examine out-of-sample performance and the sensitivity of agreement, prediction error, and item savings to alternative hyperparameter settings. The final configuration was chosen to balance predictive accuracy and administration efficiency. Analyses were performed in the statistical environment R [91] using the partykit package [104,106,107] (Fig 1).
Note. Panel A summarizes the node-wise procedure used to grow the conditional-inference tree. At each node, the algorithm tests the global null hypothesis of independence between the predictors and the outcome(s). If the null hypothesis is rejected, the predictor with the smallest adjusted p value is selected, and the admissible binary split that maximizes the test statistic is chosen; otherwise, the node becomes terminal. Panel B provides a schematic illustration of the resulting adaptive routing structure. The dark nodes and terminal node connected by dashed lines indicate one possible branch of the tree. Starting from the root node, all respondents receive the same initial item; subsequent items depend on the responses provided at each preceding node. This recursive process continues until a terminal node (TN) is reached. Each TN corresponds to a subgroup of respondents defined by a specific response pattern and stores the predicted score(s) associated with that branch.
Evaluation of CAT performance: Once the tree structure had been defined using the above-described procedure on the training dataset, it was used to instruct an adaptive assessment procedure based on respondents’ answers in the test dataset, for which complete responses were available (i.e., real data simulation). The procedure simulated the administration to participants of only the items that corresponded to the branch they were moving along based on their responses, and it produced total (sum score) and subscale (z-score) score estimates. On this dataset (not used to develop the tree), the CAT performance was evaluated in terms of efficiency and accuracy. The efficiency of the CAT procedure relative to the full-length test was measured by item savings (the average number of items administered and the corresponding reduction percentage) and evaluated using t-tests. Accuracy, interpreted as score reproduction, was quantified using Pearson’s correlation coefficients (r) and mean absolute error (MAE). Correlations indicate the degree to which CAT scores align with those obtained by administering all items (a higher r indicates better performance), while MAE captures the average absolute deviation between CAT and full-length scores (a lower MAE indicates closer alignment). Paired samples t-tests were also run to compare CAT-derived scores with full-length scores. The t-tests assess the null hypothesis that the mean of CAT scores does not differ from that of full-length scores.Agreement between CAT- and full-length scores was assessed also using intraclass correlation coefficients computed on the independent test set for each HRI-15 scale and the total score, with a two-way model, absolute agreement, single measure (ICC(A,1)). Interpretation followed [108]: ICC < .50 poor, .50–.75 moderate, .75–.90 good, > .90 excellent.
The performance of the CAT was also evaluated in terms of its accuracy in classifying risk. The cut-off score derived from the ROC curve estimated on the training data set was applied to both the total scores obtained through the CAT and the total scores obtained from the full-length administration. For both sets of scores, the following metrics were checked: sensitivity, specificity, overall accuracy, and Cohen’s κ coefficients. Cohen’s κ represent chance-corrected agreement between each method’s classifications and the criterion variable (e.g., external variable indicating the tendency to social withdrawn). Higher values indicate that the method’s predicted labels better reflect the actual condition. For the purpose of interpretation, conventional criteria were employed: 0–0.20 mild, 0.21–0.40 fair, 0.41–0.60 moderate, 0.61–0.80 substantial, and ≥ 0.81 almost perfect agreement [109]; see also [71]). To directly compare the two methods, we used McNemar’s χ² test (df = 1) on the paired classifications relative to the criterion variable. A p-value less than .05 indicates different error rates between CAT and full-length scores.
Finally, a Latent Transition Analysis (LTA) was conducted to examine the coherence of latent class assignment between the profiles estimated from the full-length and CAT versions of the HRI-15. LTA is a multi-occasion extension of LPA that enables the evaluation of latent profiles across different occasions and the estimation of transition probabilities between them [95]. In this work, the two considered conceptual “occasions” were the two set of scores (i.e., full-length and CAT scores). In the analysis the indicators were the standardized scores on the five HRI-15 dimensions. First, a configural model was specified, holding the structure and number of profiles constant across the two sets of scores, while allowing means and variances of the indicators to be freely estimated by profile and set of scores. This specification permits level differences between full-length and CAT scores while preserving the conceptual alignment of profiles across occasions [95]. Estimation was performed in Mplus 7.4 [56] using the robust maximum likelihood estimator (MLR; [101]) within a mixture-model framework, with time-2 classes regressed on time-1 classes to obtain the transition matrix (row-standardized transition probabilities from the full-length–based profile to the CAT-based profile). Entropy was used as a classification-quality index (values close to 1 indicate more precise classification), as in the preceding LPA. Concordance between the full-length and CAT solutions at the profile level was also considered. Finally, particular attention was paid to the profile transition probabilities across score type (full-length vs. CAT); low off-diagonal probabilities (i.e., high diagonal stability) were taken as evidence that the CAT procedure yields categorizations comparable to the full-length administration.
Application for adaptive assessment and automatic reporting: A Shiny application [91,110,111] was developed for the adaptive administration of the HRI-15 [11], based on the tree structure derived in the training dataset. Computerized administration follows a deterministic traversal of a pre-populated decision tree. At runtime, two external tables are loaded: (i) a tree sheet that lists the internal nodes, the splitting rules, and the outputs of the terminal leaves (i.e., sum score, subscale z-score, and LPA derived profile); and (ii) a dictionary that maps item codes to the item text and scale to which they belong.
Administration begins at the root (the first non-leaf node not referenced as a child). At each step, the engine presents the item indicated by the current node’s variable, records the response, evaluates the split, and advances to the next item based on the response and tree structure. If a subsequent node requests a variable that has already been answered, the engine automatically advances using the stored response until a new variable is requested or a terminal leaf is reached. The stopping rule is the first terminal leaf encountered. The application then returns leaf-specific predictions: the HRI-15 total score, the z-scores of the five subscales, and the latent class. The results are immediately displayed in a concise, profile-based report. The total score is interpreted based on the ROC-derived decision rule (cutoff ≥ 42), and the dimensional profile is graphically represented to support a two-step workflow in the applied settings: threshold-based triage followed by profile-driven personalization. The system is implemented in Shiny and relies on external files (Excel) for both the tree structure and the dictionary, allowing content to be updated without code changes. Main dependencies include shiny, readxl, dplyr, stringr, and ggplot2, along with custom functions for parsing and matching of path rules. The interface presents standardized instructions and, for each item, and a short guide to interpret results.
Results
Receiver operating characteristic (ROC) analysis
ROC analysis was conducted on the training dataset in order to identify the optimal cut-off score for the instrument. This analysis was then validated on the testing dataset, with the HRI-15 total score serving as the dependent variable and the item on the tendency to self-confine over the past six months as the external criterion (recoded as positive for responses 4–5: “Every week (but not every day)”/ “Every day”). According to this criterion, there were 374 positive cases in the training dataset and 184 in the testing dataset. In the training data set, the area under the ROC curve (AUC) was .846, CI [.826, .865] (Fig 2). The Youden index indicated an optimal cut-off score of ≥ 42. The performance metrics achieved using this cut-off score were .751, .791, and .748 for accuracy, sensitivity, and specificity, respectively. The corresponding 95% confidence intervals were [.740, .762] for accuracy, [.747, .832] for sensitivity, and [.737, .760] for specificity. When the same cut-off score was applied to the testing dataset, the corresponding values were AUC = .837, 95% CI [.809, .865], accuracy = .726, 95% CI [.710, .742], sensitivity = .783, 95% CI [.716, .840], and specificity = .722, 95% CI [.705, .739] (further details are reported in S1 and S2 Tables in the S1 File). In summary, a cut-off score of ≥ 42 showed acceptable ability to identify individuals at risk.
Note. Area under the curve was .846.
A sensitivity analysis using a stricter criterion definition, in which only the response “Every day” was coded as positive, yielded the same optimal integer cut-off score (≥ 42). In the training dataset, the AUC was .819, 95% CI [.780, .858], with accuracy = .725, 95% CI [.713, .736], sensitivity = .771, 95% CI [.685, .843], and specificity = .724, 95% CI [.712, .735]. In the testing dataset, the corresponding values were AUC = .794, 95% CI [.731, .858], accuracy = .700, 95% CI [.683, .717], sensitivity = .764, 95% CI [.630, .868], and specificity = .699, 95% CI [.682, .716]. Overall, the convergence of the primary and sensitivity analyses supported the robustness of the ≥ 42 threshold.
Supplementary calibration analyses based on logistic regression models estimated in the training dataset indicated acceptable agreement between predicted and observed risk probabilities in the testing dataset. In the primary analysis, the Brier score was 0.052, the calibration intercept was −0.087, and the calibration slope was 0.941, indicating only modest deviations from ideal calibration (see Supplementary Tables S3–S5 Tables and S6 Fig in S1 File).
Latent profile analysis
A total of five models, with one to five classes, were estimated. The results were evaluated using a combination of statistical and qualitative criteria (Table 1). The information criteria (i.e., AIC, BIC and SABIC) decreased sharply up to four classes, whereas the improvement from four to five classes was more modest. Entropy was highest for the four-class solution (0.877), indicating clearer classification. The VLMR/LMR tests were significant at the transitions from 2 to 3, 3–4, and 4–5 classes (p < 0.01), indicating continued improvement in fit with increasing model complexity. However, the incremental improvement from four to five classes was comparatively limited, whereas the four-class model yielded the clearest classification and an interpretable configuration with well-defined class proportions (52.4%, 17.5%, 15.0%, and 15.2%). The four-profile solution was therefore retained as the preferred solution, as it offered the best balance among fit improvement, interpretability, parsimony, and classification quality. Additional classification diagnostics for the retained solution, including average posterior probabilities and class-specific parameter estimates, are reported in (S7–S10 Tables in S1 File). Supplementary robustness analyses conducted in the training and testing subsets also supported the stability of the four-profile configuration (S11 Table and S12 Fig in S1 File).
Fig 3 illustrates the profiles derived from the four-class solution, which clearly show distinct profiles. The first group, which is by far the largest, exhibits sub-mean levels across all dimensions, indicating the absence of clinically significant signs of social withdrawal. The second group is primarily characterized by a depressed mood and low energy levels. Social avoidance (agoraphobia and anthropophobia) is a secondary characteristic, while suspicion or distrust is minimally present. This pattern can be regarded as a form of “secondary” withdrawal, where social disengagement is a consequence of pervasive sadness and reduced activity. The third profile most closely aligns with the true hikikomori pattern, which is characterized by social avoidance alongside symptoms such as paranoia, lethargy, and depressed mood, indicating a more pronounced withdrawal syndrome. The fourth group is characterized by anxiety-driven withdrawal, where the main challenges involve social interactions and environmental contexts (agoraphobia and anthropophobia), while depressed mood and low energy levels are comparatively less pronounced. In essence, the identified groups appear to distinguish between individuals without clinical indicators (Class 1, “non-at-risk”); individuals whose withdrawal is predominantly mood-driven (Class 2, “Depressive–lethargic”); individuals exhibiting a characteristic hikikomori pattern (Class 3, “at-risk”); and individuals whose withdrawal is primarily influenced by phobic and social anxiety traits (Class 4, “Social-anxiety/phobic”).
Note. Bars represent standardized scores on Anthropophobia, Agoraphobia, Paranoia, Lethargy, and Depressive Mood for the profiles identified via latent profile analysis (LPA). Positive values indicate scores above the sample means; negative values indicate scores below the sample means. The four profiles are: (1) Non-at-risk, (2) Mood-driven, (3) At-risk, and (4) Social anxiety/phobic. Scores are standardized within-sample (M = 0, SD = 1). Within-class standard deviations were 0.585 for Anthropophobia, 0.547 for Agoraphobia, 0.806 for Paranoia, 0.702 for Lethargy, and 0.499 for Depressive Mood; these values were constrained to be equal across classes in the retained model.
Differences between profiles extended systematically to the other psychological variables and risk behaviors considered. ANOVAs revealed large, significant differences for depressive symptoms (F(3, 2742) = 2496.7, p < .001, η² = .518) and social anxiety (F(3, 2688) = 2665.6, p < .001, η² = .545), as well as medium differences for impulsivity (F(3, 3055) = 301.4, p < .001, η² = .103). As shown in Fig 3, Class 1 (i.e., the non-at-risk profile) displays below-average levels of social anxiety, depression and impulsivity. Class 2 (i.e., the mood profile) shows slightly above-average anxiety and impulsivity, and moderately elevated depression. Class 3 (i.e., the at-risk profile) combines high anxiety and depression with moderately high impulsivity, constituting the highest-risk profile. Class 4 (i.e., the social anxiety profile) is characterized by above-average social anxiety. Overall, this pattern aligns with the expected profile for each class (Fig 4).
Note. Bars represent standardized (z) scores for depression, anxiety, and trait impulsivity across the four profiles identified by latent profile analysis (LPA). Higher values indicate greater symptom severity/impulsivity. Scores are standardized within-sample (M = 0, SD = 1). Error bars show 95% confidence intervals around the mean; substantial overlap suggests non-significant differences. Post hoc pairwise comparisons indicated that all between-class differences were statistically significant (p < .01).
Notable differences across profiles were also observed concerning risky behaviors and habits (Tables 2 and 3). For instance, daily engagement with social media and video games use varied systematically according to the hikikomori risk profile. Regarding gaming, the χ2 test yielded a significant result (χ² = 145, p < .001, V = .074), indicating a small effect size. For social media use, the association was more pronounced (χ² = 580, p < .001, V = .149), indicating a small to medium effect size. Examination of standardized residuals (|z| > 1.96) indicated that profiles characterized by depressed mood and those defined as “at-risk” (Classes 2–3) were more likely to report high daily social media engagement (from ≥ 3 to > 9 hours/day), whereas “non-at-risk” and “social anxiety” profiles (Classes 1 and 4) more frequently reported low levels of use. With respect to gaming, very intensive engagement (7–8 hours/day) was more likely in Class 3.
Participation in social media challenges was weakly but significantly associated with class membership (χ²(3) = 9.92, p < .05, V = .034). Standardized residuals indicated greater participation among individuals categorized on the “depressive” profile (Class 2). Regarding sexting, assessed only in the high-school subsample, small but statistically significant differences emerged across classes (χ²(3) = 81.10, p < .001, V = .126), with the “non-at-risk” group (Class 1) showing a lower propensity to send erotic/sexual content, and the depressive and “at-risk” groups (Classes 2 and 3) showing a higher propensity.
With respect to substance use, a small-to-moderate association was found between alcohol consumption and class membership (χ²(3) = 321, p < .001, V = .192), whereas cannabis use demonstrated a significant but small association (χ²(3) = 77.10, p < .001, V = .123). Standardized residuals indicated that the probability of alcohol use was higher in classes characterized by depressive mood and in the “at-risk” group (Classes 2 and 3) but lower in the “non-risk” and “anxious” profiles (Classes 1 and 4). Cannabis use was more likely in the depressive class (Class 2) and less likely in the anxious-phobic class (Class 4).
Finally, the protective tendency to seek support by talking with parents or caregivers differed markedly across profiles (χ²(3) = 1119, p < .001, V = .357). Standardized residuals indicated that Classes 2, 3, and 4 were more inclined to speak with reference figures, whereas only the “non-at-risk” profile (Class 1) was less inclined to do so.
It is interesting to note that the profile membership from LPA was associated with the risk status both relative to the criterion variable (χ²(3) = 284, p < .001, Cramér’s V = .313) and to the dichotomous classification derived from the ROC cut-off score (χ²(3) = 1833, p < .001, Cramér’s V = .796). The pattern was consistent in the testing dataset and in the total sample. In the testing dataset, standardized residuals (|z| > 1.96) indicated that Class 3 (i.e., “at-risk” profile) was more likely to be positive and to be classified as positive (ps < .001), whereas Class 1 (i.e., “non-at-risk” profile) was more likely to be negative and to be classified as negative (ps < .001). Classes 2 and 4 (i.e., “depressive” and “anxious” profiles) were more likely to be classified as at-risk (ps < .001; in the total sample Class 2 also showed a higher likelihood of being positive according to the criterion, p < .001) (Table 4).
CAT performance
The ctree algorithm was tuned on the training data set by examining a grid of plausible hyperparameter settings for mincriterion, minsplit, and minbucket, taking into account 5-fold cross-validated performance together with item savings and tree complexity. The cross-validation and sensitivity analyses indicated that the configuration mincriterion = .10, minsplit = 10, and minbucket = 10 provided a balanced compromise between predictive accuracy and administration efficiency (detailed tuning and sensitivity results are reported in Tables S13 to S16 in the S1 File).
When refitted on the full training data set with the selected hyperparameter configuration, the ctree algorithm yielded a tree structure with 349 leaves, with item 15 (“I feel a sense of inner emptiness”) at the root. Leaves stored predictions for the HRI-15 total score and the standardized scores on the five subscales; branch depth ranged from 3 to 14 nodes.
This structure was then applied to the independent testing set to drive adaptive administration. The number of administered items ranged from 3 to 14, with an average of 7.03 items, corresponding to a 53.13% saving relative to the full-length version (15 items). A t-test on the number of administered items confirmed that the mean was significantly lower than 15 (t(2888) = −143, p < .001, d = −2.66).
Regarding accuracy, agreement between scores obtained with full administration and those estimated adaptively was high: correlations ranged from .767 to .997, with good to excellent values of ICC (from .751 to .997) and low MAEs (Table 5). In addition, paired-samples t-tests indicated no significant differences between full-length and CAT scores on any scale. Moreover, applying the ROC-based cutoff score to the two sets of score yielded comparable classification performance: CAT scores achieved .73 accuracy, with sensitivity .73, specificity .73 and Cohen’s κ .17 (mild), while full-length scores achieved the same overall accuracy (.73) with sensitivity .78, specificity .72, and κ .18 (mild). Consistent with this, McNemar’s test showed no significant difference in error rates when applying the cutoff of ≥ 42 to CAT or full-length scores, χ²(1) = .194, p = .66 (further details on administration diagnostics, CAT classification, and score-discrepancy analyses between CAT-derived and full-length scores are reported in Tables S17–S23 in the S1 File).
Further evidence for the comparability of CAT and full-length scores came from a latent transition analysis (LTA). The five HRI-15 dimensions served as indicators at two occasions (full-length scores represent the first occasion; and CAT scores the second). A configural LTA model with four profiles per occasion was estimated. For interpretability, CAT-based classes were remapped to align with full-length-based profile meanings based on class-specific means. The results showed excellent classification quality (entropy = .959; AIC = 53,180.38; BIC = 53,568.35; SABIC = 53,361.82; N = 2,889). Profile-level agreement between full-length and CAT profiles was high (diagonal transition probabilities from .927 to .954; M = .943), indicating that the CAT yields categorizations comparable to the full-length administration (Table 6). Profile shapes were substantially superimposable across occasions (Fig 5), supporting conceptual alignment.
Note. Profile means from full-length test (solid) and CAT administration (dashed) across the five HRI dimensions. Four profiles are shown: (1) Non-at-risk, (2) Depressive–lethargic, (3) At-risk, and (4) Social-anxiety/phobic. Lines are plotted in black with distinct markers per profile.
Development of App
The app is provided in the S1 File. It administers the test adaptively and returns results in real time, including: (a) the total score on the general scale, (b) standardized scores for each subdimension, (c) a bar chart displaying the subdimension scores, and (d) an indicator of whether the ROC-based cut-off is exceeded, (e) along with profile membership according to the LPA classification.
The home screen presents instructions; users proceed by tapping “Start test” (Image A, Fig 6). This initiates item administration: respondents select an option, which routes them to the subsequent item (Image B, Fig 6). Instructions remain visible on the page, and a button is available to restart the assessment if needed.
Note. Screens from the adaptive assessment application (S1 File). A) Landing page with standardized instructions and “Start test.” B) Item screen showing a single prompt with five response buttons (1–5; “Not true for me” to “Very true for me”). Items are administered adaptively based on prior answers; instructions remain visible and a “Restart” option is available. C) Results screen displaying the total score, standardized subscale scores (z), LPA-based profile assignment, a ROC–cutoff risk flag, and a bar chart of subscale scores. Subscale scores are standardized within-sample (M = 0, SD = 1). The risk indicator reflects the ROC-derived cutoff (≥ 42). Color coding of the flag is not visible in grayscale of this screen.
Upon completion, a results screen displays the scores, guidance for interpretation, and a bar chart of dimensional scores (Image C, Fig 6).
Discussion
Originally described in Japan, the “hikikomori” phenomenon is now recognized as a relevant issue across cultural boundaries, spreading also in the USA and EU [3,4,7,8,10–14]. The available estimates suggest a prevalence of approximately 1.2% and a marked involvement among younger age groups, with a typical onset in adolescence [17,20,24,35,36]. Given its prevalence and the long-term repercussions on personal development, educational (occupational) trajectories, and mental health, the implementation of early identification and large-scale screening strategies have become imperative [1,16,19,37–41].
To address these challenges, this study aimed to refine the diagnostic categorization of hikikomori risk and to develop and validate an ML-based CAT for its assessment. The HRI-15 was selected as the target instrument because it is brief, developed through international collaboration, psychometrically sound, and multidimensional [11,44]. In particular, the five dimensions of the scale are derived from psychometric studies designed to measure hikikomori symptomatology and its correlates, as theorized by Saitō (1998) [8,44]. These dimensions provide a detailed description of the phenomenon, allowing the instrument to capture all the key aspects of the construct. However, this richness has not yet been fully exploited in applied contexts.
To achieve the research objectives, the study examined a representative national sample of Italian adolescents (N = 8,755) and employed a combination of methodologies: ROC analysis to determine an optimal cut-off on the HRI-15 total score, LPA to delineate multidimensional risk profiles, and ML algorithms to develop the adaptive test.
Concerning ROC analyses, the findings indicated a threshold of ≥ 42 as the optimal cut-off score. This threshold showed adequate discriminative ability (AUC = .85) and an acceptable balance between sensitivity and specificity (from .72 to .79). It is important to note that this threshold is higher than that suggested by previous studies (≥ 37; AUC = .80; [11]). However, the current estimate may be considered more robust, as it was obtained through cross-validation in a larger sample and by using a more detailed external criterion based on the frequency of social withdrawal. In addition, the threshold proved to be stable across sensitivity analyses and showed overall acceptable calibration. Nevertheless, this criterion should be interpreted with appropriate caution. Although it captures a central behavioral manifestation of hikikomori and was therefore suitable for the study’s screening purpose, it was based on a single self-report item rather than on an independent clinical assessment. Consequently, the identified threshold should be evaluated as a pragmatic screening indicator rather than as a definitive diagnostic benchmark. Future research should extend this validation using independent clinical assessments and additional external correlates of functional impairment.
The good performance obtained with ROC analyses suggest that the HRI-15 total score works well in real-world settings as an initial screen. Despite this, risk is not monolithic and depending on a single cut-off may result in the flattening of clinically meaningful heterogeneity, thereby obscuring distinct symptom constellations that could have different implications for support and prevention.
LPA is a person-centered approach that involves the classification of individuals into latent subgroups based on their indicator response patterns, thereby providing a holistic assessment of individual differences that can provide guidance in applied settings [95]. The analysis revealed the presence of four distinct profiles: (1) the non-at-risk group, with below-average levels on all dimensions; (2) the depressive–lethargic group, characterized primarily by depressed mood and low energy, where social avoidance may be secondary and consequent to low mood/energy; (3) the at-risk (hikikomori type) group, with over-average levels on all dimensions (agoraphobia, anthropophobia, paranoia, lethargy, depressed mood); and (4) the social anxiety/phobic group, distinguished by its predominant symptoms of anthropophobia and agoraphobia coupled by smaller elevations in mood/lethargy, a profile that tends to isolate due to difficulties in managing specific situations and interactions. Importantly, these profiles emerged in a sufficiently clear and reproducible manner, as also supported by the supplementary classification diagnostics and split-sample robustness analyses. Overall, these findings make optimal use of the instrument’s multidimensional design, revealing distinct profiles with clear and consistent clinical significance that provide a solid foundation for the development of targeted treatment programs. The validity and clinical utility of the profiles were further enhanced and substantiated by systematic group differences in relation to a series of external variables. In particular, differences emerged in anxiety, depression and impulsivity, with the at-risk profile demonstrating elevated anxiety and depression; the depressive–lethargic profile exhibiting moderate depression and slightly above-average anxiety and impulsivity; and the social anxiety/phobic profile primarily displaying high social anxiety. It is also notable that, as theorized, impulsivity was relatively low across all profiles, confirming theoretical expectations [44,83–86] and supporting the scale’s validity.
Risk behaviors resulted also to be sensibly differentiated across profiles. For example, depressive and at-risk groups exhibited higher daily social media use, while the at-risk profile showed a greater propensity for intensive gaming. Furthermore, depressive and at-risk profiles were associated with a higher incidence of alcohol use and, to a lesser degree, of cannabis use. Overall, these patterns indicate that similar total scores can reflect different risk profiles with distinct correlates and intervention priorities. It is, however, noteworthy that LPA profile membership was associated with risk status both with respect to the external criterion and to the dichotomous classification derived from the ROC cut-off. The “at-risk” profile was more likely to be positive and classified as positive, whereas the “not-at-risk” profile was more likely to be negative and classified as negative. Finally, the depressive” and “anxious” profiles were more likely to be classified as at risk. In practice, the ROC threshold helps decide whom to flag, whereas the profiles clarify what to prioritize clinically.
A central innovation of this work is the development of a multivariate, ML-based CAT that uses a single conditional inference tree (ctree; [102,105,107]) to predict the HRI-15 total score, the five subscale z scores, and LPA profile membership. Simulations using empirical data from the test set demonstrated that the CAT required an average of 7.03 out of 15 items, representing a 53% reduction in items needed to complete the assessment (a significant difference relative to the full-length assessment). Despite this efficiency, the procedure successfully reproduced full-length scores with a high degree of accuracy (r = .77–.997; MAE for total score 3.98, and from 0.03 to .47 for subscale). Paired-sample t tests indicated that CAT scores did not differ significantly from those obtained with the full-length test. In addition, the application of the ROC-based cut-off score to the CAT scores resulted in an accuracy that was comparable to that achieved through the full-length administration, with similar sensitivity/specificity and no differences in error rates. Beyond reproducing scores, LTA showed substantial alignment at the profile level between solutions derived from the full-length and CAT administrations (entropy = .96; diagonal probabilities = .93–.95), indicating that the adaptive procedure preserves the profile structure that gives the instrument its clinical utility.
An additional contribution of the present work is the development of a Shiny-based application for the adaptive administration of the scale, automated scoring, and reporting. From an implementation perspective, the approach is advantageous because it is fast, transparent, and lightweight since the adaptive flow follows a precompiled branch structure, with minimal runtime processing. Reporting is intentionally simple: a total score easily interpretable through the ROC-based thresholds, subscale results expressed as z-scores and visually summarized with bar graphs that improve usability in clinical practice, and automatic indication of the most likely profile for the response pattern. A particular strength of this approach lies in integrating the richness of dimensional measurement with a categorical framework that facilitates practical implementation, yielding a synthesis that is both methodologically current and empirically validated [6,112].
Overall, this work offers four concrete advances: (1) clinically useful classes derived from the HRI-15 dimensions, easily computed and interpreted, with distinct external correlates to guide targeted interventions; (2) a cross-validated threshold (≥ 42) for efficient case identification, derived from a frequency-based external criterion and showing improved discrimination; (3) a feasible, real-time adaptive app that returns the total score, subscales scores, risk flag based on ROC-derived cut-off score, and LPA profile assignment, which minimises burden and errors; (4) a generalizable ML workflow (ctree-based multivariate CAT) that illustrates how tree methods can deliver accurate, interpretable, and scalable assessments, moving beyond binary diagnosis toward continuous scores and profile-based categorisation within a single coherent framework.
In applied settings, the results support a two-step strategy: using the threshold for triage and profiles for personalization. Identification of an at-risk profile should trigger a focused approach. Management priorities should include mood and activation for the depressive–lethargic profile, guided exposure and skills training for the social anxiety/phobia profile, and multimodal support addressing withdrawal, mood, suspiciousness and digital use habits for the at-risk profile. In all cases, monitoring of risk behaviours relevant to the profile (e.g., alcohol, cannabis, gaming) is advised. The app operationalizes this flow by returning, at the end of the assessment, a concise, profile-informed report that can be integrated into school, clinical, or epidemiological pathways.
Several limitations warrant caution. First, although the external criterion for ROC analysis improves upon prior dichotomous items by leveraging frequency, it remains a single self-report indicator; future studies should incorporate multi-informant information and clinical interviews. This may have introduced some degree of criterion misclassification, with possible consequences for sensitivity, specificity, and the location of the optimal cutoff. Future studies should incorporate multi-informant information, clinical interviews, and additional external indicators of functional impairment. Second, although the sample is large, it is exclusively Italian and based on school data. Therefore, further validation using samples from other cultures and clinical populations are strongly advocated. In particular, it should be noted that hikikomori is a culturally influenced phenomenon; consequently, the generalizability of the current results to cross-cultural samples appears to be of central relevance. In this regard, some studies have shown that incorporating additional covariates into adaptive models based on tree structures derived from binary partitioning can improve the performance and personalization of the obtained assessment procedures [34]; consequently, future studies could investigate whether socio-cultural variables, gender or other characteristics can enhance the performance of the methodology. Third, although the CAT reproduced scores and profiles well, further studies should evaluate its performance in real-world adaptive administration (e.g., missingness, device variability) and fairness across key subgroups [11]. Future studies should also examine the longitudinal validity of the CAT by assessing the temporal stability of CAT-derived classifications, its sensitivity to change over time, and its performance when embedded in routine school- or clinic-based screening workflows. Future work should evaluate the utility of more advanced computerized therapy or self-monitoring applications that integrate online therapy with data from wearable devices, enabling targeted and timely support. Early identification of hikikomori risk requires tools that are not only valid but also rapid, interpretable, and scalable. By pairing a cross-validated threshold with clinically meaningful profiles, and delivering both via a ctree-based multivariate CAT integrated into an easy-to-deploy app, this work shows that administration time can be halved while preserving accuracy, maintaining alignment with full-length classifications and profiles, and returning immediately actionable information. The approach clarifies how ML can be leveraged, without sacrificing transparency, to strengthen diagnostic assessment and prevention efforts for prolonged social withdrawal. This is consistent with findings indicating that ML is particularly effective for CAT with respect to efficiency, longitudinal use, informativeness, accuracy, and personalization [34,46,52–54,67,68].
References
- 1. Matsushita Y, Yasumatsu N. A Sociological study of Hikikomori. J Human Soc Sci. 2025;7(1):1–9.
- 2. Tateno M, Teo AR, Ukai W, Kanazawa J, Katsuki R, Kubo H, et al. Internet addiction, smartphone addiction, and hikikomori trait in japanese young adult: social isolation and social network. Front Psychiatry. 2019;10:455. pmid:31354537
- 3. Kato TA, Sartorius N, Shinfuku N. Shifting the paradigm of social withdrawal: a new era of coexisting pathological and non-pathological hikikomori. Curr Opin Psychiatry. 2024;37(3):177–84. pmid:38415743
- 4. Teo AR, Gaw AC. Hikikomori, a Japanese culture-bound syndrome of social withdrawal?: a proposal for DSM-5. J Nerv Ment Dis. 2010;198(6):444–9. pmid:20531124
- 5. Zhang W, Chen M-Y, Feng Y, Su Z, Cheung T, Jackson T, et al. Epidemiology of Hikikomori: a systematic review and meta-analysis of 19 studies. Psychiatry Clin Neurosci. 2025;79(4):138–46. pmid:40016086
- 6. American Psychiatric Association. Diagnostic and statistical manual of mental disorders. 2022.
- 7. Bowker JC, Bowker MH, Santo JB, Ojo AA, Etkin RG, Raja R. Severe social withdrawal: cultural variation in past hikikomori experiences of university students in Nigeria, Singapore, and the United States. J Genet Psychol. 2019;180(4–5):217–30. pmid:31305235
- 8.
Saitō T. Shakaiteki hikikomori: owaranai shishunki. PHP Kenkyūjo; 1998.
- 9. Teo AR, Fetters MD, Stufflebam K, Tateno M, Balhara Y, Choi TY, et al. Identification of the hikikomori syndrome of social withdrawal: psychosocial features and treatment preferences in four countries. Int J Soc Psychiatry. 2015;61(1):64–72. pmid:24869848
- 10. Adamski D. The influence of new technologies on the social withdrawal (Hikikomori Syndrome) among developed communities, including Poland. Soc Commun. 2018;17(1):58–63.
- 11. Colledani D, Anselmi P, Monacis L, Robusto E, Genetti B, Andreotti A, et al. Validation of the HRI-24 on Adolescents and Development of a Short Version of the Instrument. Int J Ment Health Addiction. 2024;22(6):4072–89.
- 12. Correia Lopes F, Pinto da Costa M, Fernandez-Lazaro CI, Lara-Abelenda FJ, Pereira-Sanchez V, Teo AR, et al. Analysis of the hikikomori phenomenon - an international infodemiology study of Twitter data in Portuguese. BMC Public Health. 2024;24(1):518. pmid:38373925
- 13. Furuhashi T, Tsuda H, Ogawa T, Suzuki K, Shimizu M, Teruyama J. Current situation, commonalities and differences between socially withdrawn young adults (Hikikomori) in France and Japan. Evolution Psychiatrique. 2013;78.
- 14. Gorman G, Bacon A, May J, Minton S. Hikikomori risk in the UK. Int J Social Psych. 2025;71.
- 15. Kanai K, Kitamura Y, Zha L, Tanaka K, Ikeda M, Sobue T. Prevalence of and factors influencing Hikikomori in Osaka City, Japan: A population-based cross-sectional study. Int J Soc Psychiatry. 2024;70(5):967–80. pmid:38616515
- 16. Malagón-Amor Á, Córcoles-Martínez D, Martín-López LM, Pérez-Solà V. Hikikomori in Spain: a descriptive study. Int J Soc Psychiatry. 2015;61(5):475–83. pmid:25303955
- 17. Miriam N, Karina Z, Annamária Š. Prevalence of the Hikikomori Syndrome in the Context of Internet Addictive Behaviour Among Primary School Pupils in the Slovak Republic. TEM Journal. 2024;:476–4483.
- 18. Nonaka S, Takeda T, Sakai M. Who are hikikomori? Demographic and clinical features of hikikomori (prolonged social withdrawal): A systematic review. Aust N Z J Psychiatry. 2022;56(12):1542–54. pmid:35332798
- 19. Pozza A, Coluccia A, Kato T, Gaetani M, Ferretti F. The “Hikikomori” syndrome: worldwide prevalence and co-occurring major psychiatric disorders: a systematic review and meta-analysis protocol. BMJ Open. 2019;9(9):e025213. pmid:31542731
- 20. Hamasaki Y, Pionnié-Dax N, Dorard G, Tajan N, Hikida T. Identifying social withdrawal (Hikikomori) factors in adolescents: understanding the hikikomori spectrum. Child Psychiatry Hum Dev. 2021;52(5):808–17. pmid:32959142
- 21. Kubo H, Katsuki R, Horie K, Yamakawa I, Tateno M, Shinfuku N, et al. Risk factors of hikikomori among office workers during the COVID-19 pandemic: a prospective online survey. Curr Psychol. 2022;:1–19. pmid:35919757
- 22. Ogawa T, Shiratori Y, Midorikawa H, Aiba M, Sugawara D, Kawakami N, et al. A survey of changes in the psychological state of individuals with social withdrawal (hikikomori) in the context of the COVID pandemic. COVID. 2023;3(8):1158–72.
- 23. Kato TA, Kanba S, Teo AR. Hikikomori : multidimensional understanding, assessment, and future international perspectives. Psychiatry Clin Neurosci. 2019;73(8):427–40. pmid:31148350
- 24. Koyama A, Miyake Y, Kawakami N, Tsuchiya M, Tachimori H, Takeshima T, et al. Lifetime prevalence, psychiatric comorbidity and demographic correlates of “hikikomori” in a community population in Japan. Psychiatry Res. 2010;176(1):69–74. pmid:20074814
- 25. Fong TC, Yip PS. Prevalence of hikikomori and associations with suicidal ideation, suicide stigma, and help-seeking among 2,022 young adults in Hong Kong. Int J Soc Psychiatry. 2023;69(7):1768–80. pmid:37191282
- 26. Fong TCT, Cheng Q, Pai CY, Kwan I, Wong C, Cheung S-H, et al. Uncovering sample heterogeneity in gaming and social withdrawal behaviors in adolescent and young adult gamers in Hong Kong. Soc Sci Med. 2023;321:115774. pmid:36796169
- 27. Hu X, Fan D, Shao Y. Social Withdrawal (Hikikomori) conditions in China: a cross-sectional online survey. Front Psychol. 2022;13:826945. pmid:35360625
- 28. Liu LL, Li TM, Teo AR, Kato TA, Wong PW. Harnessing social media to explore youth social withdrawal in three major cities in China: cross-sectional web survey. JMIR Ment Health. 2018;5(2):e34. pmid:29748164
- 29. Wong PWC, Li TMH, Chan M, Law YW, Chau M, Cheng C, et al. The prevalence and correlates of severe social withdrawal (hikikomori) in Hong Kong: a cross-sectional telephone-based survey study. Int J Soc Psychiatry. 2015;61(4):330–42. pmid:25063752
- 30. Baek S-U, Yoon J-H. Prolonged social withdrawal (“hikikomori”) and its associations with depressive symptoms and suicidal ideation among young adults in Korea: findings from the 2022 Youth Life Survey. J Affect Disord. 2025;381:514–7. pmid:40203969
- 31. Amendola S, Cerutti R, von Wyl A. Estimating the prevalence and characteristics of people in severe social isolation in 29 European countries: a secondary analysis of data from the European Social Survey round 9 (2018-2020). PLoS One. 2023;18(9):e0291341. pmid:37699030
- 32. Amendola S, Presaghi F, Teo AR, Cerutti R. Psychometric properties of the italian version of the 25-item hikikomori questionnaire. Int J Environ Res Public Health. 2022;19(20):13552. pmid:36294128
- 33. Al-Sibani N, Chan MF, Al-Huseini S, Al Kharusi N, Guillemin GJ, Al-Abri M, et al. Exploring hikikomori-like idiom of distress a year into the SARS-CoV-2 pandemic in Oman: factorial validity of the 25-item hikikomori questionnaire, prevalence and associated factors. PLoS One. 2023;18(8):e0279612. pmid:37549148
- 34. Colledani D, Robusto E, Anselmi P. Shortening and personalizing psychodiagnostic assessments with decision tree-machine learning classifiers: an application example based on the patient health questionnaire-9. Int J Ment Health Addict. 2024;24:1004–24.
- 35. Yamazaki S, Ura C, Okamura T. Time regained: awareness of not young anymore is a trigger for the ageing Hikikomori person to return to society. Psychogeriatrics. 2023.
- 36. Yamazaki S, Ura C, Shimmei M, Okamura T. In search of lost time: long-term prognosis of hikikomori called 8050 crisis. Int J Geriatr Psychiatry. 2021;36(10):1590–1. pmid:34058030
- 37. Frankova I. Similar but different: psychological and psychopathological features of primary and secondary hikikomori. Front Psychiatry. 2019;10:558. pmid:31447713
- 38. Katsuki R, Tateno M, Kubo H, Kurahara K, Hayakawa K, Kuwano N, et al. Autism spectrum conditions in hikikomori: a pilot case-control study. Psychiatry Clin Neurosci. 2020;74(12):652–8. pmid:32940406
- 39. Kondo N, Sakai M, Kuroda Y, Kiyota Y, Kitabata Y, Kurosawa M. General condition of hikikomori (prolonged social withdrawal) in Japan: psychiatric diagnosis and outcome in mental health welfare centres. Int J Soc Psychiatry. 2013;59(1):79–86. pmid:22094722
- 40. Teo AR, Nelson S, Strange W, Kubo H, Katsuki R, Kurahara K, et al. Social withdrawal in major depressive disorder: a case-control study of hikikomori in japan. J Affect Disord. 2020;274:1142–6. pmid:32663943
- 41. Yasuma N, Watanabe K, Nishi D, Ishikawa H, Tachimori H, Takeshima T, et al. Psychotic experiences and hikikomori in a nationally representative sample of adult community residents in japan: a cross-sectional study. Front Psychiatry. 2021;11:602678. pmid:33584370
- 42. Uchida Y, Norasakkunkit V. The NEET and hikikomori spectrum: assessing the risks and consequences of becoming culturally marginalized. Front Psychol. 2015;6:1117. pmid:26347667
- 43. Teo AR, Chen JI, Kubo H, Katsuki R, Sato-Kasai M, Shimokawa N, et al. Development and validation of the 25-item Hikikomori Questionnaire (HQ-25). Psychiatry Clin Neurosci. 2018;72(10):780–8. pmid:29926525
- 44. Loscalzo Y, Nannicini C, Huai-Ching Liu I-T, Giannini M. Hikikomori Risk Inventory (HRI-24): a new instrument for evaluating Hikikomori in both Eastern and Western countries. Int J Soc Psychiatry. 2022;68(1):90–107. pmid:33238782
- 45. Gibbons RD, Weiss DJ, Kupfer DJ, Frank E, Fagiolini A, Grochocinski VJ, et al. Using computerized adaptive testing to reduce the burden of mental health assessment. Psychiatr Serv. 2008;59(4):361–8. pmid:18378832
- 46. Gibbons RD, Weiss DJ, Frank E, Kupfer D. Computerized adaptive diagnosis and testing of mental health disorders. Annu Rev Clin Psychol. 2016;12:83–104. pmid:26651865
- 47. McHorney CA, Tarlov AR. Individual-patient monitoring in clinical practice: are available health status surveys adequate?. Qual Life Res. 1995;4(4):293–307.
- 48. Rose M, Bjorner JB, Becker J, Fries JF, Ware JE. Evaluation of a preliminary physical function item bank supported the expected advantages of the Patient-Reported Outcomes Measurement Information System (PROMIS). J Clin Epidemiol. 2008;61(1):17–33. pmid:18083459
- 49. Gass C. MMPI-2: a practitioner’s guide. Choice Reviews Online. 2006;43.
- 50. Gibbons RD, deGruy FV. Without wasting a word: extreme improvements in efficiency and accuracy using computerized adaptive testing for mental health disorders (CAT-MH). Curr Psychiatry Rep. 2019;21(8):67. pmid:31264098
- 51. Graham AK, Minc A, Staab E, Beiser DG, Gibbons RD, Laiteerapong N. Validation of the computerized adaptive test for mental health in primary care. Ann Fam Med. 2019;17(1):23–30. pmid:30670391
- 52. Colledani D, Anselmi P, Robusto E. Machine learning-decision tree classifiers in psychiatric assessment: an application to the diagnosis of major depressive disorder. Psychiatry Res. 2023;322:115127. pmid:36842398
- 53. Colledani D, Barbaranelli C, Anselmi P. Fast, smart, and adaptive: using machine learning to optimize mental health assessment and monitor change over time. Sci Rep. 2025;15(1):6492. pmid:39987277
- 54. Gibbons RD, Chattopadhyay I, Meltzer HY, Kane JM, Guinart D. Development of a computerized adaptive diagnostic screening tool for psychosis. Schizophr Res. 2022;245:116–21. pmid:33836922
- 55. Zhou XH, Obuchowski NA, McClish DK. Statistical methods in diagnostic medicine. 2011.
- 56. Muthén L, Muthén B. Mplus version 7 user’s guide. Los Angeles, CA: Muthén & Muthén; 2012.
- 57. Nylund KL, Asparouhov T, Muthén BO. Deciding on the number of classes in latent class analysis and growth mixture modeling: a monte carlo simulation study. Struct Eq Mode A Multidiscip J. 2007;14(4):535–69.
- 58.
Muthén BO, Asparouhov T. Latent variable analysis with categorical outcomes: multiple-group and growth modeling in mplus. 2002.
- 59. Nylund-Gibson K, Choi AY. Ten frequently asked questions about latent class analysis. Transl Issues in Psychol Sci. 2018;4(4):440–61.
- 60. Thompson NA. Item selection in computerized classification testing. Educ Psychol Meas. 2009;69.
- 61. Van der Linden WJ, Glas CA. Computerized adaptive testing: theory and practice. 2000.
- 62. Cella D, Yount S, Rothrock N, Gershon R, Cook K, Reeve B, et al. The Patient-Reported Outcomes Measurement Information System (PROMIS): progress of an NIH Roadmap cooperative group during its first two years. Med Care. 2007;45(5 Suppl 1):S3–11. pmid:17443116
- 63.
Kubiszyn T, Borich GD. Educational testing and measurement. Wiley & Sons J; 2024.
- 64. Pilkonis PA, Choi SW, Reise SP, Stover AM, Riley WT, Cella D, et al. Item banks for measuring emotional distress from the Patient-Reported Outcomes Measurement Information System (PROMIS®): depression, anxiety, and anger. Assessment. 2011;18(3):263–83. pmid:21697139
- 65. Reise SP, Waller NG. Item response theory and clinical measurement. Annu Rev Clin Psychol. 2009;5:27–48. pmid:18976138
- 66. Simms LJ, Clark LA. Validation of a computerized adaptive version of the Schedule for Nonadaptive and Adaptive Personality (SNAP). Psychol Assess. 2005;17(1):28–43. pmid:15769226
- 67. Colledani D, Robusto E, Anselmi P. Machine learning–driven adaptive testing: an application for the MMPI assessment. Hum Behav Emerg Technol. 2025;2025(1):5146188.
- 68. Delgado-Gómez D, Laria JC, Ruiz-Hernández D. Computerized adaptive test and decision trees: a unifying approach. Expert Syst Appl. 2019;117:358–66.
- 69. Gonzalez O. Psychometric and machine learning approaches to reduce the length of scales. Multivariate Behav Res. 2021;56(6):903–19. pmid:32749158
- 70.
Breiman L, Friedman J, Olshen RA, Stone CJ. Classification and Regression Trees. 1st ed. Chapman and Hall/CRC; 1984.
- 71. Witten IH, Frank E, Hall MA, Pal CJ. Data Mining: practical machine learning tools and techniques. 2016.
- 72. Criminisi A, Shotton J, Konukoglu E. Decision forests: a unified framework for classification, regression, density estimation, manifold learning and semi-supervised learning. Foundations and trends® in computer graphics and vision. 2012;7(2–3):81–227.
- 73. Gupta B, Rawat A, Jain A, Arora A, Dhami N. Analysis of various decision tree algorithms for classification in data mining. Int J Comput Appl. 2017;163.
- 74. Gray RM. Entropy and Information. Entropy and information theory. Springer New York; 1990. 21–55.
- 75. Mahesh B. Machine learning algorithms—a review. Int J Sci Res. 2020;9:381–6.
- 76. Yarkoni T, Westfall J. Choosing prediction over explanation in psychology: lessons from machine learning. Perspect Psychol Sci. 2017;12(6):1100–22. pmid:28841086
- 77. Gonzalez O. Psychometric and machine learning approaches for diagnostic assessment and tests of individual classification. Psychol Methods. 2021;26(2):236–54. pmid:32614196
- 78. McArdle JJ. Adaptive testing of the number series test using standard approaches and a new decision tree analysis approach. Contemporary issues in exploratory data mining in the behavioral sciences. 2020.
- 79. Lu F, Petkova E. A comparative study of variable selection methods in the context of developing psychiatric screening instruments. Stat Med. 2014;33(3):401–21. pmid:23934941
- 80. Yan D, Lewis C, Stocking M. Adaptive testing with regression trees in the presence of multidimensionality. J Edu Behav Stat. 2004;29(3):293–316.
- 81.
Fossati A, Borroni S, Del Corno F. Scala di valutazione della gravità del disturbo d’ansia sociale (fobia sociale) – Soggetto da 11 a 17 anni. Raffaello Cortina Editore, editor; 2015.
- 82.
Fossati A, Borroni S, Del Corno F. Scala di valutazione della gravità della depressione – soggetto da 11 a 17 anni. Raffaello Cortina Editore; 2015.
- 83. Adams ZW, Kaiser AJ, Lynam DR, Charnigo RJ, Milich R. Drinking motives as mediators of the impulsivity-substance use relation: pathways for negative urgency, lack of premeditation, and sensation seeking. Addict Behav. 2012;37(7):848–55. pmid:22472524
- 84. Giannini M, Loscalzo Y. Sensation seeking. The Wiley encyclopedia of personality and individual differences. Wiley; 2020. 411–5.
- 85. Pokhrel P, Sussman S, Sun P, Kniazer V, Masagutov R. Social self-control, sensation seeking and substance use in samples of US and Russian adolescents. Am J Health Behav. 2010;34(3):374–84. pmid:20001194
- 86. Stautz K, Cooper A. Impulsivity-related personality traits and adolescent alcohol use: a meta-analytic review. Clin Psychol Rev. 2013;33(4):574–92. pmid:23563081
- 87. Maggi G, Altieri M, Ilardi CR, Santangelo G. Validation of a short Italian version of the Barratt Impulsiveness Scale (BIS-15) in non-clinical subjects: psychometric properties and normative data. Neurol Sci. 2022;43(8):4719–27. pmid:35403939
- 88. Fossati A, Di Ceglie A, Acquarini E, Barratt ES. Psychometric properties of an Italian version of the Barratt Impulsiveness Scale-11 (BIS-11) in nonclinical subjects. J Clin Psychol. 2001;57(6):815–28. pmid:11344467
- 89. Fossati A, Barratt ES, Acquarini E, Di Ceglie A. Psychometric properties of an adolescent version of the Barratt Impulsiveness Scale-11 for a sample of Italian high school students. Percept Mot Skills. 2002;95(2):621–35. pmid:12434861
- 90. Meule A. Cut-off scores for the Barratt Impulsiveness Scale-short form (BIS-15): sense and nonsense. Int J Neurosci. 2024;134(10):1149–52. pmid:37486098
- 91.
R Core Team. R: a language and environment for statistical computing. R Foundation for Statistical Computing; 2025.
- 92. Robin X, Turck N, Hainard A, Tiberti N, Lisacek F, Sanchez J-C, et al. pROC: an open-source package for R and S+ to analyze and compare ROC curves. BMC Bioinformatics. 2011;12:77. pmid:21414208
- 93. Brier GW. Verification of forecasts expressed in terms of probability. Mon Weather Rev. 1950;78:1–3.
- 94. Miller ME, Langefeld CD, Tierney WM, Hui SL, McDonald CJ. Validation of probabilistic predictions. Medical Decision Making. 1993;13:49–57.
- 95. Collins LM, Lanza ST. Latent class and latent transition analysis: with applications in the social, behavioral, and health sciences. 2010.
- 96. Dal Corso L, De Carlo A, Carluccio F, Colledani D, Falco A. Employee burnout and positive dimensions of well-being: a latent workplace spirituality profile analysis. PLoS One. 2020;15(11):e0242267. pmid:33201895
- 97. Ferguson SL, Hull DM. Personality profiles: using latent profile analysis to model personality typologies. Pers Individ Dif. 2018;122.
- 98. Meneghini AM, Colledani D, Morandini S, De France K, Hollenstein T. Emotional engagement and caring relationships: The assessment of emotion regulation repertoires of nurses. Psychol Rep. 2024;127(1):212–34.
- 99. Spurk D, Hirschi A, Wang M, Valero D, Kauffeld S. Latent profile analysis: a review and “how to” guide of its application within vocational behavior research. J Vocat Behav. 2020;120:103445.
- 100. Lo Y. Testing the number of components in a normal mixture. Biometrika. 2001;88(3):767–78.
- 101. Yuan KH, Bentler PM. Three likelihood-based methods for mean and covariance structure analysis with nonnormal missing data. Sociol Methodol. 2000;30.
- 102. Zeileis A, Hothorn T, Hornik K. Model-based recursive partitioning. J Comput Graph Stat. 2008;17(2):492–514.
- 103. Strobl C, Boulesteix A-L, Zeileis A, Hothorn T. Bias in random forest variable importance measures: illustrations, sources and a solution. BMC Bioinform. 2007;8:25. pmid:17254353
- 104. Hothorn T, Zeileis A. Partykit: a modular toolkit for recursive partitioning in R. J Mach Learn Res. 2015;16.
- 105.
Hothorn T, Hornik K, Wien W, Zeileis A. Ctree: Conditional inference trees. The Comprehensive R Archive Network; 2015.
- 106. Hothorn T, Hornik K, Zeileis A. Unbiased recursive partitioning: a conditional inference framework. J Comput Graph Stat. 2012;15.
- 107. Zeileis A, Hothorn T, Hornik K. Party with the mob: model-based recursive partitioning in R. 2006.
- 108. Koo TK, Li MY. A Guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15(2):155–63. pmid:27330520
- 109. Landis JR, Koch GG. An application of hierarchical kappa-type statistics in the assessment of majority agreement among multiple observers. Biometrics. 1977;33(2):363–74. pmid:884196
- 110. Chang W, Cheng J, Allaire J, Sievert C, Schloerke B, Xie Y. Shiny: web application framework for R.
- 111.
Wickham H. Mastering shiny. O’Reilly Media, Inc.; 2019.
- 112. Kraemer HC, Noda A, O’Hara R. Categorical versus dimensional approaches to diagnosis: methodological challenges. J Psychiatr Res. 2004;38(1):17–25. pmid:14690767