Figures
Abstract
The SF-12v2 is widely used to measure health-related quality of life (HRQoL) in the general Thai population because it is brief and places a low burden on respondents. However, it does not provide utility scores required for economic analyses. The mapping approach is a solution to estimate the utility scores from the SF-12v2 response using several regression models. Therefore, this study aimed to develop a mapping algorithm to estimate EQ-5D-5L utility scores from the SF-12v2 items using a 2022 national dataset of 2000 Thai respondents. Four predictor sets incorporating SF-12v2 items/subscales and age as a covariate were investigated using direct and indirect mapping approaches. Direct mapping approaches included ordinary least squares, Tobit, censored least absolute deviations, generalized linear model (GLM), two-part models, adjusted limited dependent variable mixture model (ALDVMM), and beta mixture model, while multinomial logistic regression (MLOGIT) was investigated for indirect mapping. Model performance was evaluated using ten-fold cross-validated (post-CV) mean absolute error (MAE) and root mean square error (RMSE). The ALDVMM-1 component with a predictor set comprising age and selected SF-12 items as categorical variables demonstrated the best predictive performance, yielding the lowest post-CV MAE and RMSE (0.0410 and 0.0625, respectively). However, MLOGIT showed poorer predictive performance. Therefore, the proposed ALDVMM-1 component-based mapping algorithm may facilitate the estimation of utility scores in studies where only the SF-12v2 was collected. Given that the mapping algorithm was developed from a predominantly healthy sample, it may be less reliable in populations with poorer health status, and further validation in populations with less healthy or clinical samples is warranted in future studies.
Citation: Kangwanrattanakul K (2026) Direct and indirect mapping of the 12-item Short Form Survey version 2 (SF-12v2) onto the EQ-5D-5L utility scores in general Thai population. PLoS One 21(6): e0351064. https://doi.org/10.1371/journal.pone.0351064
Editor: Chintal H. Shah, AstraZeneca Pharmaceuticals LP, UNITED STATES OF AMERICA
Received: March 27, 2026; Accepted: May 20, 2026; Published: June 22, 2026
Copyright: © 2026 Krittaphas Kangwanrattanakul. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: All data supporting the findings of this study are fully available without restriction. The anonymized dataset used for analysis has been uploaded as a Supporting Information file (S1_File.xlsx). This file includes all raw data with all variables necessary to replicate the analyses. Analytical methods are fully described in the manuscript. No ethical or legal restrictions prevent data sharing. For any additional inquiries, the corresponding authors can be contacted, but the data are already publicly accessible via the journal’s supplementary materials.
Funding: This study was financially supported by the Research Grant of Faculty of Pharmaceutical Sciences, Burapha University (Grant No. Rx6/2566). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.
Competing interests: The author declares no competing financial interests. For transparency, some suggested reviewers have previously collaborated with the author in previous publications within the past two years and are affiliated with the same institution.
Introduction
Health technology assessment (HTA) is widely used to evaluate and identify the most cost-effective health interventions to support decision-making in healthcare [1]. It also optimizes the allocation of limited health resources [2,3]. Several HTA guidelines recommend cost-utility analysis to guide clinical decisions and policy implementation [4–6]. This approach compares incremental health improvement, expressed in quality-adjusted life years (QALYs), across or within health conditions [7]. QALYs are calculated using utility scores and life expectancy, making accurate measurement of utility scores essential [7,8].
Utility scores range from 0 (death or worst health state) to 1 (perfect health) [9]. Preference-based instruments, such as the EuroQol 5-dimension (EQ-5D), Health Utility Index (HUI), and Short Form 6 Dimension (SF-6D), measure these scores, which are then converted using a country-specific value set [10,11]. The EQ-5D-5L is the most commonly used instrument in Thailand. However, it may not capture some dimensions considered important, such as social health [12,13], and shows limited responsiveness to changes in both general [14] and patient populations [15].
Although preference-based instruments are standard for economic analyses, they do not provide dimension-specific (health profile) scores. Health profile instruments, such as the World Health Organization Quality of Life: Brief Version (WHOQOL-BREF) and the 12-item Short Form Survey version 2 (SF-12v2), offer more sensitivity to health changes [16,17]. However, they do not generate a utility score. Mapping between preference-based and profile instruments allows retention of profile information while producing utility scores for economic analyses.
Several Thai studies have mapped WHOQOL-BREF to EQ-5D-5L [18,19]. Although WHOQOL-BREF captures physical, psychological, social, and environmental health dimensions, its 26 items may still impose a respondent burden. The SF-12v2, derived from SF-36, is shorter, less burdensome, and widely used in large population surveys, including in Thailand [20–22]. Its conceptual overlap with EQ-5D-5L enhances suitability for mapping algorithms [14].
Despite its widespread use, SF-12v2 lacks a Thai-specific value set for producing utility scores. Mapping approaches can predict EQ-5D-5L utility scores from SF-12v2 responses using regression models [23,24]. Previous studies mapped SF-12v2 to the earlier EQ-5D-3L [25–31], which has a higher ceiling effect and lower discriminative power than EQ-5D-5L [32]. Ordinary least squares (OLS) models were sometimes used [26,27,29], but they are suboptimal for bounded utility scores with high ceiling effects and predict the utility scores outside the feasible range [33,34]. Alternative models, including Tobit, Beta mixture model (Betamix), Censored Least Absolute Deviation (CLAD), Generalized linear model (GLM), Adjusted Limited Dependent Variable Mixture Model (ALDVMM), and two-part models (TPM), better accommodate utility score characteristics [34–37].
To date, no Thai studies have mapped SF-12v2 to EQ-5D-5L using these models, nor performed indirect (response-based) mappings. This study aims to develop a mapping algorithm between SF-12v2 and EQ-5D-5L, including indirect mapping based on SF-12v2 responses in the general Thai population.
Methods
Study design and samples
This study used data from the project “EQ-5D-3L and EQ-5D-5L population norms in Thailand [38].” A cross-sectional survey was conducted using face-to-face interviews with 2,000 Thai adults. To ensure national representation, a four-stage stratified random sampling method was applied to select provinces, districts, sub-districts, and villages. Data were collected from 12 provinces, representing all geographic regions of Thailand: Bangkok, Samut Prakan, Nonthaburi, Chonburi, Nakhon Pathom, Chiang Mai, Nakhon Sawan, Nakhon Ratchasima, Khon Kaen, Buriram, Nakhon Si Thammarat, and Phatthalung.
Participants were eligible if they (1) were aged 18 years or older, (2) were able to understand the interview process as assessed by the researcher or trained interviewers, and (3) were fluent in Thai. Individuals with acute or life-threatening illnesses or severe cognitive impairments were excluded.
Data collection procedures
Trained interviewers read all questions and response options verbatim, and did not provide explanations or interpretations to minimize interviewer bias. All participants completed both the SF-12v2 and EQ-5D-5L questionnaires, with no missing responses. These face-to-face interviews with eligible participants were conducted at their home residences between 1st May to 30th June 2023.
Before data collection, participants received an information sheet outlining study objectives, procedures, and their rights to withdraw without consequences. Written informed consent was obtained from all participants. Questionnaires were administered in the following order: (1) demographic information, (2) EQ-5D-5L, and (3) SF-12v2. Details of the demographic variables are reported elsewhere [38]. The protocol and data collection process were approved by the ethical committees of Burapha University Institutional Review Board (IRB1–031/2566) and adhered to the principles outlined in the Declaration of Helsinki.
Source and target measures
SF-12v2 (source measure).
The SF-12v2 consists of 12 items grouped into eight subscales: physical functioning (PF: 2 items), role limitations due to physical problems (RP, 2 items), bodily pain (BP, 1 item), general health (GH, 1 item), vitality (VT, 1 item), social functioning (SF, 1 item), role limitations due to emotional problems (RE, 2 items), and mental health (MH, 2 items) [39]. Participants rated their health over a four-week recall period. Most items use a five-point Likert scale, except PF items, which use a three-point scale. Four items (GH1, BP1, VT1, MH1) were reverse-coded to ensure scoring consistency. Table 1 displays descriptions of the SF-12v2 items.
Subscale scores were transformed to a 0–100 scale, with higher scores indicating better health [39]. The Physical Component Summary (PCS) and Mental Component Summary (MCS) scores were calculated using a proprietary scoring algorithm to generate norm-based scores. SF-12v2 responses were also reclassified to derive SF-6D health states, allowing utility score estimation using the UK algorithm, with scores ranging from 0.29 to 1.00 [40].
EQ-5D-5L (target measure).
The EQ-5D-5L comprises five dimensions: mobility, self-care, usual activities, pain/discomfort, and anxiety/depression. Each dimension is assessed across five ordered severity levels: no problems, slight problems, moderate problems, severe problems, extreme problems/unable to perform. Responses across the five dimensions generate a five-digit health-state profile that characterizes an individual’s health status. These profiles are subsequently converted into utility scores using country-specific valuation algorithms.
In this study, EQ-5D-5L utility scores were derived using the Thai-specific value set, with values ranging from −0.4212, corresponding to the worst health state (55555), to 1.0000 representing full health (11111). The second-best health state (11121) is valued at 0.9436 [11].
The WHOQOL-BREF instrument was additionally employed to compute overall quality-of-life scores by summing four domain scores: physical health (7 items), psychological health (6 items), social relationships (3 items), and environment (8 items), using the ordinal-to-interval conversion table for general Thai population [41]. WHOQOL-BREF scores range from 24 to 120, with higher scores indicating better perceived health-related quality of life. These total scores were used to stratify participants and to plot mean EQ-5D-5L utility scores, facilitating comparisons between predicted and observed utility values.
Conceptual overlap between EQ-5D-5L and SF-12v2
Conceptual overlap between the EQ-5D-5L and the SF-12v2 was examined using Spearman’s rank-order correlation analysis. Pairwise correlation coefficients (r) were calculated between each SF-12v2 item and subscale score and each EQ-5D-5L dimension, as well as the overall EQ-5D-5L utility score.
An SF-12v2 item or subscale demonstrating at least a moderate correlation (r 0.4) with any EQ-5D-5L dimension or the utility scores was deemed to exhibit conceptual overlap [42]. Items and subscales meeting the criterion were subsequently selected as candidate independent variables for inclusion in the predictor set of the regression models.
Modeling approaches
Table 2 shows four predictor sets based on SF-12v2 dimensions and items in both direct and indirect mapping approaches. The explanatory variables were initially screened based on Spearman’s correlation between the SF-12v2 items/subscale scores and utility scores. The explanatory variables with p < 0.25 in the forward stepwise regression were included in the predictor sets. Furthermore, age was included as a covariate in all models due to its significant association with EQ-5D-5L utility scores and to reduce potential confounding effects. Accordingly, four predictor sets comprising subscale scores, polynomial transformation of subscale scores, and selected SF-12v2 items modeled as either continuous or categorical variables were evaluated using both mapping approaches.
A direct mapping approach was applied to predict EQ-5D-5L utility scores from SF-12v2 items and subscale-level scores. Multiple regression models were estimated, including OLS, Tobit, CLAD, GLM, TPM, ALDVMM, and Betamix. These models were selected based on their capacity to accommodate key features of EQ-5D-5L utility data, including boundedness, skewness, multimodality, and the discontinuity between the second-best and full-health states inherent in the Thai-specific value set.
In all direct mapping models, the EQ-5D-5L utility score served as the dependent variable, while SF-12v2 item/subscale scores and participant age were included as independent variables. For GLM and TPM specifications, disutility scores (1 – utility) were modeled to satisfy the non-negativity assumption of these approaches [43,44].
The modified Park test indicated that both Poisson and Gamma variance functions were plausible for the disutility outcome, as the estimated Gamma coefficient lay between 1 and 2 [45]. Based on Box-Cox regression results, a logarithmic link function was identified as optimal for both GLM and TPM models across all predictor sets [46].
Tobit and CLAD models were estimated to account for right-censoring, with utility scores censored at the upper bound of 1.0000 [31]. In contrast, ALDVMM and beta mixture models explicitly accommodated the multimodal distribution of utility scores while constraining predictions to the feasible utility range [37]. Accordingly, these models imposed lower and intermediate constraints at −0.4212 (worst health state) and 0.9436 (second-best health state), with truncation at 1.0000 for full health. Both mixture models were restricted to single-component specification, as models with additional components failed to converge across most predictor sets.
For the indirect mapping approach, each EQ-5D-5L dimension was modeled separately as a dependent variable, with SF-12v2 item/subscale scores serving as predictors. To estimate response-level probabilities, generalized ordered logistic (GOLOGIT), ordinal logistic (OLOGIT), or multinomial logistic (MLOGIT) models were evaluated.
GOLOGIT models failed to converge, and OLOGIT models were deemed inappropriate due to violations of the proportional odds assumption, indicating that predictor effects varied across response levels. Consequently, MLOGIT was selected as the final modeling strategy.
Using the most-likely-probability method, MLOGIT generated the most probable response level for each EQ-5D-5L dimension, producing a five-digit health-state profile that was subsequently converted into a utility score using the Thai-specific value set.
Notably, predicted utility scores were truncated at the theoretical boundaries of the Thai EQ-5D-5L value set when estimated values exceeded the feasible utility range. Specifically, values greater than 1 were truncated to the upper bound of 1.0, while values less than –0.4212 were truncated to the lower bound of –0.4212.
Model performance
Model performance was evaluated using ten-fold cross-validation (CV), whereby the dataset was randomly partitioned into 10 approximately equal-sized subsets (n ≈ 200 per fold). In each iteration, nine subsets were used for model estimation, and the remaining subset served as the validation sample. This procedure was repeated 10 times so that each subset functioned as the validation set exactly once.
Predictive accuracy was assessed using the mean absolute error (MAE) and root mean square error (RMSE) across all predictor sets. The optimal model was identified based on the lowest values of these metrics, where discrepancies arose, RMSE was prioritized due to its greater sensitivity to large prediction errors [44].
Model agreement was further evaluated using the intraclass correlation coefficient (ICC) derived from a two-way mixed-effects model with absolute agreement and single-measurement specifications. ICC values were interpreted as poor (ICC < 0.50), moderate (0.50 ≤ ICC <0.75), good (0.75 ≤ ICC <0.90), or excellent (ICC ≥ 0.90) agreement [47].
In accordance with NICE guidelines [48], mean utility score plots were constructed across the full range of WHOQOL-BREF total scores to identify systematic prediction bias. Additionally, cumulative distribution plots were used to examine discrepancies between predicted and observed utility values. The optimal model was expected to demonstrate the highest agreement and minimal divergence across these visual diagnostics. Notably, this study was conducted and reported in accordance with the Mapping onto Preference-Based Measures Reporting Standards checklist [49] and the reporting standards guidance outlined in the 2017 by International Society for Pharmacoeconomics and Outcomes Research (ISPOR) Task Force Report [50], as throuroghly described in Tables in S1-S2 Table.
All statistical analyses were conducted using Stata version 17 (StataCorp LLC, College Station, TX, USA) and Microsoft Excel. A two-sided p-value <0.05 was considered statistically significant.
Results
Participant characteristics
Among the 2,000 participants drawn from the general Thai population, mean SF-12v2 subscale scores ranged from 64.07 ± 24.94 for GH to 86.83 ± 17.89 for RE. The mean PCS and MCS scores were 50.28 ± 8.44 and 53.61 ± 7.11, respectively. Most SF-12v2 subscales exhibited both ceiling effects (scores of 100) and floor effects (scores of 0). An exception was observed for the MH subscale, which demonstrated only a ceiling effect, affecting 14.50% of respondents. Detailed descriptive statistics for participant characteristics and EQ-5D-5L utility scores are reported elsewhere [38].
Conceptual overlap between EQ-5D-5L and SF-12v2
Table 3 presents the absolute Spearman’s rank correlation coefficients between SF-12v2 items and subscale-level scores and EQ-5D-5L utility scores. All pairwise correlations were statistically significant (p < 0.001), except for the correlation between MH1 and SC, MCS, and MO, and MCS and SC.
At the subscale level, PF, RP, BP, and GH demonstrated the strongest correlations with EQ-5D-5L utility scores (|r| = 0.5904–0.6109) and showed moderate to strong correlations with at least one EQ-5D-5L dimension. These subscales were therefore retained for inclusion in the predictor sets used in the mapping models. In contrast, VT exhibited the weakest correlation with EQ-5D-5L utility scores (|r| = 0.2879) and was excluded from further analyses.
At the item level, GH1, PF1, PF2, RP1, RP2, RE1, RE2, BP1, and MH2 showed at least moderate correlations with EQ-5D-5L utility scores (|r| > 0.4). However, the inclusion of RP1, RP2, RE1, and RE2 concurrently raised concerns regarding multicollinearity, given their conceptual and statistical overlap within subscales. Consequently, RP2 and RE1 were excluded from the predictor sets, as their correlation coefficients were lower than those of RP1 and RE2, respectively.
Model selection
Table 4 summarizes predictive performance metrics across all regression models for both direct and indirect mapping approaches. For direct mapping, predictor set 4 demonstrated superior performance relative to the other sets, yielding the lowest average MAE and RMSE.
Although both Tobit and ALDVMM-1 component models achieved the lowest post-cross-validation (CV) MAE (0.0402 and 0.0410, respectively), the ALDVMM-1 component model produced the lowest post-CV RMSE (0.0625). In contrast, the Tobit model ranked 16th in terms of RMSE among all evaluated models.
Based on overall predicted accuracy, the ALDVMM-1 component model was selected as the best-performing direct mapping model. As illustrated in Fig 1, this model exhibited less variability in predicted utility scores across the full distribution compared with the Tobit model. Although cumulative distribution plots indicated that the Tobit model more closely regressed observed values at the upper boundary (1.0000), this was primarily attributable to truncation at the maximum utility value. In contrast, the ALDVMM-1 component model generated predictions without truncation at either end of the scale. Similar to the distribution plots, ALDVMM-1 component model outperformed the other models because it could predict the utility scores more closely aligned with the observed utility scores, particularly at the maximum utility score of 1.00 (50%) and for values below zero as shown in Figure in S1 Fig.
ALDVMM-1 component adjusted limited dependent variable mixture model with 1 component CLAD censored least absolute deviation GLM generalized linear model MLOGIT multinomial logistic regression OLS ordinary least squares Trunc truncated predicted utility values for the OLS, Tobit, and CLAD models. Their plots were generated with predicted values truncated at theoretical boundaries of the Thai EQ-5D-5L value set: an upper bound of 1.0 for values exceeding 1 and a lower bound of –0.4212 for values less than –0.4212. Truncation did not affect the comparative assessment of model performance. Non-truncated predictions are shown for GLM-poisson, ALDVMM-1 component, and MLOGIT, which generate predicted utility values within the Thai-specific value set by model structures.
Consistent with these findings, the ALDVMM-1 component model demonstrated the highest agreement between predicted and observed utility scores, with an ICC of 0.8192, exceeding that of all alternative regression models.
For indirect mapping, all five predictor sets were initially evaluated. However, the GOLOGIT model failed to converge, and the OLOGIT model violated the proportional odds assumption across all predictor sets. Consequently, MLOGIT was implemented.
Model convergence was achieved only for predictor set 1, yielding post-CV MAE and RMSE values of 0.0401 and 0.0703, respectively. Agreement between predicted and observed utility scores was good, with an ICC of 0.7810.
Fig 2 presents plots comparing predicted and observed mean utility scores stratified by WHOQOL-BREF total scores across all regression models. All models exhibited a consistent upward trend in predicted utility scores with increasing WHOQOL-BREF scores. For WHOQOL-BREF scores below 80, the GLM-poisson produced the most accurate predictions, whereas other models tended to overestimate utility.
ALDVMM-1 component adjusted limited dependent variable mixture model with 1 component CLAD censored least absolute deviation GLM generalized linear model MLOGIT multinomial logistic regression OLS ordinary least squares Trunc truncated predicted utility values for the OLS, Tobit, and CLAD models. Their plots were generated with predicted values truncated at theoretical boundaries of the Thai EQ-5D-5L value set: an upper bound of 1.0 for values exceeding 1 and a lower bound of –0.4212 for values less than –0.4212. Truncation did not affect the comparative assessment of model performance. Non-truncated predictions are shown for GLM-poisson, ALDVMM-1 component, and MLOGIT, which generate predicted utility values within the Thai-specific value set by model structures.
For WHOQOL-BREF scores between 80 and 105, the ALDVMM-1 component model demonstrated superior predictive performance compared to other regression models. In contrast, for scores exceeding 105, where data density was lower, the OLS model yielded more accurate predictions than other approaches. Notably, the MLOGIT model systematically overestimated utility scores across the entire WHOQOL-BREF score range.
Algorithm for direct and indirect mapping
Based on the above evaluations, the ALDVMM-1 component model with predictor set 4 and the MLOGIT model with predictor set 1 were identified as the best-predicting models for direct and indirect mapping, respectively.
The mathematical specification for predicting EQ-5D-5L utility scores using the selected direct mapping model is shown below:
where:
( denotes response to SF-12v2 items GH1, PF1, PF2, RP1, RE2, BP1, and SF1 (response levels for PF1 and PF2 ranged from 1–3; all others ranged from 1–5);
represents the estimated coefficient corresponding to the item
Age refers to the respondent’s age in years.
Estimated utility scores were subsequently used to calculate predicted utility values using algorithms derived from the StataCorp LLC implementation [51]. The predicted utility values were subsequently calculated using the following formula.
Where:
: The predicted value of dependent variable
(utility scores) for each individual
: The summation operator over C possible categories; however, it is one for this study because the best-predicting model is ALDVMM with one component.
: Cumulative distribution function of the standard normal distribution for upper limit
: Cumulative distribution function of the standard normal distribution for lower limit
: Upper bound is 0.9436 (the second highest values for the Thai-specific value set)
: Lower bound is -0.4212 (the lowest values for the Thai-specific value set)
: The standard deviation of the error term for category c where it is 0.0795847 for
this study
: Estimated utility value from the exploratory variable and regression coefficients
of the STATA output
: Probability density function of the standard normal distribution for upper limit
: Probability density function of the standard normal distribution for lower limit
Detailed coefficients and standard errors are provided in Table in S3 Table, with a worked example illustrated in Text in S1 Text.
Table in S4 Table reports the coefficients for the selected SF-12v2 items used in indirect mapping (predictor set 1), while Text in S2 Text provides a step-by-step example of utility scores calculated via the indirect mapping approach.
The predicted utility scores may be applied in future economic evaluations. The variance-covariance matrices are available in Microsoft Excel format to facilitate probabilistic sensitivity analysis in S2 File.
Discussion
This study represents the first Thai mapping investigation to convert SF-12v2 responses into EQ-5D-5L utility scores using a Thai-specific value in a general population sample, employing both direct and indirect mapping approaches. In doing so, it addresses several methodological limitations of earlier mapping studies, many of which relied on EQ-5D-3L utilities [25–31] and predominantly applied OLS regression [26,27,29]. OLS is widely recognized as suboptimal for modeling health utility data because it fails to adequately accommodate ceiling effects and distributional irregularities [34].
Correlation analyses demonstrated substantial conceptual overlap between SF-12v2 items or subscale scores and EQ-5D-5L utility scores and dimensions. Most SF-12v2 subscales exhibited at least moderate correlations with EQ-5D-5L outcomes, except for the VT subscale and item. On this basis, the majority of SF-12v2 subscales and items were retained in the predictor sets. These findings are consistent with previous mapping studies conducted in both general and patient populations internationally, including studies from Thailand [14,52–54].
Although the MH subscale showed moderate correlations with AD and EQ-5D-5L utility scores, the MH1 item exhibited only weak correlations with EQ-5D-5L dimensions and utility. Moreover, inclusion of the MH subscale did not improve overall model fit. Consequently, only the MH2 item was retained in the final predictor set. PCS and MCS were excluded from all predictor sets because their inclusions introduced multicollinearity with SF-12v2 subscales (VIF > 0.40) and did not improve predictive performance.
For direct mapping, predictor set 4, which comprised categorical SF-12v2 items (GH1, PF1, PF2, RP1, RE2, BP1, and SF1), demonstrated the best predictive performance across models. Based on post-CV MAE and RMSE, the ALDVMM-1 component, OLS, and Tobit models emerged as the strongest candidates. However, OLS was excluded because it did not adequately account for the bounded, skewed, and heteroscedastic nature of utility data [55,56]. Although Tobit regression is theoretically suitable for censored outcomes, it did not yield the lowest model prediction errors, violated assumptions of normality and homoscedasticity [57], and produced predicted utility values exceeding 1.000, requiring artificial truncation. Consequently, the ALDVMM-1 component model was selected as the best direct mapping approach.
The ALDVMM-1 component model consistently demonstrated superior predictive performance. Its advantage lies in its ability to address key distributional challenges inherent in health utility data, including pronounced ceiling effects, multimodality, and the discontinuity between perfect health and the next-best health state [58]. These findings are consistent with previous mapping studies [25,37,58,59], indicating that mixture models outperform conventional linear models in both general and patient populations. Notably, the present ALDVMM-1 component outperformed earlier SF-12v2 mapping studies that relied on CLAD and simple OLS models based on PCS and MCS scores alone.
Although ALDVMM and beta mixture models are designed to accommodate the multimodality of utility distributions, the ALDVMM-2 component and beta mixture models failed to fully converge for the predictor set 4, which consisted of several SF-12v2 items coded as categorical variables in this study. This may be attributable to the substantial increase in the number of estimated parameters in these mixture models, resulting in instability and convergence difficulties. Furthermore, this mapping algorithm was developed from the predominantly healthy samples leading to sparse responses in certain categorical levels of the SF-12v2 items, particularly for poorer health status. Therefore, utility distribution might not have sufficient multimodality for complex mixture model specifications. ALDVMM-1 component can provide a more stable utility estimation while maintaining good predictive performance.
Whereas previous studies reported a mean prediction error of approximately 0.0744 and an MAE value around 0.14 in general [31] and socioeconomically disadvantaged US populations [27], the current study achieved markedly improved accuracy, with a mean error of 0.0008 and an MAE of 0.0410. These results support the inclusion of SF-12v2 items as exploratory variables rather than relying exclusively on summary component scores. Furthermore, the ALDVMM-1 component outperformed mixture models reported in earlier SF-12v2 mapping studies, despite those models being developed to predict EQ-5D-3L utility in US national samples. Specifically, the present model yielded a substantially lower post-CV RMSE (0.0625) compared with previously reported values ranging from 0.146 to 0.166 [25].
Because SF-12v2 items can be reclassified to derive the SF-6D instrument, the UK valuation algorithm is commonly applied in Thailand in the absence of a Thai-specific value set. Previous Thai studies have supported the use of SF-6D utilities for HRQoL measurement and economic evaluations in both the general [14] and chronic disease populations [52]. However, the present findings did not demonstrate conceptual overlap between the VT1 item and EQ-5D-5L utilities or dimensions. Moreover, agreement between SF-6D–derived and observed EQ-5D-5L utilities was lower (ICC = 0.7300) than that observed between predicted and observed EQ-5D-5L utilities using the proposed algorithm (ICC = 0.8192) with the general Thai samples in this present dataset.
Importantly, the proposed mapping algorithm showed greater sensitivity in poorer health states, predicting utility values as low as −0.2214. This value is substantially lower than the minimum utility score of 0.29 obtainable using the UK SF-6D algorithm [40]. In the absence of a Thai-specific SF-6D value set, the ALDVMM-based mapping equation is therefore recommended for estimating EQ-5D-5L utility scores from SF-12v2 data in general Thai populations.
For indirect mapping, the MLOGIT model with predictor set 1 is recommended. However, indirect mapping did not outperform direct mapping, as reflected by a higher post-CV RMSE (0.0703 vs 0.0625). Distributional and cumulative plots further indicated systematic overprediction of utility values above 0.6. Although OLOGIT is theoretically appropriate for ordinal outcomes [60], it was unsuitable in this study due to violations of the proportional odds assumption and convergence issues. Consequently, MLOGIT was employed to estimate EQ-5D-5L dimensions response probabilities.
The relatively weaker performance of indirect mapping may be attributable to sparse extreme responses within the EQ-5D-5L dimension in this generally healthy sample, leading to inflated standard errors. Similar challenges have been documented in previous mapping studies [30,44]. Despite these limitations, the MLOGIT model demonstrated superior performance compared with earlier US-based studies, achieving a lower MAE (0.0401 vs 0.0480) [30]. This improvement may reflect the enhanced discriminative capacity and reduced ceiling effects of the previous mapping study estimating the EQ-5D-3L utility scores, which are noted to have higher ceiling effects of EQ-5D-5L relative to EQ-5D-3L [30].
Although there are several mapping algorithm generated from general representative samples from other countries, they might not be applicable to Thai population because mapping algorithms are largely population-dependent because both predictor distributions and health-state valuations may vary across countries and cultural contexts. Differences in demographic composition and baseline health status between the present Thai sample and the other populations used to develop their country -based mapping algorithms may therefore influence model estimation and predictive accuracy. In the previous US-based mapping algorithm from the US representative samples [31], two key differences may contribute to the observed performance gap between the two algorithms. First, the US-derived algorithms were developed using a sample in which the majority of respondents presenting with more chronic conditions (average number of chronic conditions = 1.90), compared with approximately 0.37 in the Thai derivation sample. Therefore, the US-derived algorithms may predict utility scores more accurately among unhealthy respondents or those at the lower end of the health spectrum than the Thai mapping algorithm. This explanation is supported by the lower mean utility score estimated using the US-based algorithm (0.8912), compared with that estimated using the Thai mapping algorithm (0.9222). Second, cultural differences between these two countries can influence how respondents interpret and report HRQoL items because health perception is inherently shaped by sociocultural context. Previous study has demonstrated that certain culture-specific health dimensions perceived by the Thai population differ from those reported in Western populations [61]. As a result, these factors may help explain why a Thai-specific mapping algorithm is required for estimating utility scores in the Thai population. Nevertheless, the performance of the Thai mapping algorithm should be further validated using an independent dataset from the general Thai population, particularly among unhealthy individuals or clinical population.
Several limitations should be acknowledged. First, this study sample comprised predominantly healthy individuals from the general population, resulting in a sparse representation of severe health states across EQ-5D-5L dimensions. Consequently, the mapping algorithm may be less accurate for populations with poorer health. Second, although 10-fold CV was conducted to evaluate the model performance and reduce the overfitting of the mapping algorithm, the external validation using an independent dataset was not performed. Future studies should validate and refine this proposed mapping algorithm in less healthy samples or clinical population before broader application in economic analyses.
Conclusions
This study developed a novel mapping algorithm to convert SF-12v2 responses into EQ-5D-5L utility scores for the general Thai population. While earlier Thai studies have supported the application of the UK valuation algorithm to derive SF-6D utility score, the present findings challenge this practice. Specifically, correlation analyses did not demonstrate sufficient conceptual overlap between the vitality subscale or item and EQ-5D-5L dimension or scores. In addition, agreement between SF-6D–derived and observed EQ-5D-5L utility scores was lower than that observed between the mapped and observed EQ-5D-5L utility generated by the proposed algorithm.
Despite these strengths, an important limitation should be acknowledged. The proposed mapping algorithm may have reduced accuracy in predicting utility scores among individuals with poor health status, reflecting the limited representation of severe health states in the study sample. Accordingly, caution is warranted when applying this algorithm to populations with substantial morbidity, and further validation in less healthy or clinical samples is recommended.
Supporting information
S1 Table. Guidelines and checklist for mapping onto Preference-Based Measures Standard (MAPS) checklist.
https://doi.org/10.1371/journal.pone.0351064.s001
(DOCX)
S2 Table. Checklist for 2017 ISPOR Good Practices Report: Mapping to Estimate Health-State Utility from Non-Preference-Based Outcome Measures.
Summary of reporting of mapping studies recommendations.
https://doi.org/10.1371/journal.pone.0351064.s002
(DOCX)
S3 Table. Coefficients and standard errors of the final model used for direct mapping based on the Thai value set.
https://doi.org/10.1371/journal.pone.0351064.s003
(DOCX)
S4 Table. Coefficients and standard errors of the final model used for indirect mapping.
https://doi.org/10.1371/journal.pone.0351064.s004
(DOCX)
S1 Text. Instructions for predicting utility scores using the direct mapping algorithm.
https://doi.org/10.1371/journal.pone.0351064.s005
(DOCX)
S2 Text. Instructions for predicting utility scores using the indirect mapping algorithm.
https://doi.org/10.1371/journal.pone.0351064.s006
(DOCX)
S1 Fig. Comparison of the distributions of observed and predicted utility scores across all regression models for direct and indirect mapping.
https://doi.org/10.1371/journal.pone.0351064.s007
(DOCX)
S2 File. Variance-Covariance matrix for both direct and indirect mapping algorithms.
https://doi.org/10.1371/journal.pone.0351064.s009
(XLSX)
References
- 1. Chen Y. Health technology assessment and economic evaluation: Is it applicable for the traditional medicine?. Integr Med Res. 2022;11(1):100756. pmid:34401322
- 2. J Asunmonu O. Strategic Resource Allocation in Modern Health Systems: A Systematic Review of the Role and Impact of Health Technology Assessments (HTAs) on Healthcare Financing and Policy Decisions. IJMCR. 2025;4(3):51–5.
- 3. Tanvejsilp P, Taychakhoonavudh S, Chaikledkaew U, Chaiyakunapruk N, Ngorsuraches S. Revisiting Roles of Health Technology Assessment on Drug Policy in Universal Health Coverage in Thailand: Where Are We? And What Is Next? Value Health Reg Issues. 2019;18:78–82. https://doi.org/10.1016/j.vhri.2018.11.004
- 4. Wani S, Alsabti H, Allamki S, Almandhari A, Alhajji S, Alrashdi I, et al. Methodological guidelines for Health Technology Assessment in Oman. J Pharm Policy Pract. 2025;18(1):2596523. pmid:41384030
- 5. Rowen D, Azzabi Zouraq I, Chevrou-Severac H, van Hout B. International regulations and recommendations for utility data for health technology assessment. Pharmacoeconomics. 2017;35(Suppl 1):11–9. https://doi.org/10.1007/s40273-017-0544-y
- 6. Botwright S, Sittimart M, Chavarina KK, Bayani DB, Merlin T, Surgey G, et al. Good Practices for Health Technology Assessment Guideline Development: A Report of the Health Technology Assessment International, HTAsiaLink, and ISPOR Special Task Force. Value in Health. 2025;28(1):1–15. https://doi.org/10.1016/j.jval.2024.09.001
- 7.
Rai M, Goyal R. Pharmacoeconomics in Healthcare. Pharmaceutical Medicine and Translational Clinical Research. Boston: Academic Press. 2018. 465–72.
- 8. Finch AP, Brazier JE, Mukuria C. What is the evidence for the performance of generic preference-based measures? A systematic overview of reviews. Eur J Health Econ. 2018;19(4):557–70. pmid:28560520
- 9. Wolowacz SE, Briggs A, Belozeroff V, Clarke P, Doward L, Goeree R, et al. Estimating Health-State Utility for Economic Models in Clinical Studies: An ISPOR Good Research Practices Task Force Report. Value Health. 2016;19(6):704–19. pmid:27712695
- 10. Meregaglia M, Nicod E, Drummond M. The estimation of health state utility values in rare diseases: do the approaches in submissions for NICE technology appraisals reflect the existing literature? A scoping review. Eur J Health Econ. 2023;24(7):1151–216. pmid:36335234
- 11. Pattanaphesaj J, Thavorncharoensap M, Ramos-Goni JM, Tongsiri S, Ingsrisawang L, Teerawattananon Y. The EQ-5D-5L Valuation study in Thailand. Expert Rev Pharmacoecon Outcomes Res. 2018;18(5):551–8. https://doi.org/10.1080/14737167.2018.1494574
- 12. Németh G. Health related quality of life outcome instruments. Eur Spine J. 2006;15 Suppl 1(Suppl 1):S44-51. pmid:16320032
- 13. Kangwanrattanakul K, Phimarn W. A systematic review of the development and testing of additional dimensions for the EQ-5D descriptive system. Expert Rev Pharmacoecon Outcomes Res. 2019;19(4):431–43. pmid:31244348
- 14. Kangwanrattanakul K. A comparison of measurement properties between UK SF-6D and English EQ-5D-5L and Thai EQ-5D-5L value sets in general Thai population. Expert Rev Pharmacoecon Outcomes Res. 2021;21(4):765–74. pmid:32981380
- 15. Sakthong P, Sonsa-Ardjit N, Sukarnjanaset P, Munpan W. Psychometric properties of the EQ-5D-5L in Thai patients with chronic diseases. Qual Life Res. 2015;24(12):3015–22. pmid:26048348
- 16. Turner N, Campbell J, Peters TJ, Wiles N, Hollinghurst S. A comparison of four different approaches to measuring health utility in depressed patients. Health Qual Life Outcomes. 2013;11:81. pmid:23659557
- 17. Haywood KL, Garratt AM, Dziedzic K, Dawes PT. Generic measures of health-related quality of life in ankylosing spondylitis: reliability, validity and responsiveness. Rheumatology (Oxford). 2002;41(12):1380–7. pmid:12468817
- 18. Kangwanrattanakul K. Mapping of the World Health Organization Quality of Life Brief (WHOQOL-BREF) to the EQ-5D-5L in the General Thai Population. Pharmacoecon Open. 2023;7(1):139–48. https://doi.org/10.1007/s41669-022-00380-0
- 19. Sakthong P. Mapping World Health Organization Quality of Life-BREF Onto 5-Level EQ-5D in Thai Patients With Chronic Diseases. Value Health. 2021;24(8):1089–94. pmid:34372973
- 20. Młyńczak K, Golicki D. Psychometric properties of the Polish version of SF-12v2 in the general population survey. Expert Rev Pharmacoecon Outcomes Res. 2022;22(3):465–72. pmid:33941017
- 21. Younsi M. Health-Related Quality-of-Life Measures: Evidence from Tunisian Population Using the SF-12 Health Survey. Value Health Reg Issues. 2015;7:54–66. pmid:29698153
- 22. Kangwanrattanakul K. Factor structure and trends in SF-12v2 health-related quality of life scores among pre-and post-pandemic samples in Thailand: confirmatory factor analysis and Rasch analysis. Health Qual Life Outcomes. 2025;23(1):74. pmid:40691581
- 23. Brazier JE, Kolotkin RL, Crosby RD, Williams GR. Estimating a preference-based single index for the Impact of Weight on Quality of Life-Lite (IWQOL-Lite) instrument from the SF-6D. Value Health. 2004;7(4):490–8. pmid:15449641
- 24. Oliveira Gonçalves AS, Werdin S, Kurth T, Panteli D. Mapping Studies to Estimate Health-State Utilities From Nonpreference-Based Outcome Measures: A Systematic Review on How Repeated Measurements are Taken Into Account. Value Health. 2023;26(4):589–97. pmid:36371289
- 25. Coca Perraillon M, Shih Y-CT, Thisted RA. Predicting the EQ-5D-3L Preference Index from the SF-12 Health Survey in a National US Sample: A Finite Mixture Approach. Med Decis Making. 2015;35(7):888–901. pmid:25840902
- 26. Franks P, Lubetkin EI, Gold MR, Tancredi DJ, Jia H. Mapping the SF-12 to the EuroQol EQ-5D Index in a national US sample. Med Decis Making. 2004;24(3):247–54. pmid:15185716
- 27. Franks P, Lubetkin EI, Gold MR, Tancredi DJ. Mapping the SF-12 to preference-based instruments: convergent validity in a low-income, minority population. Med Care. 2003;41(11):1277–83. pmid:14583690
- 28. Gray AM, Rivero-Arias O, Clarke PM. Estimating the association between SF-12 responses and EQ-5D utility values by response mapping. Med Decis Making. 2006;26(1):18–29. pmid:16495197
- 29. Lawrence WF, Fleishman JA. Predicting EuroQoL EQ-5D preference scores from the SF-12 Health Survey in a nationally representative sample. Med Decis Making. 2004;24(2):160–9. pmid:15090102
- 30. Le QA. Probabilistic mapping of the health status measure SF-12 onto the health utility measure EQ-5D using the US-population-based scoring models. Qual Life Res. 2014;23(2):459–66. pmid:24026631
- 31. Sullivan PW, Ghushchyan V. Mapping the EQ-5D index from the SF-12: US general population preferences in a nationally representative sample. Med Decis Making. 2006;26(4):401–9. pmid:16855128
- 32. Buchholz I, Janssen MF, Kohlmann T, Feng Y-S. A Systematic Review of Studies Comparing the Measurement Properties of the Three-Level and Five-Level Versions of the EQ-5D. Pharmacoeconomics. 2018;36(6):645–61. pmid:29572719
- 33. Yu L, Yang H, Lu L, Fang Y, Zhang X, Li S, et al. Developing mapping algorithms to predict EQ-5D health utility values from Bath Ankylosing Spondylitis Disease Activity Index and Bath Ankylosing Spondylitis Functional Index among patients with Ankylosing Spondylitis. Health Qual Life Outcomes. 2024;22(1):61. pmid:39113080
- 34. Huang I-C, Frangakis C, Atkinson MJ, Willke RJ, Leite WL, Vogel WB, et al. Addressing ceiling effects in health status measures: a comparison of techniques applied to measures for people with HIV disease. Health Serv Res. 2008;43(1 Pt 1):327–39. pmid:18211533
- 35. Fang H, Hong T, Liu X, Luo C, Hou Y, Xie S. Mapping the ADDQoL to the EQ-5D-5L and SF-6Dv2 among Chinese patients with type 2 diabetes mellitus. Health Qual Life Outcomes. 2025;23(1):46. pmid:40307906
- 36. Chen Z, Yang L, Zhang Y. Mapping KDQOL-36 Onto EQ-5D-5L and SF-6Dv2 in Patients Undergoing Dialysis in China. Value Health Reg Issues. 2025;48:101103. https://doi.org/10.1016/j.vhri.2025.101103
- 37. Aghdaee M, Gu Y, Sinha K, Parkinson B, Sharma R, Cutler H. Mapping the Patient-Reported Outcomes Measurement Information System (PROMIS-29) to EQ-5D-5L. Pharmacoeconomics. 2023;41(2):187–98. pmid:36336773
- 38. Kangwanrattanakul K, Krägeloh CU. EQ-5D-3L and EQ-5D-5L population norms for Thailand. BMC Public Health. 2024;24(1):1108. https://doi.org/10.1186/s12889-024-18391-3
- 39.
Ware JE Jr, Kosinski M, Keller SD. SF-12: how to score the SF-12 physical and mental health summary scales Second ed. Boston, MA: The Health Institute, New England Medical Center; 1995.
- 40. Brazier JE, Roberts J, Deverill M. The estimation of a preference-based measure of health from the SF-36. J Health Econ. 2002;2(21):271–92. pmid:11939242
- 41. Kangwanrattanakul K, Krägeloh CU. Psychometric evaluation of the WHOQOL-BREF and its shorter versions for general Thai population: confirmatory factor analysis and Rasch analysis. Qual Life Res. 2024;33(2):335–48. pmid:37906345
- 42. Akoglu H. User’s guide to correlation coefficients. Turk J Emerg Med. 2018;18(3):91–3. pmid:30191186
- 43. Nguyen LH, Tran BX, Hoang Le QN, Tran TT, Latkin CA. Quality of life profile of general Vietnamese population using EQ-5D-5L. Health Qual Life Outcomes. 2017;15(1):199. pmid:29020996
- 44. Tan YJ, Ong SC. Direct and indirect mapping of assessment of quality of life - 6 dimensions (AQoL-6D) onto EQ-5D-5L utilities using data from a multicenter, cross-sectional study of Malaysians with chronic heart failure. Value in Health. 2024;27(12):1762–70. https://doi.org/10.1016/j.jval.2024.07.016
- 45. Zhou J, Williams C, Keng MJ, Wu R, Mihaylova B. Estimating costs associated with disease model states using generalized linear models: a tutorial. Pharmacoeconomics. 2024;42(3):261–73. https://doi.org/10.1007/s40273-023-01319-x
- 46. Riani M, Atkinson AC, Corbellini A. Automatic robust Box–Cox and extended Yeo–Johnson transformations in regression. Stat Methods Appl. 2022;32(1):75–102.
- 47. Koo TK, Li MY. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. J Chiropr Med. 2016;15(2):155–63. pmid:27330520
- 48.
Allan WMHA, Pudney S. Mapping to Estimate Health State Utilities 2023 https://www.sheffield.ac.uk/media/42422/download?attachment
- 49. Petrou S, Rivero-Arias O, Dakin H, Longworth L, Oppe M, Froud R, et al. Preferred reporting items for studies mapping onto preference-based outcome measures: The MAPS statement. Health Qual Life Outcomes. 2015;13:106. pmid:26232268
- 50. Wailoo AJ, Hernandez-Alava M, Manca A, Mejia A, Ray J, Crawford B, et al. Mapping to Estimate Health-State Utility from Non-Preference-Based Outcome Measures: An ISPOR Good Practices for Outcomes Research Task Force Report. Value Health. 2017;20(1):18–27. pmid:28212961
- 51. Alava MH, Wailoo A. Fitting Adjusted Limited Dependent Variable Mixture Models to EQ-5D. The Stata Journal: Promoting communications on statistics and Stata. 2015;15(3):737–50.
- 52. Sakthong P, Munpan W. A Head-to-Head Comparison of UK SF-6D and Thai and UK EQ-5D-5L Value Sets in Thai Patients with Chronic Diseases. Appl Health Econ Health Policy. 2017;15(5):669–79. pmid:28290106
- 53. De Smedt D, Clays E, Annemans L, De Bacquer D. EQ-5D versus SF-12 in coronary patients: are they interchangeable?. Value Health. 2014;17(1):84–9. 10.1016/j.jval.2013.10.010
- 54. Xie S, Wu J, Chen G. Comparative performance and mapping algorithms between EQ-5D-5L and SF-6Dv2 among the Chinese general population. Eur J Health Econ. 2024;25(1):7–19. pmid:36709458
- 55. Longworth L, Rowen D. Mapping to obtain EQ-5D utility values for use in NICE health technology assessments. Value Health. 2013;16(1):202–10. pmid:23337232
- 56. Lamu AN, Olsen JA. Testing alternative regression models to predict utilities: mapping the QLQ-C30 onto the EQ-5D-5L and the SF-6D. Qual Life Res. 2018;27(11):2823–39. pmid:30173314
- 57.
Cunillera O. Tobit models. In: Maggino F. Encyclopedia of Quality of Life and Well-Being Research. Cham: Springer International Publishing. 2023. 7237–42.
- 58. Worboys HM, Gray L, Burton J, Alava MH, Greenwood S, Cooper N. Mapping the Kidney Disease Quality-of-Life Questionnaire Onto the EQ-5D-5L Utility Index in Patients Undergoing Hemodialysis. Value Health. 2025;28(7):1091–9. pmid:40246068
- 59. Gray LA, Hernandez Alava M, Wailoo AJ. Mapping the EORTC QLQ-C30 to EQ-5D-3L in patients with breast cancer. BMC Cancer. 2021;21(1):1237. pmid:34794404
- 60. Freitas Souza R de, Lima FG, Corrêa HL. Multilevel Ordinal Logit Models: A Proportional Odds Application Using Data from Brazilian Higher Education Institutions. Axioms. 2024;13(1):47.
- 61. Mao Z, Ahmed S, Graham C, Kind P, Sun Y-N, Yu C-H. Similarities and Differences in Health-Related Quality-of-Life Concepts Between the East and the West: A Qualitative Analysis of the Content of Health-Related Quality-of-Life Measures. Value Health Reg Issues. 2021;24:96–106. pmid:33524902