Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Direct and indirect mapping of the 12-item Short Form Survey version 2 (SF-12v2) onto the EQ-5D-5L utility scores in general Thai population

Abstract

The SF-12v2 is widely used to measure health-related quality of life (HRQoL) in the general Thai population because it is brief and places a low burden on respondents. However, it does not provide utility scores required for economic analyses. The mapping approach is a solution to estimate the utility scores from the SF-12v2 response using several regression models. Therefore, this study aimed to develop a mapping algorithm to estimate EQ-5D-5L utility scores from the SF-12v2 items using a 2022 national dataset of 2000 Thai respondents. Four predictor sets incorporating SF-12v2 items/subscales and age as a covariate were investigated using direct and indirect mapping approaches. Direct mapping approaches included ordinary least squares, Tobit, censored least absolute deviations, generalized linear model (GLM), two-part models, adjusted limited dependent variable mixture model (ALDVMM), and beta mixture model, while multinomial logistic regression (MLOGIT) was investigated for indirect mapping. Model performance was evaluated using ten-fold cross-validated (post-CV) mean absolute error (MAE) and root mean square error (RMSE). The ALDVMM-1 component with a predictor set comprising age and selected SF-12 items as categorical variables demonstrated the best predictive performance, yielding the lowest post-CV MAE and RMSE (0.0410 and 0.0625, respectively). However, MLOGIT showed poorer predictive performance. Therefore, the proposed ALDVMM-1 component-based mapping algorithm may facilitate the estimation of utility scores in studies where only the SF-12v2 was collected. Given that the mapping algorithm was developed from a predominantly healthy sample, it may be less reliable in populations with poorer health status, and further validation in populations with less healthy or clinical samples is warranted in future studies.

Introduction

Health technology assessment (HTA) is widely used to evaluate and identify the most cost-effective health interventions to support decision-making in healthcare [1]. It also optimizes the allocation of limited health resources [2,3]. Several HTA guidelines recommend cost-utility analysis to guide clinical decisions and policy implementation [46]. This approach compares incremental health improvement, expressed in quality-adjusted life years (QALYs), across or within health conditions [7]. QALYs are calculated using utility scores and life expectancy, making accurate measurement of utility scores essential [7,8].

Utility scores range from 0 (death or worst health state) to 1 (perfect health) [9]. Preference-based instruments, such as the EuroQol 5-dimension (EQ-5D), Health Utility Index (HUI), and Short Form 6 Dimension (SF-6D), measure these scores, which are then converted using a country-specific value set [10,11]. The EQ-5D-5L is the most commonly used instrument in Thailand. However, it may not capture some dimensions considered important, such as social health [12,13], and shows limited responsiveness to changes in both general [14] and patient populations [15].

Although preference-based instruments are standard for economic analyses, they do not provide dimension-specific (health profile) scores. Health profile instruments, such as the World Health Organization Quality of Life: Brief Version (WHOQOL-BREF) and the 12-item Short Form Survey version 2 (SF-12v2), offer more sensitivity to health changes [16,17]. However, they do not generate a utility score. Mapping between preference-based and profile instruments allows retention of profile information while producing utility scores for economic analyses.

Several Thai studies have mapped WHOQOL-BREF to EQ-5D-5L [18,19]. Although WHOQOL-BREF captures physical, psychological, social, and environmental health dimensions, its 26 items may still impose a respondent burden. The SF-12v2, derived from SF-36, is shorter, less burdensome, and widely used in large population surveys, including in Thailand [2022]. Its conceptual overlap with EQ-5D-5L enhances suitability for mapping algorithms [14].

Despite its widespread use, SF-12v2 lacks a Thai-specific value set for producing utility scores. Mapping approaches can predict EQ-5D-5L utility scores from SF-12v2 responses using regression models [23,24]. Previous studies mapped SF-12v2 to the earlier EQ-5D-3L [2531], which has a higher ceiling effect and lower discriminative power than EQ-5D-5L [32]. Ordinary least squares (OLS) models were sometimes used [26,27,29], but they are suboptimal for bounded utility scores with high ceiling effects and predict the utility scores outside the feasible range [33,34]. Alternative models, including Tobit, Beta mixture model (Betamix), Censored Least Absolute Deviation (CLAD), Generalized linear model (GLM), Adjusted Limited Dependent Variable Mixture Model (ALDVMM), and two-part models (TPM), better accommodate utility score characteristics [3437].

To date, no Thai studies have mapped SF-12v2 to EQ-5D-5L using these models, nor performed indirect (response-based) mappings. This study aims to develop a mapping algorithm between SF-12v2 and EQ-5D-5L, including indirect mapping based on SF-12v2 responses in the general Thai population.

Methods

Study design and samples

This study used data from the project “EQ-5D-3L and EQ-5D-5L population norms in Thailand [38].” A cross-sectional survey was conducted using face-to-face interviews with 2,000 Thai adults. To ensure national representation, a four-stage stratified random sampling method was applied to select provinces, districts, sub-districts, and villages. Data were collected from 12 provinces, representing all geographic regions of Thailand: Bangkok, Samut Prakan, Nonthaburi, Chonburi, Nakhon Pathom, Chiang Mai, Nakhon Sawan, Nakhon Ratchasima, Khon Kaen, Buriram, Nakhon Si Thammarat, and Phatthalung.

Participants were eligible if they (1) were aged 18 years or older, (2) were able to understand the interview process as assessed by the researcher or trained interviewers, and (3) were fluent in Thai. Individuals with acute or life-threatening illnesses or severe cognitive impairments were excluded.

Data collection procedures

Trained interviewers read all questions and response options verbatim, and did not provide explanations or interpretations to minimize interviewer bias. All participants completed both the SF-12v2 and EQ-5D-5L questionnaires, with no missing responses. These face-to-face interviews with eligible participants were conducted at their home residences between 1st May to 30th June 2023.

Before data collection, participants received an information sheet outlining study objectives, procedures, and their rights to withdraw without consequences. Written informed consent was obtained from all participants. Questionnaires were administered in the following order: (1) demographic information, (2) EQ-5D-5L, and (3) SF-12v2. Details of the demographic variables are reported elsewhere [38]. The protocol and data collection process were approved by the ethical committees of Burapha University Institutional Review Board (IRB1–031/2566) and adhered to the principles outlined in the Declaration of Helsinki.

Source and target measures

SF-12v2 (source measure).

The SF-12v2 consists of 12 items grouped into eight subscales: physical functioning (PF: 2 items), role limitations due to physical problems (RP, 2 items), bodily pain (BP, 1 item), general health (GH, 1 item), vitality (VT, 1 item), social functioning (SF, 1 item), role limitations due to emotional problems (RE, 2 items), and mental health (MH, 2 items) [39]. Participants rated their health over a four-week recall period. Most items use a five-point Likert scale, except PF items, which use a three-point scale. Four items (GH1, BP1, VT1, MH1) were reverse-coded to ensure scoring consistency. Table 1 displays descriptions of the SF-12v2 items.

Subscale scores were transformed to a 0–100 scale, with higher scores indicating better health [39]. The Physical Component Summary (PCS) and Mental Component Summary (MCS) scores were calculated using a proprietary scoring algorithm to generate norm-based scores. SF-12v2 responses were also reclassified to derive SF-6D health states, allowing utility score estimation using the UK algorithm, with scores ranging from 0.29 to 1.00 [40].

EQ-5D-5L (target measure).

The EQ-5D-5L comprises five dimensions: mobility, self-care, usual activities, pain/discomfort, and anxiety/depression. Each dimension is assessed across five ordered severity levels: no problems, slight problems, moderate problems, severe problems, extreme problems/unable to perform. Responses across the five dimensions generate a five-digit health-state profile that characterizes an individual’s health status. These profiles are subsequently converted into utility scores using country-specific valuation algorithms.

In this study, EQ-5D-5L utility scores were derived using the Thai-specific value set, with values ranging from −0.4212, corresponding to the worst health state (55555), to 1.0000 representing full health (11111). The second-best health state (11121) is valued at 0.9436 [11].

The WHOQOL-BREF instrument was additionally employed to compute overall quality-of-life scores by summing four domain scores: physical health (7 items), psychological health (6 items), social relationships (3 items), and environment (8 items), using the ordinal-to-interval conversion table for general Thai population [41]. WHOQOL-BREF scores range from 24 to 120, with higher scores indicating better perceived health-related quality of life. These total scores were used to stratify participants and to plot mean EQ-5D-5L utility scores, facilitating comparisons between predicted and observed utility values.

Conceptual overlap between EQ-5D-5L and SF-12v2

Conceptual overlap between the EQ-5D-5L and the SF-12v2 was examined using Spearman’s rank-order correlation analysis. Pairwise correlation coefficients (r) were calculated between each SF-12v2 item and subscale score and each EQ-5D-5L dimension, as well as the overall EQ-5D-5L utility score.

An SF-12v2 item or subscale demonstrating at least a moderate correlation (r 0.4) with any EQ-5D-5L dimension or the utility scores was deemed to exhibit conceptual overlap [42]. Items and subscales meeting the criterion were subsequently selected as candidate independent variables for inclusion in the predictor set of the regression models.

Modeling approaches

Table 2 shows four predictor sets based on SF-12v2 dimensions and items in both direct and indirect mapping approaches. The explanatory variables were initially screened based on Spearman’s correlation between the SF-12v2 items/subscale scores and utility scores. The explanatory variables with p < 0.25 in the forward stepwise regression were included in the predictor sets. Furthermore, age was included as a covariate in all models due to its significant association with EQ-5D-5L utility scores and to reduce potential confounding effects. Accordingly, four predictor sets comprising subscale scores, polynomial transformation of subscale scores, and selected SF-12v2 items modeled as either continuous or categorical variables were evaluated using both mapping approaches.

thumbnail
Table 2. Four predictors sets based on SF-12v2 dimensions and items for direct and indirect mapping.

https://doi.org/10.1371/journal.pone.0351064.t002

A direct mapping approach was applied to predict EQ-5D-5L utility scores from SF-12v2 items and subscale-level scores. Multiple regression models were estimated, including OLS, Tobit, CLAD, GLM, TPM, ALDVMM, and Betamix. These models were selected based on their capacity to accommodate key features of EQ-5D-5L utility data, including boundedness, skewness, multimodality, and the discontinuity between the second-best and full-health states inherent in the Thai-specific value set.

In all direct mapping models, the EQ-5D-5L utility score served as the dependent variable, while SF-12v2 item/subscale scores and participant age were included as independent variables. For GLM and TPM specifications, disutility scores (1 – utility) were modeled to satisfy the non-negativity assumption of these approaches [43,44].

The modified Park test indicated that both Poisson and Gamma variance functions were plausible for the disutility outcome, as the estimated Gamma coefficient lay between 1 and 2 [45]. Based on Box-Cox regression results, a logarithmic link function was identified as optimal for both GLM and TPM models across all predictor sets [46].

Tobit and CLAD models were estimated to account for right-censoring, with utility scores censored at the upper bound of 1.0000 [31]. In contrast, ALDVMM and beta mixture models explicitly accommodated the multimodal distribution of utility scores while constraining predictions to the feasible utility range [37]. Accordingly, these models imposed lower and intermediate constraints at −0.4212 (worst health state) and 0.9436 (second-best health state), with truncation at 1.0000 for full health. Both mixture models were restricted to single-component specification, as models with additional components failed to converge across most predictor sets.

For the indirect mapping approach, each EQ-5D-5L dimension was modeled separately as a dependent variable, with SF-12v2 item/subscale scores serving as predictors. To estimate response-level probabilities, generalized ordered logistic (GOLOGIT), ordinal logistic (OLOGIT), or multinomial logistic (MLOGIT) models were evaluated.

GOLOGIT models failed to converge, and OLOGIT models were deemed inappropriate due to violations of the proportional odds assumption, indicating that predictor effects varied across response levels. Consequently, MLOGIT was selected as the final modeling strategy.

Using the most-likely-probability method, MLOGIT generated the most probable response level for each EQ-5D-5L dimension, producing a five-digit health-state profile that was subsequently converted into a utility score using the Thai-specific value set.

Notably, predicted utility scores were truncated at the theoretical boundaries of the Thai EQ-5D-5L value set when estimated values exceeded the feasible utility range. Specifically, values greater than 1 were truncated to the upper bound of 1.0, while values less than –0.4212 were truncated to the lower bound of –0.4212.

Model performance

Model performance was evaluated using ten-fold cross-validation (CV), whereby the dataset was randomly partitioned into 10 approximately equal-sized subsets (n ≈ 200 per fold). In each iteration, nine subsets were used for model estimation, and the remaining subset served as the validation sample. This procedure was repeated 10 times so that each subset functioned as the validation set exactly once.

Predictive accuracy was assessed using the mean absolute error (MAE) and root mean square error (RMSE) across all predictor sets. The optimal model was identified based on the lowest values of these metrics, where discrepancies arose, RMSE was prioritized due to its greater sensitivity to large prediction errors [44].

Model agreement was further evaluated using the intraclass correlation coefficient (ICC) derived from a two-way mixed-effects model with absolute agreement and single-measurement specifications. ICC values were interpreted as poor (ICC < 0.50), moderate (0.50 ≤ ICC <0.75), good (0.75 ≤ ICC <0.90), or excellent (ICC ≥ 0.90) agreement [47].

In accordance with NICE guidelines [48], mean utility score plots were constructed across the full range of WHOQOL-BREF total scores to identify systematic prediction bias. Additionally, cumulative distribution plots were used to examine discrepancies between predicted and observed utility values. The optimal model was expected to demonstrate the highest agreement and minimal divergence across these visual diagnostics. Notably, this study was conducted and reported in accordance with the Mapping onto Preference-Based Measures Reporting Standards checklist [49] and the reporting standards guidance outlined in the 2017 by International Society for Pharmacoeconomics and Outcomes Research (ISPOR) Task Force Report [50], as throuroghly described in Tables in S1-S2 Table.

All statistical analyses were conducted using Stata version 17 (StataCorp LLC, College Station, TX, USA) and Microsoft Excel. A two-sided p-value <0.05 was considered statistically significant.

Results

Participant characteristics

Among the 2,000 participants drawn from the general Thai population, mean SF-12v2 subscale scores ranged from 64.07 ± 24.94 for GH to 86.83 ± 17.89 for RE. The mean PCS and MCS scores were 50.28 ± 8.44 and 53.61 ± 7.11, respectively. Most SF-12v2 subscales exhibited both ceiling effects (scores of 100) and floor effects (scores of 0). An exception was observed for the MH subscale, which demonstrated only a ceiling effect, affecting 14.50% of respondents. Detailed descriptive statistics for participant characteristics and EQ-5D-5L utility scores are reported elsewhere [38].

Conceptual overlap between EQ-5D-5L and SF-12v2

Table 3 presents the absolute Spearman’s rank correlation coefficients between SF-12v2 items and subscale-level scores and EQ-5D-5L utility scores. All pairwise correlations were statistically significant (p < 0.001), except for the correlation between MH1 and SC, MCS, and MO, and MCS and SC.

thumbnail
Table 3. Spearman’s correlation coefficients between SF-12v2 items or subscale scores and the EQ-5D-5L dimensions and utility scores.

https://doi.org/10.1371/journal.pone.0351064.t003

At the subscale level, PF, RP, BP, and GH demonstrated the strongest correlations with EQ-5D-5L utility scores (|r| = 0.5904–0.6109) and showed moderate to strong correlations with at least one EQ-5D-5L dimension. These subscales were therefore retained for inclusion in the predictor sets used in the mapping models. In contrast, VT exhibited the weakest correlation with EQ-5D-5L utility scores (|r| = 0.2879) and was excluded from further analyses.

At the item level, GH1, PF1, PF2, RP1, RP2, RE1, RE2, BP1, and MH2 showed at least moderate correlations with EQ-5D-5L utility scores (|r| > 0.4). However, the inclusion of RP1, RP2, RE1, and RE2 concurrently raised concerns regarding multicollinearity, given their conceptual and statistical overlap within subscales. Consequently, RP2 and RE1 were excluded from the predictor sets, as their correlation coefficients were lower than those of RP1 and RE2, respectively.

Model selection

Table 4 summarizes predictive performance metrics across all regression models for both direct and indirect mapping approaches. For direct mapping, predictor set 4 demonstrated superior performance relative to the other sets, yielding the lowest average MAE and RMSE.

thumbnail
Table 4. Predictive performance of all regression models for predicting EQ-5D-5L utility scores from ten-fold cross-validation.

https://doi.org/10.1371/journal.pone.0351064.t004

Although both Tobit and ALDVMM-1 component models achieved the lowest post-cross-validation (CV) MAE (0.0402 and 0.0410, respectively), the ALDVMM-1 component model produced the lowest post-CV RMSE (0.0625). In contrast, the Tobit model ranked 16th in terms of RMSE among all evaluated models.

Based on overall predicted accuracy, the ALDVMM-1 component model was selected as the best-performing direct mapping model. As illustrated in Fig 1, this model exhibited less variability in predicted utility scores across the full distribution compared with the Tobit model. Although cumulative distribution plots indicated that the Tobit model more closely regressed observed values at the upper boundary (1.0000), this was primarily attributable to truncation at the maximum utility value. In contrast, the ALDVMM-1 component model generated predictions without truncation at either end of the scale. Similar to the distribution plots, ALDVMM-1 component model outperformed the other models because it could predict the utility scores more closely aligned with the observed utility scores, particularly at the maximum utility score of 1.00 (50%) and for values below zero as shown in Figure in S1 Fig.

thumbnail
Fig 1. Comparison of the cumulative plots of observed and predicted utility scores across all regression models for direct and indirect mapping.

ALDVMM-1 component adjusted limited dependent variable mixture model with 1 component CLAD censored least absolute deviation GLM generalized linear model MLOGIT multinomial logistic regression OLS ordinary least squares Trunc truncated predicted utility values for the OLS, Tobit, and CLAD models. Their plots were generated with predicted values truncated at theoretical boundaries of the Thai EQ-5D-5L value set: an upper bound of 1.0 for values exceeding 1 and a lower bound of –0.4212 for values less than –0.4212. Truncation did not affect the comparative assessment of model performance. Non-truncated predictions are shown for GLM-poisson, ALDVMM-1 component, and MLOGIT, which generate predicted utility values within the Thai-specific value set by model structures.

https://doi.org/10.1371/journal.pone.0351064.g001

Consistent with these findings, the ALDVMM-1 component model demonstrated the highest agreement between predicted and observed utility scores, with an ICC of 0.8192, exceeding that of all alternative regression models.

For indirect mapping, all five predictor sets were initially evaluated. However, the GOLOGIT model failed to converge, and the OLOGIT model violated the proportional odds assumption across all predictor sets. Consequently, MLOGIT was implemented.

Model convergence was achieved only for predictor set 1, yielding post-CV MAE and RMSE values of 0.0401 and 0.0703, respectively. Agreement between predicted and observed utility scores was good, with an ICC of 0.7810.

Fig 2 presents plots comparing predicted and observed mean utility scores stratified by WHOQOL-BREF total scores across all regression models. All models exhibited a consistent upward trend in predicted utility scores with increasing WHOQOL-BREF scores. For WHOQOL-BREF scores below 80, the GLM-poisson produced the most accurate predictions, whereas other models tended to overestimate utility.

thumbnail
Fig 2. Mean observed and predicted utility scores by WHOQOL-BREF Total Score as the conditioning variable across all regression models for direct and indirect mapping.

ALDVMM-1 component adjusted limited dependent variable mixture model with 1 component CLAD censored least absolute deviation GLM generalized linear model MLOGIT multinomial logistic regression OLS ordinary least squares Trunc truncated predicted utility values for the OLS, Tobit, and CLAD models. Their plots were generated with predicted values truncated at theoretical boundaries of the Thai EQ-5D-5L value set: an upper bound of 1.0 for values exceeding 1 and a lower bound of –0.4212 for values less than –0.4212. Truncation did not affect the comparative assessment of model performance. Non-truncated predictions are shown for GLM-poisson, ALDVMM-1 component, and MLOGIT, which generate predicted utility values within the Thai-specific value set by model structures.

https://doi.org/10.1371/journal.pone.0351064.g002

For WHOQOL-BREF scores between 80 and 105, the ALDVMM-1 component model demonstrated superior predictive performance compared to other regression models. In contrast, for scores exceeding 105, where data density was lower, the OLS model yielded more accurate predictions than other approaches. Notably, the MLOGIT model systematically overestimated utility scores across the entire WHOQOL-BREF score range.

Algorithm for direct and indirect mapping

Based on the above evaluations, the ALDVMM-1 component model with predictor set 4 and the MLOGIT model with predictor set 1 were identified as the best-predicting models for direct and indirect mapping, respectively.

The mathematical specification for predicting EQ-5D-5L utility scores using the selected direct mapping model is shown below:

where:

( denotes response to SF-12v2 items GH1, PF1, PF2, RP1, RE2, BP1, and SF1 (response levels for PF1 and PF2 ranged from 1–3; all others ranged from 1–5);

represents the estimated coefficient corresponding to the item

Age refers to the respondent’s age in years.

Estimated utility scores were subsequently used to calculate predicted utility values using algorithms derived from the StataCorp LLC implementation [51]. The predicted utility values were subsequently calculated using the following formula.

Where:

: The predicted value of dependent variable (utility scores) for each individual

: The summation operator over C possible categories; however, it is one for this study because the best-predicting model is ALDVMM with one component.

: Cumulative distribution function of the standard normal distribution for upper limit

: Cumulative distribution function of the standard normal distribution for lower limit

: Upper bound is 0.9436 (the second highest values for the Thai-specific value set)

: Lower bound is -0.4212 (the lowest values for the Thai-specific value set)

: The standard deviation of the error term for category c where it is 0.0795847 for

this study

: Estimated utility value from the exploratory variable and regression coefficients

of the STATA output

: Probability density function of the standard normal distribution for upper limit

: Probability density function of the standard normal distribution for lower limit

Detailed coefficients and standard errors are provided in Table in S3 Table, with a worked example illustrated in Text in S1 Text.

Table in S4 Table reports the coefficients for the selected SF-12v2 items used in indirect mapping (predictor set 1), while Text in S2 Text provides a step-by-step example of utility scores calculated via the indirect mapping approach.

The predicted utility scores may be applied in future economic evaluations. The variance-covariance matrices are available in Microsoft Excel format to facilitate probabilistic sensitivity analysis in S2 File.

Discussion

This study represents the first Thai mapping investigation to convert SF-12v2 responses into EQ-5D-5L utility scores using a Thai-specific value in a general population sample, employing both direct and indirect mapping approaches. In doing so, it addresses several methodological limitations of earlier mapping studies, many of which relied on EQ-5D-3L utilities [2531] and predominantly applied OLS regression [26,27,29]. OLS is widely recognized as suboptimal for modeling health utility data because it fails to adequately accommodate ceiling effects and distributional irregularities [34].

Correlation analyses demonstrated substantial conceptual overlap between SF-12v2 items or subscale scores and EQ-5D-5L utility scores and dimensions. Most SF-12v2 subscales exhibited at least moderate correlations with EQ-5D-5L outcomes, except for the VT subscale and item. On this basis, the majority of SF-12v2 subscales and items were retained in the predictor sets. These findings are consistent with previous mapping studies conducted in both general and patient populations internationally, including studies from Thailand [14,5254].

Although the MH subscale showed moderate correlations with AD and EQ-5D-5L utility scores, the MH1 item exhibited only weak correlations with EQ-5D-5L dimensions and utility. Moreover, inclusion of the MH subscale did not improve overall model fit. Consequently, only the MH2 item was retained in the final predictor set. PCS and MCS were excluded from all predictor sets because their inclusions introduced multicollinearity with SF-12v2 subscales (VIF > 0.40) and did not improve predictive performance.

For direct mapping, predictor set 4, which comprised categorical SF-12v2 items (GH1, PF1, PF2, RP1, RE2, BP1, and SF1), demonstrated the best predictive performance across models. Based on post-CV MAE and RMSE, the ALDVMM-1 component, OLS, and Tobit models emerged as the strongest candidates. However, OLS was excluded because it did not adequately account for the bounded, skewed, and heteroscedastic nature of utility data [55,56]. Although Tobit regression is theoretically suitable for censored outcomes, it did not yield the lowest model prediction errors, violated assumptions of normality and homoscedasticity [57], and produced predicted utility values exceeding 1.000, requiring artificial truncation. Consequently, the ALDVMM-1 component model was selected as the best direct mapping approach.

The ALDVMM-1 component model consistently demonstrated superior predictive performance. Its advantage lies in its ability to address key distributional challenges inherent in health utility data, including pronounced ceiling effects, multimodality, and the discontinuity between perfect health and the next-best health state [58]. These findings are consistent with previous mapping studies [25,37,58,59], indicating that mixture models outperform conventional linear models in both general and patient populations. Notably, the present ALDVMM-1 component outperformed earlier SF-12v2 mapping studies that relied on CLAD and simple OLS models based on PCS and MCS scores alone.

Although ALDVMM and beta mixture models are designed to accommodate the multimodality of utility distributions, the ALDVMM-2 component and beta mixture models failed to fully converge for the predictor set 4, which consisted of several SF-12v2 items coded as categorical variables in this study. This may be attributable to the substantial increase in the number of estimated parameters in these mixture models, resulting in instability and convergence difficulties. Furthermore, this mapping algorithm was developed from the predominantly healthy samples leading to sparse responses in certain categorical levels of the SF-12v2 items, particularly for poorer health status. Therefore, utility distribution might not have sufficient multimodality for complex mixture model specifications. ALDVMM-1 component can provide a more stable utility estimation while maintaining good predictive performance.

Whereas previous studies reported a mean prediction error of approximately 0.0744 and an MAE value around 0.14 in general [31] and socioeconomically disadvantaged US populations [27], the current study achieved markedly improved accuracy, with a mean error of 0.0008 and an MAE of 0.0410. These results support the inclusion of SF-12v2 items as exploratory variables rather than relying exclusively on summary component scores. Furthermore, the ALDVMM-1 component outperformed mixture models reported in earlier SF-12v2 mapping studies, despite those models being developed to predict EQ-5D-3L utility in US national samples. Specifically, the present model yielded a substantially lower post-CV RMSE (0.0625) compared with previously reported values ranging from 0.146 to 0.166 [25].

Because SF-12v2 items can be reclassified to derive the SF-6D instrument, the UK valuation algorithm is commonly applied in Thailand in the absence of a Thai-specific value set. Previous Thai studies have supported the use of SF-6D utilities for HRQoL measurement and economic evaluations in both the general [14] and chronic disease populations [52]. However, the present findings did not demonstrate conceptual overlap between the VT1 item and EQ-5D-5L utilities or dimensions. Moreover, agreement between SF-6D–derived and observed EQ-5D-5L utilities was lower (ICC = 0.7300) than that observed between predicted and observed EQ-5D-5L utilities using the proposed algorithm (ICC = 0.8192) with the general Thai samples in this present dataset.

Importantly, the proposed mapping algorithm showed greater sensitivity in poorer health states, predicting utility values as low as −0.2214. This value is substantially lower than the minimum utility score of 0.29 obtainable using the UK SF-6D algorithm [40]. In the absence of a Thai-specific SF-6D value set, the ALDVMM-based mapping equation is therefore recommended for estimating EQ-5D-5L utility scores from SF-12v2 data in general Thai populations.

For indirect mapping, the MLOGIT model with predictor set 1 is recommended. However, indirect mapping did not outperform direct mapping, as reflected by a higher post-CV RMSE (0.0703 vs 0.0625). Distributional and cumulative plots further indicated systematic overprediction of utility values above 0.6. Although OLOGIT is theoretically appropriate for ordinal outcomes [60], it was unsuitable in this study due to violations of the proportional odds assumption and convergence issues. Consequently, MLOGIT was employed to estimate EQ-5D-5L dimensions response probabilities.

The relatively weaker performance of indirect mapping may be attributable to sparse extreme responses within the EQ-5D-5L dimension in this generally healthy sample, leading to inflated standard errors. Similar challenges have been documented in previous mapping studies [30,44]. Despite these limitations, the MLOGIT model demonstrated superior performance compared with earlier US-based studies, achieving a lower MAE (0.0401 vs 0.0480) [30]. This improvement may reflect the enhanced discriminative capacity and reduced ceiling effects of the previous mapping study estimating the EQ-5D-3L utility scores, which are noted to have higher ceiling effects of EQ-5D-5L relative to EQ-5D-3L [30].

Although there are several mapping algorithm generated from general representative samples from other countries, they might not be applicable to Thai population because mapping algorithms are largely population-dependent because both predictor distributions and health-state valuations may vary across countries and cultural contexts. Differences in demographic composition and baseline health status between the present Thai sample and the other populations used to develop their country -based mapping algorithms may therefore influence model estimation and predictive accuracy. In the previous US-based mapping algorithm from the US representative samples [31], two key differences may contribute to the observed performance gap between the two algorithms. First, the US-derived algorithms were developed using a sample in which the majority of respondents presenting with more chronic conditions (average number of chronic conditions = 1.90), compared with approximately 0.37 in the Thai derivation sample. Therefore, the US-derived algorithms may predict utility scores more accurately among unhealthy respondents or those at the lower end of the health spectrum than the Thai mapping algorithm. This explanation is supported by the lower mean utility score estimated using the US-based algorithm (0.8912), compared with that estimated using the Thai mapping algorithm (0.9222). Second, cultural differences between these two countries can influence how respondents interpret and report HRQoL items because health perception is inherently shaped by sociocultural context. Previous study has demonstrated that certain culture-specific health dimensions perceived by the Thai population differ from those reported in Western populations [61]. As a result, these factors may help explain why a Thai-specific mapping algorithm is required for estimating utility scores in the Thai population. Nevertheless, the performance of the Thai mapping algorithm should be further validated using an independent dataset from the general Thai population, particularly among unhealthy individuals or clinical population.

Several limitations should be acknowledged. First, this study sample comprised predominantly healthy individuals from the general population, resulting in a sparse representation of severe health states across EQ-5D-5L dimensions. Consequently, the mapping algorithm may be less accurate for populations with poorer health. Second, although 10-fold CV was conducted to evaluate the model performance and reduce the overfitting of the mapping algorithm, the external validation using an independent dataset was not performed. Future studies should validate and refine this proposed mapping algorithm in less healthy samples or clinical population before broader application in economic analyses.

Conclusions

This study developed a novel mapping algorithm to convert SF-12v2 responses into EQ-5D-5L utility scores for the general Thai population. While earlier Thai studies have supported the application of the UK valuation algorithm to derive SF-6D utility score, the present findings challenge this practice. Specifically, correlation analyses did not demonstrate sufficient conceptual overlap between the vitality subscale or item and EQ-5D-5L dimension or scores. In addition, agreement between SF-6D–derived and observed EQ-5D-5L utility scores was lower than that observed between the mapped and observed EQ-5D-5L utility generated by the proposed algorithm.

Despite these strengths, an important limitation should be acknowledged. The proposed mapping algorithm may have reduced accuracy in predicting utility scores among individuals with poor health status, reflecting the limited representation of severe health states in the study sample. Accordingly, caution is warranted when applying this algorithm to populations with substantial morbidity, and further validation in less healthy or clinical samples is recommended.

Supporting information

S1 Table. Guidelines and checklist for mapping onto Preference-Based Measures Standard (MAPS) checklist.

https://doi.org/10.1371/journal.pone.0351064.s001

(DOCX)

S2 Table. Checklist for 2017 ISPOR Good Practices Report: Mapping to Estimate Health-State Utility from Non-Preference-Based Outcome Measures.

Summary of reporting of mapping studies recommendations.

https://doi.org/10.1371/journal.pone.0351064.s002

(DOCX)

S3 Table. Coefficients and standard errors of the final model used for direct mapping based on the Thai value set.

https://doi.org/10.1371/journal.pone.0351064.s003

(DOCX)

S4 Table. Coefficients and standard errors of the final model used for indirect mapping.

https://doi.org/10.1371/journal.pone.0351064.s004

(DOCX)

S1 Text. Instructions for predicting utility scores using the direct mapping algorithm.

https://doi.org/10.1371/journal.pone.0351064.s005

(DOCX)

S2 Text. Instructions for predicting utility scores using the indirect mapping algorithm.

https://doi.org/10.1371/journal.pone.0351064.s006

(DOCX)

S1 Fig. Comparison of the distributions of observed and predicted utility scores across all regression models for direct and indirect mapping.

https://doi.org/10.1371/journal.pone.0351064.s007

(DOCX)

S2 File. Variance-Covariance matrix for both direct and indirect mapping algorithms.

https://doi.org/10.1371/journal.pone.0351064.s009

(XLSX)

References

  1. 1. Chen Y. Health technology assessment and economic evaluation: Is it applicable for the traditional medicine?. Integr Med Res. 2022;11(1):100756. pmid:34401322
  2. 2. J Asunmonu O. Strategic Resource Allocation in Modern Health Systems: A Systematic Review of the Role and Impact of Health Technology Assessments (HTAs) on Healthcare Financing and Policy Decisions. IJMCR. 2025;4(3):51–5.
  3. 3. Tanvejsilp P, Taychakhoonavudh S, Chaikledkaew U, Chaiyakunapruk N, Ngorsuraches S. Revisiting Roles of Health Technology Assessment on Drug Policy in Universal Health Coverage in Thailand: Where Are We? And What Is Next? Value Health Reg Issues. 2019;18:78–82. https://doi.org/10.1016/j.vhri.2018.11.004
  4. 4. Wani S, Alsabti H, Allamki S, Almandhari A, Alhajji S, Alrashdi I, et al. Methodological guidelines for Health Technology Assessment in Oman. J Pharm Policy Pract. 2025;18(1):2596523. pmid:41384030
  5. 5. Rowen D, Azzabi Zouraq I, Chevrou-Severac H, van Hout B. International regulations and recommendations for utility data for health technology assessment. Pharmacoeconomics. 2017;35(Suppl 1):11–9. https://doi.org/10.1007/s40273-017-0544-y
  6. 6. Botwright S, Sittimart M, Chavarina KK, Bayani DB, Merlin T, Surgey G, et al. Good Practices for Health Technology Assessment Guideline Development: A Report of the Health Technology Assessment International, HTAsiaLink, and ISPOR Special Task Force. Value in Health. 2025;28(1):1–15. https://doi.org/10.1016/j.jval.2024.09.001
  7. 7. Rai M, Goyal R. Pharmacoeconomics in Healthcare. Pharmaceutical Medicine and Translational Clinical Research. Boston: Academic Press. 2018. 465–72.
  8. 8. Finch AP, Brazier JE, Mukuria C. What is the evidence for the performance of generic preference-based measures? A systematic overview of reviews. Eur J Health Econ. 2018;19(4):557–70. pmid:28560520
  9. 9. Wolowacz SE, Briggs A, Belozeroff V, Clarke P, Doward L, Goeree R, et al. Estimating Health-State Utility for Economic Models in Clinical Studies: An ISPOR Good Research Practices Task Force Report. Value Health. 2016;19(6):704–19. pmid:27712695
  10. 10. Meregaglia M, Nicod E, Drummond M. The estimation of health state utility values in rare diseases: do the approaches in submissions for NICE technology appraisals reflect the existing literature? A scoping review. Eur J Health Econ. 2023;24(7):1151–216. pmid:36335234
  11. 11. Pattanaphesaj J, Thavorncharoensap M, Ramos-Goni JM, Tongsiri S, Ingsrisawang L, Teerawattananon Y. The EQ-5D-5L Valuation study in Thailand. Expert Rev Pharmacoecon Outcomes Res. 2018;18(5):551–8. https://doi.org/10.1080/14737167.2018.1494574
  12. 12. Németh G. Health related quality of life outcome instruments. Eur Spine J. 2006;15 Suppl 1(Suppl 1):S44-51. pmid:16320032
  13. 13. Kangwanrattanakul K, Phimarn W. A systematic review of the development and testing of additional dimensions for the EQ-5D descriptive system. Expert Rev Pharmacoecon Outcomes Res. 2019;19(4):431–43. pmid:31244348
  14. 14. Kangwanrattanakul K. A comparison of measurement properties between UK SF-6D and English EQ-5D-5L and Thai EQ-5D-5L value sets in general Thai population. Expert Rev Pharmacoecon Outcomes Res. 2021;21(4):765–74. pmid:32981380
  15. 15. Sakthong P, Sonsa-Ardjit N, Sukarnjanaset P, Munpan W. Psychometric properties of the EQ-5D-5L in Thai patients with chronic diseases. Qual Life Res. 2015;24(12):3015–22. pmid:26048348
  16. 16. Turner N, Campbell J, Peters TJ, Wiles N, Hollinghurst S. A comparison of four different approaches to measuring health utility in depressed patients. Health Qual Life Outcomes. 2013;11:81. pmid:23659557
  17. 17. Haywood KL, Garratt AM, Dziedzic K, Dawes PT. Generic measures of health-related quality of life in ankylosing spondylitis: reliability, validity and responsiveness. Rheumatology (Oxford). 2002;41(12):1380–7. pmid:12468817
  18. 18. Kangwanrattanakul K. Mapping of the World Health Organization Quality of Life Brief (WHOQOL-BREF) to the EQ-5D-5L in the General Thai Population. Pharmacoecon Open. 2023;7(1):139–48. https://doi.org/10.1007/s41669-022-00380-0
  19. 19. Sakthong P. Mapping World Health Organization Quality of Life-BREF Onto 5-Level EQ-5D in Thai Patients With Chronic Diseases. Value Health. 2021;24(8):1089–94. pmid:34372973
  20. 20. Młyńczak K, Golicki D. Psychometric properties of the Polish version of SF-12v2 in the general population survey. Expert Rev Pharmacoecon Outcomes Res. 2022;22(3):465–72. pmid:33941017
  21. 21. Younsi M. Health-Related Quality-of-Life Measures: Evidence from Tunisian Population Using the SF-12 Health Survey. Value Health Reg Issues. 2015;7:54–66. pmid:29698153
  22. 22. Kangwanrattanakul K. Factor structure and trends in SF-12v2 health-related quality of life scores among pre-and post-pandemic samples in Thailand: confirmatory factor analysis and Rasch analysis. Health Qual Life Outcomes. 2025;23(1):74. pmid:40691581
  23. 23. Brazier JE, Kolotkin RL, Crosby RD, Williams GR. Estimating a preference-based single index for the Impact of Weight on Quality of Life-Lite (IWQOL-Lite) instrument from the SF-6D. Value Health. 2004;7(4):490–8. pmid:15449641
  24. 24. Oliveira Gonçalves AS, Werdin S, Kurth T, Panteli D. Mapping Studies to Estimate Health-State Utilities From Nonpreference-Based Outcome Measures: A Systematic Review on How Repeated Measurements are Taken Into Account. Value Health. 2023;26(4):589–97. pmid:36371289
  25. 25. Coca Perraillon M, Shih Y-CT, Thisted RA. Predicting the EQ-5D-3L Preference Index from the SF-12 Health Survey in a National US Sample: A Finite Mixture Approach. Med Decis Making. 2015;35(7):888–901. pmid:25840902
  26. 26. Franks P, Lubetkin EI, Gold MR, Tancredi DJ, Jia H. Mapping the SF-12 to the EuroQol EQ-5D Index in a national US sample. Med Decis Making. 2004;24(3):247–54. pmid:15185716
  27. 27. Franks P, Lubetkin EI, Gold MR, Tancredi DJ. Mapping the SF-12 to preference-based instruments: convergent validity in a low-income, minority population. Med Care. 2003;41(11):1277–83. pmid:14583690
  28. 28. Gray AM, Rivero-Arias O, Clarke PM. Estimating the association between SF-12 responses and EQ-5D utility values by response mapping. Med Decis Making. 2006;26(1):18–29. pmid:16495197
  29. 29. Lawrence WF, Fleishman JA. Predicting EuroQoL EQ-5D preference scores from the SF-12 Health Survey in a nationally representative sample. Med Decis Making. 2004;24(2):160–9. pmid:15090102
  30. 30. Le QA. Probabilistic mapping of the health status measure SF-12 onto the health utility measure EQ-5D using the US-population-based scoring models. Qual Life Res. 2014;23(2):459–66. pmid:24026631
  31. 31. Sullivan PW, Ghushchyan V. Mapping the EQ-5D index from the SF-12: US general population preferences in a nationally representative sample. Med Decis Making. 2006;26(4):401–9. pmid:16855128
  32. 32. Buchholz I, Janssen MF, Kohlmann T, Feng Y-S. A Systematic Review of Studies Comparing the Measurement Properties of the Three-Level and Five-Level Versions of the EQ-5D. Pharmacoeconomics. 2018;36(6):645–61. pmid:29572719
  33. 33. Yu L, Yang H, Lu L, Fang Y, Zhang X, Li S, et al. Developing mapping algorithms to predict EQ-5D health utility values from Bath Ankylosing Spondylitis Disease Activity Index and Bath Ankylosing Spondylitis Functional Index among patients with Ankylosing Spondylitis. Health Qual Life Outcomes. 2024;22(1):61. pmid:39113080
  34. 34. Huang I-C, Frangakis C, Atkinson MJ, Willke RJ, Leite WL, Vogel WB, et al. Addressing ceiling effects in health status measures: a comparison of techniques applied to measures for people with HIV disease. Health Serv Res. 2008;43(1 Pt 1):327–39. pmid:18211533
  35. 35. Fang H, Hong T, Liu X, Luo C, Hou Y, Xie S. Mapping the ADDQoL to the EQ-5D-5L and SF-6Dv2 among Chinese patients with type 2 diabetes mellitus. Health Qual Life Outcomes. 2025;23(1):46. pmid:40307906
  36. 36. Chen Z, Yang L, Zhang Y. Mapping KDQOL-36 Onto EQ-5D-5L and SF-6Dv2 in Patients Undergoing Dialysis in China. Value Health Reg Issues. 2025;48:101103. https://doi.org/10.1016/j.vhri.2025.101103
  37. 37. Aghdaee M, Gu Y, Sinha K, Parkinson B, Sharma R, Cutler H. Mapping the Patient-Reported Outcomes Measurement Information System (PROMIS-29) to EQ-5D-5L. Pharmacoeconomics. 2023;41(2):187–98. pmid:36336773
  38. 38. Kangwanrattanakul K, Krägeloh CU. EQ-5D-3L and EQ-5D-5L population norms for Thailand. BMC Public Health. 2024;24(1):1108. https://doi.org/10.1186/s12889-024-18391-3
  39. 39. Ware JE Jr, Kosinski M, Keller SD. SF-12: how to score the SF-12 physical and mental health summary scales Second ed. Boston, MA: The Health Institute, New England Medical Center; 1995.
  40. 40. Brazier JE, Roberts J, Deverill M. The estimation of a preference-based measure of health from the SF-36. J Health Econ. 2002;2(21):271–92. pmid:11939242
  41. 41. Kangwanrattanakul K, Krägeloh CU. Psychometric evaluation of the WHOQOL-BREF and its shorter versions for general Thai population: confirmatory factor analysis and Rasch analysis. Qual Life Res. 2024;33(2):335–48. pmid:37906345
  42. 42. Akoglu H. User’s guide to correlation coefficients. Turk J Emerg Med. 2018;18(3):91–3. pmid:30191186
  43. 43. Nguyen LH, Tran BX, Hoang Le QN, Tran TT, Latkin CA. Quality of life profile of general Vietnamese population using EQ-5D-5L. Health Qual Life Outcomes. 2017;15(1):199. pmid:29020996
  44. 44. Tan YJ, Ong SC. Direct and indirect mapping of assessment of quality of life - 6 dimensions (AQoL-6D) onto EQ-5D-5L utilities using data from a multicenter, cross-sectional study of Malaysians with chronic heart failure. Value in Health. 2024;27(12):1762–70. https://doi.org/10.1016/j.jval.2024.07.016
  45. 45. Zhou J, Williams C, Keng MJ, Wu R, Mihaylova B. Estimating costs associated with disease model states using generalized linear models: a tutorial. Pharmacoeconomics. 2024;42(3):261–73. https://doi.org/10.1007/s40273-023-01319-x
  46. 46. Riani M, Atkinson AC, Corbellini A. Automatic robust Box–Cox and extended Yeo–Johnson transformations in regression. Stat Methods Appl. 2022;32(1):75–102.
  47. 47. Koo TK, Li MY. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. J Chiropr Med. 2016;15(2):155–63. pmid:27330520
  48. 48. Allan WMHA, Pudney S. Mapping to Estimate Health State Utilities 2023 https://www.sheffield.ac.uk/media/42422/download?attachment
  49. 49. Petrou S, Rivero-Arias O, Dakin H, Longworth L, Oppe M, Froud R, et al. Preferred reporting items for studies mapping onto preference-based outcome measures: The MAPS statement. Health Qual Life Outcomes. 2015;13:106. pmid:26232268
  50. 50. Wailoo AJ, Hernandez-Alava M, Manca A, Mejia A, Ray J, Crawford B, et al. Mapping to Estimate Health-State Utility from Non-Preference-Based Outcome Measures: An ISPOR Good Practices for Outcomes Research Task Force Report. Value Health. 2017;20(1):18–27. pmid:28212961
  51. 51. Alava MH, Wailoo A. Fitting Adjusted Limited Dependent Variable Mixture Models to EQ-5D. The Stata Journal: Promoting communications on statistics and Stata. 2015;15(3):737–50.
  52. 52. Sakthong P, Munpan W. A Head-to-Head Comparison of UK SF-6D and Thai and UK EQ-5D-5L Value Sets in Thai Patients with Chronic Diseases. Appl Health Econ Health Policy. 2017;15(5):669–79. pmid:28290106
  53. 53. De Smedt D, Clays E, Annemans L, De Bacquer D. EQ-5D versus SF-12 in coronary patients: are they interchangeable?. Value Health. 2014;17(1):84–9. 10.1016/j.jval.2013.10.010
  54. 54. Xie S, Wu J, Chen G. Comparative performance and mapping algorithms between EQ-5D-5L and SF-6Dv2 among the Chinese general population. Eur J Health Econ. 2024;25(1):7–19. pmid:36709458
  55. 55. Longworth L, Rowen D. Mapping to obtain EQ-5D utility values for use in NICE health technology assessments. Value Health. 2013;16(1):202–10. pmid:23337232
  56. 56. Lamu AN, Olsen JA. Testing alternative regression models to predict utilities: mapping the QLQ-C30 onto the EQ-5D-5L and the SF-6D. Qual Life Res. 2018;27(11):2823–39. pmid:30173314
  57. 57. Cunillera O. Tobit models. In: Maggino F. Encyclopedia of Quality of Life and Well-Being Research. Cham: Springer International Publishing. 2023. 7237–42.
  58. 58. Worboys HM, Gray L, Burton J, Alava MH, Greenwood S, Cooper N. Mapping the Kidney Disease Quality-of-Life Questionnaire Onto the EQ-5D-5L Utility Index in Patients Undergoing Hemodialysis. Value Health. 2025;28(7):1091–9. pmid:40246068
  59. 59. Gray LA, Hernandez Alava M, Wailoo AJ. Mapping the EORTC QLQ-C30 to EQ-5D-3L in patients with breast cancer. BMC Cancer. 2021;21(1):1237. pmid:34794404
  60. 60. Freitas Souza R de, Lima FG, Corrêa HL. Multilevel Ordinal Logit Models: A Proportional Odds Application Using Data from Brazilian Higher Education Institutions. Axioms. 2024;13(1):47.
  61. 61. Mao Z, Ahmed S, Graham C, Kind P, Sun Y-N, Yu C-H. Similarities and Differences in Health-Related Quality-of-Life Concepts Between the East and the West: A Qualitative Analysis of the Content of Health-Related Quality-of-Life Measures. Value Health Reg Issues. 2021;24:96–106. pmid:33524902