Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

An evaluation of strategies commonly used by health advocate programs

  • Jingyao Huang ,

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Validation, Visualization, Writing – original draft, Writing – review & editing

    jhhkq@umkc.edu

    Affiliation Bloch School of Business, University of Missouri-Kansas City, Kansas City, Missouri, United States of America

  • Diwakar Gupta

    Roles Conceptualization, Data curation, Formal analysis, Investigation, Methodology, Project administration, Resources, Software, Supervision, Validation, Visualization, Writing – original draft, Writing – review & editing

    Affiliation McCombs School of Business, University of Texas, Austin, Texas, United States of America

Abstract

Many non-urgent and routine medical procedures, such as MRI and CT Scans, are performed under standardized industry protocols with minimal process variations. However, large price variations exist among providers offering these services within the same geographic region. Many insurers and independent vendors have implemented health advocate/concierge programs to steer beneficiaries to high-quality and low-cost providers. These programs also aim to reduce costs. The Blue Cross Blue Shield (BCBS) of Texas’ Benefits-Value-Advisor (BVA) program is one of the most cited examples. It uses three key strategies: recommendation, unconditional monetary rewards, and persuasion. However, their effectiveness in influencing beneficiaries’ choices, particularly for routine medical procedures, has not been tested. This study fills the gap through a behavioral experiment, which provides insights into how beneficiaries may respond to these strategies and providing guidance for health plan managers. A full-factorial-between-subjects experiment was conducted with 500 subjects recruited through Amazon Mechanical Turk. The experiment design included three treatments: recommendation, copay waiver, and persuasion, each with two levels (Yes or No), resulting in eight distinct scenarios. Subjects were randomly assigned to one of eight scenarios. The survey data were analyzed using a logit regression model. We found that subjects who received recommendations were 32.7% more likely to select the lowest-cost provider and had a 27.3% greater chance of choosing lower-cost providers. However, the effectiveness of recommendations diminished when subjects mistrusted their insurance companies. Neither copay waiver nor persuasion had a significant impact on provider choice. We acknowledge that these findings are derived from a hypothetical experimental setting using an online sample, and that actual beneficiary behavior in clinical settings may differ due to additional factors such as physician referrals, appointment availability, urgency of medical needs and convenience. Our results should be interpreted as providing suggestive evidence rather than directly generalizable predictions of real-world behavior.

Introduction

Many non-urgent and routine medical procedures, such as MRI and CT Scans, adhere to standardized industry protocols, resulting in minimal variation in quality across providers. However, large price variations exist among providers that offer these services within the same geographic region [1,2]. For example, in Austin, Texas area, the price of a cervical spine MRI procedure, varied from $869 to $5,955 in 2025 [3]. These characteristics of the health services market lead to high costs for insurers, while beneficiaries do not necessarily receive higher quality care. Insurers in recent years have implemented strategies to encourage beneficiaries to utilize high-value (low-cost, high-quality) care providers. Similar programs are also offered by independent vendors. This study concerns a health advocate program called the Benefits-Value-Advisor (BVA) program, which has been introduced by the Blue Cross Blue Shield (BCBS) of Texas on behalf of self-insured employers like the University of Texas (UT). It is one of the most cited insurer-led program of its type. The program targets non-urgent, shoppable medical services such as MRI and CT scans, colorectal screenings, and ultrasounds, with the goal of lowering long-term health plan costs for such procedures [4]. We provide a summary table of other similar programs by insurers and independent vendors in S1 File. The BVA program steers beneficiaries to utilize high-value providers, thereby putting pressure on all the providers to reduce prices and improve quality.

While the BVA program has utilized recommendation, copay waiver, and persuasion to achieve its goal, the effectiveness of strategies it employs remains untested. The authors obtained aggregate data from BCBS for MRI and CT scans, which is described in S2 File. The data show that beneficiaries choose lower cost providers after the program was implemented. However, the data are not sufficiently granular to allow a comparison of the effectiveness of different strategies, either independently or in combination with other strategies. This paper concerns an online experiment that was conducted through Amazon’s Mechanical Turk (MTurk) platform to evaluate the effectiveness of the three approaches commonly used by BVA agents to affect beneficiaries’ choices.

The goal of this study is to provide insights into how insurer-led interventions may influence beneficiary decision-making in BVA-like health advocate programs, thereby filling a gap in knowledge. The specific research questions that this study answers are as follows.

  • Keeping provider quality and beneficiaries’ out-of-pocket costs the same across all providers, what proportion of beneficiaries change their initial choice of provider by switching to either the lowest-cost or one of lower-cost providers?
  • Which of the three strategies, either independently or in combination with others, is most effective in increasing the proportion of beneficiaries switching to lower-cost providers?
  • Among individual characteristics, such as beneficiaries’ familiarity with health insurers costs, and their level of trust of their insurers, which factors have the greatest impact on beneficiaries’ choices?
  • What insights do the experiment’s results provide for improving the effectiveness of BVA-like programs?

The functioning of the BVA program is illustrated in Fig 1. Beneficiaries often select high-cost providers, in part because such providers have high brand visibility, while BVA agents steer them toward low-cost options (see S2 File for details). Beneficiaries are encouraged to utilize the BVA program by calling a program agent (hereafter agent) before selecting their service providers for non-urgent medical procedures. They may have an initial preference for a provider, say Provider A, which is referred to as the requested provider. Based on the beneficiary’s needs, location, and travel constraints, the agent may recommend a different provider, say Provider B, referred to as the recommended provider. This type of intervention is referred to as the recommendation treatment in this paper. However, beneficiaries retain full autonomy: they may select the requested provider, the recommended provider, or an entirely different option, with no penalty for disregarding advice from the agent. The final choice is referred to as the selected provider.

thumbnail
Fig 1. An example of the interaction between an agent and a beneficiary in the BVA program.

https://doi.org/10.1371/journal.pone.0350645.g001

BCBS has offered the BVA service to beneficiaries enrolled in the UT Select Plan since 2014. UT pays BCBS a fixed fee of $10 per member per month to subscribe to the BVA services, regardless of usage, and it is also responsible for the actual medical expenditures of the beneficiaries. As a self-insured employer, UT assumes the financial risk, while BCBS serves as the plan administrator. In this paper, the term insurer is used to refer to both entities. In addition to making recommendations, the BVA may offer copay waiver and/or utilize persuasion by testimonials. These strategies are described next.

UT Select beneficiaries typically pay a fixed copay (e.g., $100) for diagnostic imaging. Alternative payment structures are possible. Those are described in S3 File. To incentivize participation, BCBS waives the copay for plan members who consult an agent before selecting an MRI provider. This waiver is unconditional, i.e., beneficiaries receive it simply by using the BVA program, regardless of their chosen provider. It is offered only for MRI procedures. In addition to inducing a higher participation rate, insurers may anticipate that beneficiaries will reciprocate by selecting low-cost providers, offsetting the program’s expense. This type of intervention is referred to as the copay waiver treatment in this paper. Persuasion by testimonials is the third type of treatment. It entails the agents informing beneficiaries about the positive experience of other beneficiaries with the recommended providers. All three strategies are described in detail in the Methods section.

There have been attempts in the literature to explain the large spread in prices for shoppable medical services. It has been argued that either due to lack of awareness, or the absence of price transparency tools, or high search costs, beneficiaries are often unaware of the fees providers charge their insurers. Additionally, the moral hazard inherent in health insurance provides little incentive for beneficiaries to choose low-cost providers. Therefore, they may select providers based on the strength of their doctors’ referrals and brand recognition. Absent pressure from beneficiaries, prices are mostly influenced by factors such as market concentration and providers’ bargaining power in negotiations with insurers [5,6]. Viewed from this perspective, the BVA program helps bring about awareness of prices that insurers pay, reduces search costs, and employs a combination of recommendation, copay waiver, and persuasion to influence those who might not have a strong preference for any particular provider. There are also studies in the literature of other approaches used by insurers to lower costs for shoppable medical services. We describe those studies and their findings in comparison with the findings of our paper in the Discussion section. To the best of our knowledge, no previous study compares and contrasts the effectiveness of strategies utilized by the BVA program.

For standardized medical procedures, higher prices do not necessarily reflect higher quality. Prices are negotiated across a multitude of services that a provider offers. A provider may offer a low overall price for a bundle, e.g., bundled charges for knee joint replacement surgery, while having a high price for a particular procedure such as MRI scan that is included in the bundle. When an agent recommends a low-cost provider for a shoppable medical procedure, that provider is not necessarily an inferior option. It is likely that the provider has a lower negotiated price for that procedure (see S4 File for more details).

Methods

Ethical approval

This study was reviewed and exempted by the University of Texas (UT) Institutional Review Board (IRB) (Study number: 2019030104). The IRB stated “The IRB determined that this protocol meets the criteria for exemption from IRB review under 45 CFR 46.104 (4) Secondary research on data or specimens (no consent required).”.

Study sample & inclusion criteria

We recruited 500 subjects through Amazon Mechanical Turk (MTurk) between June 3–8, 2020. To ensure a consistent health insurance context, participants were restricted to be U.S. residents aged 25 or older. The age restriction was imposed to exclude individuals likely covered under parental insurance plans and to ensure familiarity with personal health insurance decisions.

To maximize sample heterogeneity, the experiment was conducted in four batches at different times of day and on different days of the week. Participants were compensated $2 for completing a 10–15 minute survey. The study had a completion rate of 88.75%, indicating high data quality and reliability [7,8]. Additional details on sample size selection and experiment completion rate are provided in S6 File.

Experiment design

We employ a full-factorial between-subjects design with three treatments: recommendation, copay waiver, and persuasion. Each treatment has two levels: 1 (Treatment = Yes) and 0 (Treatment = No), resulting in eight experimental scenarios shown in Fig 2. Subjects were randomly assigned to one scenario.

Experimental procedures

The experiment followed a structured sequence. All survey materials and detailed question wording are provided in S5 File to facilitate replication.

  1. Consent form. Subjects provided informed consent electronically prior to participating in the study. Subjects who did not provide consent were not asked to provide additional responses.
  2. Background information. Subjects received a brief explanation of MRI procedures.
  3. Demographics. Participants reported demographic information including gender, and categories of income, age, race, education status, employment status, English comprehension and whether subjects had insurance or not. Details on the categories of the demographic variables are provided in the survey instrument.
  4. Insurance cost-sharing comprehension quiz. Subjects completed a quiz composed of one or two questions assessing their understanding of copay and coinsurance. Subjects who answered the first question correctly passed the quiz immediately. If incorrect, they were shown the correct answer and given a second question. A correct response on the second question resulted in a “Pass”; otherwise, the subject was designated as “Not Pass.” All subjects proceeded with the survey regardless of designation (Pass or Not Pass), which was not revealed to them.
  5. Attention check. Subject was randomly assigned to one of four Winograd schema questions [9], which served as an attention check to identify careless respondents and assess data quality prior to analysis.
  6. Scenario assignment & subjects’ provider selection decision. Subjects were randomly assigned to one of eight scenarios, which used text to present the information provided by the agent and explain the provider selection problem. Participants selected a provider based on presented information. Descriptions of the treatments and scenarios are shown in Table 1.
  7. Post-scenario questions. Subjects reported reasons for their choice by answering the question “Why did you choose X clinic?.” In addition, they rated their level of trust in their insurance company on a 7-point Likert scale by indicating how strongly they agreed with the statement “I trust my insurance company,” where 1 represents strongly disagree and 7 represents strongly agree.

The interactions between agents and beneficiaries were modeled as text-based scenarios describing the provider selection problem and information provided by agents. The scenarios were based on the actual setting of the BVA program, with the patient’s out-of-pocket for the MRI set at $100. MRI procedures were chosen as they were the shoppable medical services targeted by the BVA program because of large variation in provider costs for these procedures.

Additionally, neutral names (Peach, Orange, Apple and Pear Clinic) were used for service providers to ensure that subjects would neither recognize nor associate provider names with service quality. The four clinics differed only in price, which followed the order: Peach Clinic > Orange Clinic > Apple Clinic > Pear Clinic. A five-star rating system was used to capture the service quality, as it is the most widely adopted format for evaluating healthcare providers. Quality ratings were held constant across all providers, as the BCBS web site had many providers with nearly identical five-star ratings despite substantial price differences. The provider names, prices, and quality ratings did not vary across all eight scenarios. In addition, the requested provider was set as a relatively high-cost provider (i.e., Orange Clinic, the second highest provider), which was common in BCBS data and consistent with the pattern that high cost providers tend to have greater market visibility.

Dependent measures

The analysis focuses on two binary measures: 1) whether subjects chose the lowest-cost provider, i.e., Pear Clinic, and 2) whether subjects chose a provider that was lower-cost than the requested provider, i.e., whether it was either Apple or Pear Clinic. The second measure reflects the real-world success metric used by BCBS, where any cost reduction relative to beneficiaries’ initial requests is deemed a success. The first measure is aligned with the recommendation treatment, which explicitly guides patients toward the best-value option. Note that the Pear Clinic is the best-value option given the identical quality ratings across all four providers. We refer to the first measure as M1 and the second as M2.

Descriptive statistical analysis

The descriptive analysis employed unpaired sample t-tests, tests, and Fisher’s exact tests. To ensure a balanced sample across the eight scenarios, tests were conducted for each demographic variable. Next, provider selection outcomes were compared across treatment levels to descriptively examine the impact of treatments on the proportions of subjects selecting either the lowest-cost provider or any lower-cost provider. The analysis also measured subjects’ trust in their insurance companies using their responses to the statement, “I trust my insurance company.” Based on these responses, subjects were categorized into two groups: (1) Mistrust (i.e., Strongly disagree, Disagree, Somewhat disagree), and (2) Do Not Mistrust (i.e., Neither agree or disagree, Somewhat agree, Agree, Strongly agree). Using this classification, we examined whether mistrust moderates the treatment effect by analyzing the association between recommendations and the selection of a lowest- or lower-cost provider. Additionally, we analyzed the association between subjects’ understanding of cost-sharing terms (as measured by quiz performance) and their provider selection.

Statistical methods

The primary statistical analysis tool employed is logit regression. A logit specification is common in settings where the dependent variable is binary. The analysis begins with a preliminary model examining the overall effects of the three treatments. The specification is shown in Eq (1). All regression variables are defined in Table 2. Coefficients , and correspond to the treatments: recommendation, copay waiver and persuasion by testimonial, respectively. Parameter is a vector of coefficients associated with control variables Ci, which includes subject i’s gender, insurance status, income, age, race, education, employment status, and English proficiency.

(1)
thumbnail
Table 2. Definitions of variables used in the regression analysis.

https://doi.org/10.1371/journal.pone.0350645.t002

Next, the model in Eq (2) is fitted to the data. This model incorporates all interaction terms. Coefficients and are associated with the subject’s mistrust level and whether they pass the quiz on the meaning of health insurance cost-sharing terms (i.e., co-pay and co-insurance). Parameters , and are the coefficients for the interaction terms between mistrust level and the three treatments, respectively. For subject i, the vector includes the between-treatment interaction terms, i.e., , , and . The vector is a vector of coefficients associated with interaction terms in .

(2)

The inclusion of subjects’ mistrust level and their quiz performance in the comprehension of co-pay and co-insurance is based on the observed correlation between these variables and provider selection decision (see Descriptive statistics).

Robustness checks

Four robustness checks are implemented to validate the findings. The first analysis uses results of the Winograd questions. The primary analysis does not screen out the respondents who fail the Winograd attention check, thus two additional analyses are performed as a robustness check: (1) fitting the model only to responses from subjects who pass the attention check, and (2) including a control variable indicating whether a subject passes the attention check. Second, to address the issue of imbalanced data caused by much fewer respondents mistrusting their insurance company, an additional analysis employs an oversampling approach via bootstrapping. To mitigate overfitting concerns in this exercise, the minority class (respondents expressing mistrust) is progressively oversampled at twice, three times, four times, and five times its original size, with the regression models re-estimated at each increment. Third, a counterfactual analysis compares the experimental results and the claims data in aggregate (see S12 File). We estimate a multinomial logit model using the experimental data, with provider choice among the four providers as the dependent variable and the same set of covariates as in Eq (2). We then use the estimated model to predict the proportion of subjects selecting lower-cost providers and compare these predictions with the actual proportion of beneficiaries choosing lower-cost providers observed in the claims data. Lastly, in addition to the logit specification, we also estimate probit specification of Eqs (1) and (2), which are otherwise identically specified.

Results

Descriptive statistics

Demographic profiles Two observations from subjects with the same IP address were dropped, resulting in 498 observations for analysis. The number of observations for Scenarios 1–8 were 58, 59, 65, 63, 62, 69, 60, and 62, respectively. In accordance with ethical guidelines, we report only aggregated summary statistics for the demographic variables. Among the subjects, 50.6% were female, 81.5% were white, and 86.94% were aged 25–54. Additionally, 78.9% were employed full-time or part-time, 86% had healthcare insurance, and 65.7% held a bachelor’s degree or higher. The demographic profile of the online sample aligned with other reported survey samples from MTurk [10], which were younger and more educated than national average.

The tests were performed for each demographic variable across the eight scenarios. The results showed no significant differences in subjects’ demographic characteristics, confirming homogeneous subject populations across the eight scenarios.

Subjects’ provider selection Provider selection results are summarized in Table 3. We focus on two measures: (i) the proportion of subjects choosing the lowest-cost provider (Pear Clinic, M1), and (ii) the proportion choosing lower-cost providers (Apple or Pear Clinic, M2).

The descriptive results reveal several patterns. First, recommendation has a strong positive effect on the choice of the lowest-cost provider. Relative to the control group (Scenario 1), the proportion of subjects selecting the lowest-cost provider increases substantially when there is a recommendation. A similar pattern is observed for M2. In contrast, neither copay waiver nor persuasion by itself appears to meaningfully influence either measure. In both cases, the proportion of subjects selecting the lowest-cost provider or lower-cost options is lower than that in the control group.

Level of mistrust Table 4 summarizes the subjects’ attitude towards their insurance companies. Among all subjects, 15.66% had a negative attitude towards their insurance company in terms of trust level (i.e., they responded Strongly disagree, Disagree, Somewhat disagree to the question concerning trust level). After reclassifying subjects’ trust level into two groups (i.e., Mistrust and Do Not Mistrust), it was found that the effect of recommendation on steering subjects’ choice to the lowest-cost provider (and lower-cost providers) was related to their mistrust level. Table 5 illustrates the association between recommendation and the count of subjects who choose the lowest-cost provider (and lower-cost providers) under each mistrust level. The and Fisher’s exact tests were conducted to examine whether there were differences in two dependent measures under recommendation versus no recommendation. The results showed that while there was no association between the differences in dependent measures for subjects who mistrusted their insurance companies, there was a significant correlation for those who did not mistrust their insurance company. This suggests that mistrust may be an important moderator of subjects’ choice behavior.

thumbnail
Table 4. Subjects’ response to the statement “I trust my insurance company”.

https://doi.org/10.1371/journal.pone.0350645.t004

thumbnail
Table 5. Heterogeneity of recommendations effect across subject subgroups with different mistrust levels in insurance companies.

https://doi.org/10.1371/journal.pone.0350645.t005

Understanding of insurance cost sharing In the study cohort, 80.92% of the subjects passed the quiz (see Table 6, part (1)). The analysis categorized those subjects who passed the quiz as “Yes” in the Table 6, part (2), and the rest as “No,” regardless of their treatment assignments. There was a strong correlation between subjects’ quiz performance and their provider choices (see Table 6, part (2)). Additionally, passing the quiz was strongly associated with choosing the lowest-cost provider and lower-cost providers (see Table 6, parts (3) and (4)). This suggested that subjects with a good understanding of the cost sharing were more likely to choose the lowest-cost provider and lower-cost providers.

thumbnail
Table 6. Description of quiz performance and its correlation with provider choice.

https://doi.org/10.1371/journal.pone.0350645.t006

Results of regression analysis

Preliminary results Regression results for Model (1) are summarized in Table 7. The results from an approach in which control variables are added sequentially are presented in S7 File. The reported coefficients represent changes in the log-odds of the outcomes (i.e., M1, M2) associated with a one-unit change in the treatment, which are not directly interpretable in probability terms. To facilitate interpretation, we report Average Marginal Effects (AMEs), which reflects the average change in the probability of subjects choosing the lowest-cost provider (or lower-cost providers) when the treatment changes from 0 to 1, holding other variables constant. The results confirm the earlier finding that recommendation has a positive impact on subjects choosing the lowest-cost provider and lower-cost providers, but copay waiver and persuasion do not have a significant impact on either measure. The AME indicates that the recommendation treatment increases the probability of subjects choosing the lowest-cost provider by 28.6% and the probability of selecting lower-cost providers by 23.5%.

thumbnail
Table 7. Regression results for the preliminary model.

https://doi.org/10.1371/journal.pone.0350645.t007

Comprehensive regression analysis results The results for the comprehensive model are presented in Table 8, displaying only the significant interactions. The complete results including all interactions can be found in S8 File.

thumbnail
Table 8. Regression Results for Comprehensive Model.

https://doi.org/10.1371/journal.pone.0350645.t008

The findings are as follows. First, recommendation has a positive impact on subjects’ choice of the lowest-cost provider. Looking at AMEs, subjects receiving recommendations are predicted to have 32.7% higher likelihood of selecting the lowest-cost provider and 27.3% higher chance of choosing lower-cost providers. Second, the recommendation effect decreases when subjects mistrust their insurance companies. When receiving recommendations, subjects who mistrust their insurance company are predicted to be 27.4% less likely to choose the lowest-cost one and 27.6% less likely to choose lower-cost providers. However, even with mistrust, recommendation maintains a net positive effect on choosing the lowest-cost provider, though not for selecting lower-cost providers (i.e., 32.7%−27.4% = 5.3% for M1; 27.3%−27.6% = −0.3% for M2). Third, copay and persuasion do not have a significant impact on the choice of either the lowest-cost providers or lower-cost providers. Lastly, whether subjects pass the quiz pertaining to the knowledge of health insurance cost-sharing has a significant impact on both the choice of the lowest-cost provider and lower-cost providers. Subjects who pass the quiz are predicted to be 18.8% more likely to choose the lowest-cost provider and 12.4% more likely to select lower-cost providers, holding all else constant.

While our models explain a modest portion of the total variance in the outcome (Pseudo R2 ranges from 8.29% to 14.85% in Tables 7 and 8), the primary focus of this study is on the consistent and theoretically-grounded relationship between the three treatments and two dependent measures. The low R-squared is common in studies involving human behavior and decision making, as a great deal of the variation is inherently idiosyncratic and unobservable. We provide further investigation on the subjects’ provider choices in the Discussion section and S13 File. The significant coefficients and AMEs nonetheless provide valuable insights into the directional effect and relative importance of these factors.

Robustness check results

First, among all subjects, 79.52% passed the Winograd check and the remaining 20.48% failed. Subjects were more likely to have a good understanding of the cost sharing if they passed the attention check (see S10 File) although the Winograd questions had nothing to do with provider choice whatsoever. The two additional analyses, mentioned earlier, produced results similar to the original analysis (see S10 File). That is, only recommendation significantly increased the likelihood of subjects selecting lowest- and lower-cost providers and the effect was undermined if subjects mistrusted the insurance company. Second, through the oversampling, it was found that the key results remain qualitatively consistent throughout (see S11 File). Again, recommendation was the only effective strategy in steering subjects’ choice to lowest- and lower-cost providers, with its effect moderated by the level of mistrust. Third, the counterfactual analysis showed that the results were consistent with observed claims data in the aggregate (see S12 File). Recall that we compared the estimated proportion of subjects choosing lower-cost providers with the actual proportion of beneficiaries choosing a lower-cost provider in the claims data. The predicted proportion of subjects choosing a lower-cost provider was 53.2%, closely matching the 53% observed in claims data. Note that this consistency should be interpreted as supportive, rather than conclusive, evidence for the validity of our findings, as the study remains subject to the behavioral limitations outlined in the Behavioral Limitations & External Validity of the Discussion section. Lastly, we fitted the probit version of both the preliminary and comprehensive regression models. All the results remained consistent (see S9 File).

Discussion

Summary of findings

Our experiment sheds light on the four research questions we highlighted in the Introduction section. A summary of these findings is presented below.

  1. In response to the question “what proportion of beneficiaries change their initial choice of provider?,” we find that without any intervention (i.e., Scenario 1 in Table 3), 31.03% subjects chose the lowest-cost provider and 37.93% chose one of the lower-cost providers.
  2. In response to the question “Which strategy is most effective in increasing the proportion of beneficiaries switching to lower-cost providers?,” we find that recommendation affects subjects’ choices the most. Subjects receiving recommendations are predicted to have 32.7% higher likelihood of selecting the lowest-cost provider and 27.3% higher chance of choosing lower-cost providers. Moreover, copay waiver and testimonials do not significantly affect beneficiaries’ choices.
  3. In response to the question “Which individual characteristics have the greatest impact on beneficiaries choices?,” we find that mistrust and comprehension of cost sharing concepts, i.e., copay and coinsurance, are the highest impact factors. Mistrust undermines the recommendation effect, whereas greater comprehension of cost sharing increases the likelihood that a subject will choose the lowest-cost provider and lower-cost providers.
  4. First, the findings suggest that recommendation may be more effective than copay waivers or testimonial persuasion in influencing provider choice. Second, the results highlight the potential importance of reducing mistrust and increasing beneficiaries’ comprehension of cost-sharing. Improving transparency around how provider choices affect both individual out-of-pocket expenses and broader insurance costs may further improve the effectiveness of such interventions.

We next utilize insights we gleaned from the literature to explain the theoretical underpinnings of our findings.

Possible mechanisms The impact of word-of-mouth (WOM) recommendations on consumer decision-making is well-documented, with effectiveness influenced by social ties and perceived risk [11,12]. Social ties refer to interactions between individuals [13], and tie strength reflects the closeness of these relationships [14]. Strong ties, such as family and close friends, are typically more influential than weak ties, which include strangers or casual acquaintances [11]. In this context, an agent’s recommendation resembles a WOM recommendation with weak social ties. Consumers often prefer weak-tie sources for technical products requiring specialized knowledge and with higher perceived risks [12,15]. Health care products are complex with higher perceived outcome risks and information asymmetry, which drives beneficiaries to seek recommendations from various sources to assess provider quality and avoid service failures [12,16]. An agent’s recommendation can serve as expert guidance, aiding their decision-making process. This explains the effectiveness of the recommendation treatment in our experiment. Meanwhile, people scrutinize source trustworthiness more closely in the recommendation seeking process when perceived risks are high [15]. This aligns with our finding that the credibility of the agent’s recommendation diminishes when subjects mistrust their insurance company.

Focusing next on the lack of effectiveness of copay waiver, we first postulate that the insurer’s expectation that a copay waiver may be effective is rooted in the reciprocity effect which refers to the tendency to respond to others’ intentions by rewarding kindness and punishing unkindness [1720]. That is, in our experiment the unconditional copay waiver is designed to be a gesture of goodwill from the insurer. However, this backfired. Drawn from the experiment’s results, we conjecture that subjects may interpret the waiver not as genuine kindness but as a strategic maneuver to direct them toward lower-quality providers. The potential suspicion of insurer’s intention could explain the lack of the reciprocity effect.

The testimonial effect is tied to persuasion, defined as “human communication designed to influence the judgments and actions of others” [21]. Unlike inducement or coercion, persuasion does not restrict options, or significantly alter economic incentives, or impose penalties. Its effectiveness depends on three key factors: persuasion knowledge (how individuals cope with persuasion), agent knowledge (beliefs about the persuader’s traits, tactics, and goals), and topic familiarity (prior experience with the subject) [22]. Similar to copay waiver, we conjecture that the suspicion of the insurer’s motives undermines the subjects’ perception of agents’ knowledge, which explains the lack of effectiveness of persuasion treatment in the experiment.

A limitation of our experiment in this context is that we did not test other testimonials. While sharing patient experiences in one-on-one beneficiary-agent interactions seem to be ineffective, other approaches remain unexplored and may produce positive outcomes.

A possible explanation for the association between a better understanding of the cost sharing and a greater propensity to choose the lowest-cost providers or lower-cost providers is that subjects who understand their costs are more likely to recognize the connection between their choices and the long-term cost of health services (see S13 File). Therefore, they are more inclined to choose a low-cost provider as long as the quality is not compromised.

Behavioral limitations & external validity

While the experimental design allows for controlled identification of treatment effects, subjects’ choices in this hypothetical setting may differ from real-world beneficiary behavior in several important ways.

First, the decision context is simplified. In practice, provider choice is influenced by additional factors such as physician referrals, appointment availability, provider capacity, and urgency of medical needs. These constraints are intentionally abstracted away in our design to isolate the causal impact of the three strategies, but may limit behavioral realism. In addition, while we limit the providers’ location to within 3 miles distance from a participant’ residence, individuals in real settings may prefer the most conveniently located provider.

Second, the experiment does not fully capture physician influence. While our design uses the requested provider to approximate potential physician recommendations, in practice such influence often takes the form of physician referrals, which exerts a stronger effect and could be a driver of provider selection.

Third, subjects’ responses may be subject to hypothetical bias [23]. Because choices are not tied to real consequences, participants may behave differently than they would in practice. For example, some individuals may select lower-cost providers to signal altruism and present themselves as “good Samaritans,” even if they would prioritize other factors such as convenience or appointment availability in real decision contexts.

While our counterfactual analysis suggests that the experimental patterns are broadly consistent with observed aggregate claims data (see S11 File), still, the results should be interpreted as providing suggestive evidence rather than directly generalizable estimates of real-world behavior because of the aforementioned limitations. Future research using field data or real-world interventions would be valuable to further assess their external validity.

Qualitative analysis of the underlying reasons for subjects’ choices

The survey also asked subjects to report their reasons for selecting a particular provider after making their final decision. Their responses highlighted the multitude of motivations that influence beneficiaries choice of lowest-cost and lower-cost providers. Details of the qualitative analysis of these self-reported reasons are provided in S13 File. We summarize the main observations here.

  • Subjects care about the value delivered by providers, and are willing to select lower-cost options among those with similar quality. Among the 191 subjects who selected Pear Clinic (lowest-cost provider) across eight scenarios, 59% mentioned it as the best-value provider.
  • A significant proportion of subjects exhibit altruism by selecting options that reduce costs for health insurance companies. For instance, 44% chose Pear Clinic to save costs for their insurer.
  • Many subjects tend to stick with their initial choice. Among the 215 subjects who selected Orange Clinic, the second-highest priced one and the requested provider, 63% cited that they preferred to stick to their initial choice.
  • A substantial number of subjects believe in a positive correlation between provider prices and quality, leading them to choose expensive options. For example, 5% of subjects who chose Orange Clinic believed it had higher quality. Meanwhile, among the 69 subjects selecting Peach Clinic, the highest-cost provider, 86% believed it had the best quality.

Other payment innovations

In addition to BVA-like health advocacy programs, insurers have utilized other approaches to help lower costs and improve quality [24,25]. For example, Reference Pricing (RP) and Rewards Program are two alternative approaches that steer beneficiary choices to low-cost providers. In the RP scheme, the insurer sets the maximum amount, called reference price, that can be reimbursed for a procedure. If a patient selects a provider that charges more than the reference price, then the patient is responsible for the portion in excess of the reference price in addition to the copayment or the coinsurance portion of the reference price, as applicable. The Rewards Program pays a monetary reward to beneficiaries who choose a low-cost provider [24]. Unlike the RP Scheme and the Rewards Program, beneficiaries’ choices do not directly influence their out-of-pocket costs under the BVA program. This characteristic retains the moral hazard inherent in health insurance services, making the program unique. Therefore, the insights do not directly apply to RP and Rewards Program. More details on the healthcare payment innovations are provided in S14 File.

Use of MTurk

MTurk was chosen for several reasons. First, an experimental approach is necessary to establish a causal relationship. Even if observational data were available, numerous confounding factors, such as variations in interactions between agents and beneficiaries, as well as the tone and manner of communication would be difficult to control for. Our text-based experimental design allows for control over these variables. Furthermore, due to ethical concerns, IRB typically prohibit such experiments on real patients, leading researchers to conduct similar studies in behavioral labs with university students. However, college students are typically under 25 and lack the demographic and socioeconomic diversity that is necessary for this study. Lastly, although generative AI could be used to mimic human-like responses to questions pertaining to preference, such experiments do not replicate real participants’ responses [26]. Therefore, directly using large language models in place of human respondents for preference elicitation may be misleading [27]. For these reasons we chose MTurk, which has been widely used in health services, psychology, political science, and marketing research, consistently providing representative and reliable samples [2833]. Additional justification for the choice of MTurk for conducting the experiment is presented in S15 File.

Other limitations

Besides limitations in external validity, our study has several other limitations. First, in terms of the experiment format, we only utilized text-based scenarios to avoid the confounding factors inherent in other formats (e.g., accent and tone in audio format, and identification with the narrator’s race and skin color in the video format). Additionally, we did not test all forms of testimonial and did not vary the copayment amount. These would be the avenues for future research. Lastly, the sample obtained from the MTurk is younger than national population. Future studies may include more elderly subjects to improve the generalizability.

Supporting information

S1 Fig. (1) Requested versus Recommended amount. (2) Percentage of times providers belonged to different price quartile by requested versus recommended. The price quartile is determined based on requested amount.

https://doi.org/10.1371/journal.pone.0350645.s022

(TIF)

S2 Fig. An example of MRI Scan.

Computer-generated image by the authors using Claude. Note that the figure is similar but not identical to the image we used in the survey and is therefore for illustrative purposes only.

https://doi.org/10.1371/journal.pone.0350645.s001

(TIF)

S4 Fig. Provider information with recommendation.

https://doi.org/10.1371/journal.pone.0350645.s003

(TIF)

S5 Fig. Provider information with copay waiver.

https://doi.org/10.1371/journal.pone.0350645.s004

(TIF)

S6 Fig. Provider information with copay waiver & recommendation.

https://doi.org/10.1371/journal.pone.0350645.s005

(TIF)

S7 Fig. Top reasons for choosing pear & orange clinic.

https://doi.org/10.1371/journal.pone.0350645.s006

(TIF)

S1 File. S1 Appendix.

Similar health advocate programs.

https://doi.org/10.1371/journal.pone.0350645.s007

(PDF)

S2 File. S2 Appendix.

Preliminary Analysis of BCBS Claims Data.

https://doi.org/10.1371/journal.pone.0350645.s008

(PDF)

S3 File. S3 Appendix.

Payment structure in the BVA program.

https://doi.org/10.1371/journal.pone.0350645.s009

(PDF)

S4 File. S4 Appendix.

Price negotiations for non-urgent and shoppable procedures & misperception of price-quality relationship.

https://doi.org/10.1371/journal.pone.0350645.s010

(PDF)

S5 File. S5 Appendix.

The survey instrument.

https://doi.org/10.1371/journal.pone.0350645.s011

(PDF)

S6 File. S6 Appendix.

Sample size selection & mturk experiment completion rate.

https://doi.org/10.1371/journal.pone.0350645.s012

(PDF)

S7 File. S7 Appendix.

Results of Preliminary model – logit regression.

https://doi.org/10.1371/journal.pone.0350645.s013

(PDF)

S8 File. S8 Appendix.

Results of comprehensive model – logit regression.

https://doi.org/10.1371/journal.pone.0350645.s014

(PDF)

S9 File. S9 Appendix.

Results of preliminary and comprehensive models – probit regression.

https://doi.org/10.1371/journal.pone.0350645.s015

(PDF)

S10 File. S10 Appendix.

Additional analysis involving winograd Check.

https://doi.org/10.1371/journal.pone.0350645.s016

(PDF)

S11 File. S11 Appendix.

Oversampling results.

https://doi.org/10.1371/journal.pone.0350645.s017

(PDF)

S13 File. S13 Appendix.

Qualitative analysis of the underlying reasons for subjects’ choices.

https://doi.org/10.1371/journal.pone.0350645.s019

(PDF)

S14 File. S14 Appendix.

Healthcare payment innovations.

https://doi.org/10.1371/journal.pone.0350645.s020

(PDF)

S15 File. S15 Appendix.

Use of mechanical turk.

https://doi.org/10.1371/journal.pone.0350645.s021

(PDF)

References

  1. 1. Robinson JC. Variation in hospital costs, payments, and profitabilty for cardiac valve replacement surgery. Health Serv Res. 2011;46(6pt1):1928–45. pmid:21762141
  2. 2. Baker LC, Bundorf MK, Royalty AB, Levin Z. Physician practice competition and prices paid by private insurers for office visits. JAMA. 2014;312(16):1653–62. pmid:25335147
  3. 3. Procedure Price Data com. MRI Cervical Spine - Austin, TX. Available from: https://procedurepricedata.com/procedures/mri-spine-cervical/austin?page=1. Accessed 2026 April 20.
  4. 4. BCBS of Texas. Benefits value advisors: helping you maximize your benefit plan. Available from: https://connect.bcbstx.com/understanding-benefits/b/weblog/posts/benefits-value-advisor-helping-navigate-health-care. 2025. Accessed 2026 April 20.
  5. 5. Hussey PS, Wertheimer S, Mehrotra A. The association between health care quality and cost: a systematic review. Ann Intern Med. 2013;158(1):27–34. pmid:23277898
  6. 6. Whaley C. Searching for health: the effects of online price transparency. 2015.
  7. 7. Roddy J, Robinson S. An exploration of stress: leveraging online data from crowdsourcing platforms. Front Artif Intell. 2021;4:591529. pmid:33733231
  8. 8. Eysenbach G. Improving the quality of web surveys: the Checklist for Reporting Results of Internet E-Surveys (CHERRIES). J Med Internet Res. 2004;6(3):e34. pmid:15471760
  9. 9. Levesque HJ, Davis E, Morgenstern L. The Winograd schema challenge. In: Proceedings of thirteenth international conference on the principles of knowledge representation and reasoning, 2012. Available from: https://cdn.aaai.org/ocs/4492/4492-21843-1-PB.pdf
  10. 10. Difallah D, Filatova E, Ipeirotis P. Demographics and dynamics of mechanical turk workers. In: Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining; 2018. pp. 135–43. http://dx.doi.org/10.1145/3159652.3159661
  11. 11. Brown JJ, Reingen PH. Social ties and word-of-mouth referral behavior. J CONSUM RES. 1987;14(3):350.
  12. 12. Duhan DF, Johnson SD, Wilcox JB, Harrell GD. Influences on consumer use of word-of-mouth recommendation sources. J Acad Mark Sci. 1997;25(4):283–95.
  13. 13. Wang J-C, Chang C-H. How online social ties and product-related risks influence purchase intentions: a facebook experiment. Electron Commer Res Appl. 2013;12(5):337–46.
  14. 14. Granovetter MS. The strength of weak ties. Am J Sociol. 1973;78(6):1360–80.
  15. 15. Dowling GR, Staelin R. A model of perceived risk and intended risk-handling activity. J Consum Res. 1994;21(1):119.
  16. 16. Akçura MT, Ozdemir ZD. A strategic analysis of multi-channel expert services. J Manag Inf Syst. 2017;34(1):206–31.
  17. 17. Falk A, Fischbacher U. A theory of reciprocity. Games Econ Behav. 2006;54(2):293–315.
  18. 18. Fehr E, Kirchsteiger G, Riedl A. Does fairness prevent market clearing? An experimental investigation. Q J Econ. 1993;108(2):437–59.
  19. 19. Fehr E, Gächter S. Fairness and retaliation: the economics of reciprocity. J Econ Perspect. 2000;14(3):159–82.
  20. 20. Rabin M. Incorporating fairness into game theory and economics. Am Econ Rev. 1993;:1281–302.
  21. 21. Simons HW, Jones J. Persuasion in society. 3rd ed. Taylor & Francis; 2011.
  22. 22. Friestad M, Wright P. The persuasion knowledge model: how people cope with persuasion attempts. J Consum Res. 1994;21(1):1–31.
  23. 23. Hensher DA. Hypothetical bias, choice experiments and willingness to pay. Transp Res Part B Methodol. 2010;44(6):735–52.
  24. 24. Whaley CM, Vu L, Sood N, Chernew ME, Metcalfe L, Mehrotra A. Paying patients to switch: impact of a rewards program on choice of providers, prices, and utilization. Health Aff (Millwood). 2019;38(3):440–7. pmid:30830823
  25. 25. Gupta D, Mehrotra M, Tang X. Gainsharing contracts for CMS’ Episode‐based payment models. Prod Oper Manag. 2021;30(5):1290–312.
  26. 26. Stanford Institute for Human-Centered Artificial Intelligence. Simulating human behavior with AI agents. Available from: https://hai.stanford.edu/policy/simulating-human-behavior-with-ai-agents. 2023. Accessed 2025 November 10.
  27. 27. Goli A, Singh A. Can large language models capture human preferences? Mark Sci. 2024;43(4):709–22.
  28. 28. Yeh VM, Schnur JB, Margolies L, Montgomery GH. Dense breast tissue notification: impact on women’s perceived risk, anxiety, and intentions for future breast cancer screening. J Am Coll Radiol. 2015;12(3):261–6. pmid:25556313
  29. 29. Bardos J, Hercz D, Friedenthal J, Missmer SA, Williams Z. A national survey on public perceptions of miscarriage. Obstet Gynecol. 2015;125(6):1313–20. pmid:26000502
  30. 30. Bardos J, Friedenthal J, Spiegelman J, Williams Z. Cloud based surveys to assess patient perceptions of health care: 1000 respondents in 3 days for US $300. JMIR Res Protoc. 2016;5(3):e166. pmid:27554915
  31. 31. Crump MJC, McDonnell JV, Gureckis TM. Evaluating Amazon’s mechanical turk as a tool for experimental behavioral research. PLoS One. 2013;8(3):e57410. pmid:23516406
  32. 32. Liu N, Finkelstein SR, Kruk ME, Rosenthal D. When waiting to see a doctor is less irritating: understanding patient preferences and choice behavior in appointment scheduling. Manage Sci. 2018;64(5):1975–96.
  33. 33. Kim S-H, Tong J, Peden C. Admission control biases in hospital unit capacity management: how occupancy information hurdles and decision noise impact utilization. Manage Sci. 2020;66(11):5151–70.