Skip to main content
Advertisement
Browse Subject Areas
?

Click through the PLOS taxonomy to find articles in your field.

For more information about PLOS Subject Areas, click here.

  • Loading metrics

Examining the joint impact of missing data mechanisms and item parameter drift on the accuracy of item response theory-based test equating

Abstract

Maintaining score comparability across different test administrations is essential in large-scale educational and psychological assessment. Two major threats to equating accuracy are item parameter drift (IPD), reflecting changes in anchor item characteristics over time, and missing item responses, which commonly occur in operational testing. Although each factor has been studied separately, their joint impact on equating accuracy has not been systematically evaluated. A Monte Carlo simulation was conducted using a nonequivalent groups with anchor test design. Data were generated under a three-parameter logistic model with 2,000 examinees per form and 500 replications per condition. The Stocking–Lord method was used to estimate equating constants across 25 conditions, varying IPD rate (0%, 10%, 20%), IPD magnitude (0, 0.25, 0.50), missing data mechanism (MCAR vs. MAR), and missing data rate (0%, 10%, 20%). Missing responses were addressed using multiple imputation by chained equations. Equating accuracy was evaluated using bias and root mean square error. Results indicated that the B constant was highly sensitive to the missing data mechanism: MCAR conditions produced near-zero bias, whereas MAR conditions introduced substantial positive bias, particularly at higher missing rates. IPD alone did not result in meaningful bias but increased estimation variability. When IPD and MAR co-occurred, equating error exceeded the sum of their individual effects, indicating an interaction rather than a purely additive relationship. Although multiple imputation reduced error under MCAR, it did not fully eliminate bias under MAR conditions. In contrast, the scale constant remained stable across all conditions. Overall, the missing data mechanism had a stronger impact on equating accuracy than either the rate of missingness or the degree of item drift alone. When both factors were present, their combined influence led to increased error that could not be fully corrected through imputation. These findings underscore the importance of carefully evaluating missing data mechanisms, selecting appropriate imputation strategies, such as MICE, and monitoring anchor item stability to ensure accurate score equating, particularly in psychological assessment contexts where even small errors may affect individual-level decisions.

Introduction

In educational and psychological measurement, the use of multiple test forms that assess the same construct is a common practice, particularly to prevent item exposure, ensure test security, and administer large-scale tests. When different forms are administered to different groups of examinees, it is necessary to place the resulting scores on a common scale to enable meaningful comparisons. In psychological assessment, this issue of comparability goes beyond a simple technical concern. When the same construct is measured across different settings or groups, a lack of score consistency can seriously weaken the validity of any clinical or diagnostic interpretations based on those results. This process is referred to in the literature as test equating [1] and can also be understood as expressing examinees’ ability parameters on the same latent trait metric [2]. Through equating, it is possible to monitor individuals’ performance across administrations and to compare item parameter estimates obtained from different samples. In psychological assessment contexts, such comparability is essential, as test scores are often used to inform decisions about individuals’ cognitive, emotional, or behavioral functioning.

Item Response Theory (IRT) based equating approaches are widely used in measurement practice because they model item and person parameters separately and are relatively less sensitive to sampling fluctuations [35]. This is particularly important in psychological measurement, where ensuring comparability across administrations supports valid inferences about latent traits. One of the most employed designs in this context is the nonequivalent groups with anchor test (NEAT) design [1], in which a set of common items is used to link a new form to a base form and place both on the same scale [6]. This approach is supported by evidence showing that characteristic-curve methods based on separate calibration yield more stable results than moment methods under the Common-Item Equating Design [7].

In IRT-based equating, anchor items play a central role in establishing the linking, and the accuracy of the linking constants largely depends on the stability of their item parameters. Accordingly, the validity of the equating results rests on the assumption that anchor items function equivalently across forms. Item parameter drift (IPD) refers to the phenomenon in which the difficulty or discrimination parameters of anchor items shift across administrations [8,9], thereby posing one of the most fundamental threats to the assumption of parameter invariance. The primary causes of IPD include curriculum changes, item exposure, and cultural, educational, or technological developments that may render anchor items easier or harder for examinees over time [10,11]. When item parameter drift occurs, scale transformations may become biased, which can in turn affect ability estimates and compromise the comparability of scores across test forms [12]. Among the various forms of IPD, drift in the difficulty (b) parameter represents the most encountered and practically consequential form in large-scale assessments [13,14]. Although drift in the discrimination (a) and pseudo-guessing (c) parameters also occurs in practice, b-parameter drift appears to be more prevalent in operational testing contexts and has been shown to have a more direct impact on equating accuracy. Drift in a parameter tends to be less frequent and has comparatively more localized effects on scale transformation, while c-parameter drift is particularly difficult to detect reliably due to the inherently lower precision of its estimates [13,14]. For these reasons, the present study focuses specifically on b-parameter drift, reflecting conditions most likely to arise in operational testing contexts. In psychological testing contexts, such distortions may compromise the accuracy of individual-level interpretations, especially when test scores are used for diagnostic or evaluative purposes.

Another major challenge in equating studies is missing data. Missing responses are ubiquitous in social and educational research and directly affect the validity and reliability of statistical analyses. Examinees may omit items for various reasons, such as difficulty in choosing among response options, lack of motivation to respond, concerns about data security, or time limitations that prevent them from reaching all items [15]. The impact of missing data is determined not only by the proportion of missingness but also by its pattern and underlying mechanism [16]. As the amount of missing data increases, the representativeness of the sample decreases, generalizability is compromised, and statistical bias tends to increase [17]. In psychological research and assessment, such biases may disproportionately affect certain groups of individuals, thereby raising concerns about fairness and validity. Following Rubin [18], missing data mechanisms are categorized as missing completely at random (MCAR), missing at random (MAR), and missing not at random (MNAR). Under MCAR, missingness is independent of both observed and unobserved variables; under MAR, it depends only on observed variables such as ability level [19]. In this study, only MCAR and MAR were simulated, since MNAR introduces complex methodological challenges in large-scale test equating [20,21]. MCAR represents completely random missingness, whereas MAR reflects missingness dependent on observed variables, consistent with patterns observed in empirical data [15].

To address missing data, a variety of traditional and modern approaches have been proposed. Traditional methods such as listwise and pairwise deletion may lead to biased estimates and inflated standard errors, particularly when the proportion of missing data is high [22]. Among modern techniques, multiple imputation (MI) is a flexible and efficient approach that involves generating multiple plausible versions of the incomplete dataset and combining results from these completed datasets. Allison [23] noted that MI is applicable across a wide range of statistical models, making it one of the most practical imputation methods. In recent years, MI has been increasingly recommended in the IRT context to reduce bias due to missing data and to provide more realistic representations of estimation uncertainty [24]. In the present study, MICE was applied to generate multiple plausible datasets and to combine estimates using Rubin’s rules [23,25]. The number of imputations for each condition is reported in the Method section, where the rationale for the selected imputation strategy is described in detail.

A review of the literature indicates that, although test equating has been examined in relation to missing data [2628] and item parameter drift [2933] separately, to our knowledge, no study has yet examined their interactive effects within a single controlled framework, a gap that has direct implications for operational equating practice.

Given that missing responses are unavoidable in operational testing programs, understanding how missing data and anchor item parameter drift jointly affect equating accuracy is of both theoretical and practical importance. In particular, it remains unclear how different missing data mechanisms and rates influence the standard errors of the equating constants in the presence of parameter drift, to what extent MI can reduce this uncertainty, and how the performance of the commonly used equating method is affected under these non-ideal conditions.

The current simulation study was designed to systematically manipulate missing data mechanism (MCAR, MAR), missing data rate (0%, 10%, 20%), IPD rate (0%, 10%, 20%), and IPD magnitude (0, 0.25, 0.50), allowing for a controlled examination of equating accuracy under realistic operational scenarios using the Stocking–Lord (SL) method and IRT true score equating. The SL method was selected over competing linking methods, notably the Mean/Sigma and Haebara approaches, due to its demonstrated robustness in the presence of IPD. Unlike the Mean/Sigma method, which relies on the mean and standard deviation of item parameter estimates and may therefore be more sensitive to parameter drift, the SL method minimizes discrepancies across the full test characteristic curve, thereby reducing sensitivity to localized parameter instability. The Haebara method similarly operates at the item level but has, in some simulation studies, shown sensitivity to outlying item parameters [29]. Previous studies have also suggested that the SL method can exhibit relatively greater stability under conditions such as small sample sizes and noisy data [34,35]. In combination with IRT true score equating, the SL method facilitates the transformation of item parameters and ability estimates across test forms. This is particularly important when the 3PL model is used, as non-zero guessing parameters (c) require accurate scale alignment.

In sum, the study integrates NEAT design, multiple imputation for missing data, and the SL scaling approach to investigate the joint effects of IPD and missing responses on equating constants (A and B), where A represents the scale transformation constant, and B represents the location transformation constant used to place the new form on the base-form metric [1].

To the authors’ knowledge, research examining the joint effects of item parameter drift and missing data within a unified simulation framework in IRT-based equating remains limited. The findings are intended to provide evidence-based guidance to practitioners and researchers who must make methodological decisions regarding anchor item management and missing data handling in operational testing programs.

Purpose of the study

The present study aims to examine the joint effects of item parameter drift and missing data on the accuracy of IRT true score equating within a unified simulation framework. Specifically, the study investigates how varying levels of anchor item drift (differing in both proportion and magnitude) interact with different missing data mechanisms and rates to influence the bias and precision of SL equating constants. A secondary aim is to evaluate the extent to which multiple imputation mitigates equating error under these non-ideal conditions. By systematically manipulating these factors, the study seeks to provide evidence-based guidance for methodological decisions in operational equating practice. Beyond its methodological contribution, the study aims to inform best practices in psychological assessment by identifying conditions under which equating procedures may produce biased or unstable results.

Based on this purpose, the study addresses the following research questions:

  1. How does anchor item parameter drift, in terms of both the proportion of drifting items and drift magnitude, affect the bias and Root Mean Square Error (RMSE) of SL equating constants?
  2. How do the missing data mechanisms (MCAR vs. MAR) and the missing data rate influence the bias and RMSE of equating constants?
  3. What are the joint effects of item parameter drift and missing data on equating accuracy, and do these combined effects exceed those observed when each source of error is present in isolation?
  4. To what extent does MI reduce bias and improve the precision of equating constants under missing data conditions, and does this effect vary in the presence of item parameter drift?

Method

Research design

This study employed a simulation-based research design as described by Dooley [36]. Unlike empirical studies, which are constrained by observed data, simulation approaches allow researchers to systematically manipulate the parameters and conditions of interest and examine their behavior under controlled scenarios. In this context, a Monte Carlo simulation framework was used to investigate the effects of common item parameter drift and missing data conditions on IRT-based test equating results.

Ethics statement. Ethics approval was not required because this study used simulated data and did not involve human participants.

Data generation

Data sets were generated in R (version 4.5.1) to evaluate the performance of the test equating method under varying missing-data mechanisms and IPD conditions. A Monte Carlo simulation approach was adopted, with 500 replications following common practice in the simulation literature [37]. A sample of 2,000 examinees was simulated for each form in each replication. The data generation code was reviewed and test-run prior to the main simulation to confirm that the generated parameters were consistent with the specified distributions. Data generation was based on the three-parameter logistic (3PL) model within the IRT framework.

Item parameters

Item difficulty parameters (b) were generated from a uniform distribution U(−2, 2), a range commonly used in simulation studies to represent tests of moderate difficulty [38]. Item discrimination parameters (a) were generated from a log-normal distribution, log(a) ~ N(0, 0.5), where 0.5 denotes the standard deviation of the log-normal distribution, with the additional constraint that no item had a discrimination value below 0.4 [5]. The use of a log-normal distribution ensures that discrimination parameters remain positive and reflect the asymmetric distributions typically observed in real test items [38].

The guessing parameter (c) was fixed at 0.20 for all items, consistent with chance performance on five-option multiple-choice items. In the 3PL model, the pseudo-guessing parameter is often treated as a fixed item-level constant, and fixing it is common due to instability in its estimation [39]. It should be acknowledged that fixing c at a single value across all items is a simplifying assumption; in practice, pseudo-guessing parameters vary across items and are rarely known with certainty. Additional design assumptions, including the use of a fixed pseudo-guessing parameter and unidirectional drift, may limit the generalizability of the findings to more complex operational testing conditions.

Ability distribution

Examinee ability parameters (θ) were generated from a standard normal distribution, N(0, 1) [40]. Sass et al. [41] noted that a normally distributed ability parameter is commonly assumed for stable estimation in IRT models. This assumption is also consistent with many previous simulation studies [14,30]. Ability parameters were estimated using the Expected a Posteriori (EAP) method, which has been shown to provide more accurate estimates compared to maximum likelihood estimation in unidimensional IRT frameworks [41,42].

Test design

The test length for both the base and new forms was set to 40 items, consistent with conditions commonly used in IRT simulation studies [14,39]. This length also reflects the typical number of items found in large-scale standardized assessments, such as certification and licensure examinations. An additional reason for selecting this length is that it conveniently meets the minimum anchor-item proportion required in a NEAT design.

Each form included 8 internal anchor items (20%). Kolen and Brennan [1] recommend that, in a NEAT design with 40 items, at least 20% of the test items be anchor items.

Within the scope of the study, the proportion of drifting anchor items was set at 10% and 20%. Accordingly, when the data were manipulated, drift was introduced to the 1 and 2 anchor items, respectively. These numbers are consistent with the number of drifting items observed in operational testing settings [13].

Experimental conditions

IPD conditions.

IPD conditions were defined based on factors expected to influence equating error. Three main dimensions were considered: (1) the proportion of drifting items, (2) the type, direction, and magnitude of drift, and (3) the parameter in which the drift occurs.

In this study, the proportion of drifting items in the anchor set was manipulated at three levels: 0% (control), 10%, and 20%. This range is consistent with both simulation studies and empirical findings on the number of misbehaving common items in operational testing settings [13,14].

IPD was applied only to the difficulty (b) parameter, consistent with previous simulation studies that focused on b-parameter drift in isolation [30,43]. This approach allows for a cleaner examination of the effect of difficulty drift on equating results.

IPD was simulated as unidirectional (positive), with all drifting anchor items having their b values increased in the same direction. Wells et al. [14] stated that unidirectional drift has a greater impact on ability estimation compared to bidirectional drift. In bidirectional cases, positive and negative deviations may partially cancel each other out, potentially masking the true effect of IPD. The decision to apply drift exclusively in the positive direction was made to facilitate the detectability of equating error and to isolate the effect of drift magnitude. Accordingly, the findings may be most applicable to testing conditions in which drift occurs predominantly in a single direction.

The magnitude of IPD was set at b = 0.0, 0.25, and 0.50, representing no drift, small drift, and moderate drift, respectively [14,30].

Missing data conditions

In this study, MCAR and MAR mechanisms were simulated as missing data conditions. Unlike MNAR, in which missing observations follow a systematic pattern associated with the latent levels of the measured variable, MCAR and MAR are considered more tractable mechanisms that allow for valid inference when appropriately handled [19,44]. MNAR was excluded due to the methodological complexity associated with modeling and interpreting it in large-scale test equating contexts [20,21].

Under the MAR condition, examinees were first divided into three groups based on their latent ability (θ) (low, medium, and high) using tertile splits. The probability of missing responses was then assigned differentially across these groups. For the 10% missing rate condition, the probabilities were set at 0.15, 0.10, and 0.05 for the low, medium, and high ability groups, respectively. For the 20% missing rate condition, the corresponding probabilities were 0.30, 0.20, and 0.10. Although MAR is formally defined with respect to observed variables, latent ability (θ) was used in this simulation as a proxy for observed performance indicators, such as total score or prior achievement, which are typically highly correlated with ability in operational testing contexts [15,21]. This approach was intended to approximate realistic omission behavior, in which lower-performing examinees are more likely to omit items. Accordingly, the present implementation is best interpreted as a near-MAR condition rather than a strictly observed-variable MAR process.

The missing data rate was set at 0%, 10%, and 20%. These rates are consistent with typical missing data patterns observed in large-scale assessments such as Programme for International Student Assessment (PISA), where missing item response rates have been reported to average approximately 8% and reach up to 19% across countries [15,45], and were further supported by simulation studies examining similar conditions in IRT and equating contexts [21,26,46].

Simulation design structure

The simulation design in this study included the factors of IPD rate (0%, 10%, 20%), IPD magnitude (0, 0.25, 0.50), missing data mechanism (MCAR, MAR), and missing data rate (0%, 10%, 20%).

However, the design was not fully crossed. Certain combinations were excluded based on logical constraints. For example, when the missing data rate was 0%, the missing data mechanism (MCAR/MAR) was not applicable. Similarly, when the IPD rate was 0%, the IPD magnitude factor was not implemented. Therefore, instead of generating all possible combinations, only theoretically meaningful and interpretable conditions were included. In total, 25 experimental data sets were created. Detailed information on these conditions is provided in Table 1.

thumbnail
Table 1. Simulation conditions (N = 25 conditions × 500 replications).

https://doi.org/10.1371/journal.pone.0353665.t001

Data analysis

Parameter estimation.

Item parameters for both the base and new forms were estimated using the mirt package in R [47]. A unidimensional 3PL model was specified in all conditions. To ensure consistency across conditions and avoid estimation instability, the guessing parameter (c) was fixed at 0.20 in all analyses.

Examinee ability parameters (θ) were estimated using the Expected A Posteriori (EAP) method. EAP is one of the most widely used methods for ability estimation in unidimensional IRT [42] and has been shown to yield more accurate estimates than maximum likelihood (ML) methods [41].

Missing data handling

Missing data were handled using Multiple Imputation by Chained Equations (MICE, [25]), implemented via the mice package in R [48]. Given the binary (0/1) nature of item response data, logistic regression was used as the imputation model for each item within the MICE framework. This choice is consistent with recommendations for imputing dichotomous variables [25].

The number of imputations (M) was set equal to the percentage of missing data (e.g., M = 10 for 10% missingness). This approach is consistent with the rule M ≥ percentage of missing data suggested by Graham et al. [49] and White et al. [50]. Following multiple imputation, item parameters were estimated separately for each imputed dataset using the mirt package.

Equating procedure

Equating was conducted using the SL method within a NEAT design. Analyses were performed using the equateIRT package [51]. Within the SL framework, the equating constants A and B are estimated by minimizing the weighted sum of squared differences between the test characteristic curves of the base and new forms. Item parameters from the new form are then transformed onto the base-form metric. Under the 3PL model, this transformation rescales the a and b parameters while maintaining consistency in the interpretation of the c parameter across forms [1,52].

In each replication, the new form was equated to the base form. Common items (Q1–Q8) were automatically identified using the item name matching feature of the equateIRT package. In the presence of missing data, pooled parameter estimates were used.

Evaluation metrics

The estimation of the A and B transformation constants obtained from the SL method was evaluated using Bias and RMSE. These metrics are widely used as standard performance indicators in IRT equating simulation studies [53]. Bias and RMSE were calculated as follows:

where λ̂_r is the estimated parameter in replication r, λ is the true parameter value, and R is the total number of replications.

In addition to absolute Bias and RMSE, Relative Bias (RB) and Relative RMSE (RRMSE) were also computed to assess the magnitude of deviation relative to the true parameter values [37,54]:

In calculating relative metrics, conditions with |λ| < 0.05 were excluded to avoid division-by-zero problems; for these cases, only absolute Bias and RMSE were reported. The threshold of |λ| < 0.05 was adopted as a practical cutoff to prevent numerically unstable relative metrics when the true parameter value approaches zero; this decision is consistent with common practice in equating simulation studies [54]. The criterion |RB| > 5%, proposed by Hoogland and Boomsma [55], was used as a threshold for identifying practically significant bias.

Results

This section presents the bias and RMSE values for the A and B equating constants obtained using the SL method across four research questions. The values for all conditions are reported in Table 2, and the differences relative to the reference condition (K1) are presented in Table 3. Relative bias (RB_B) is reported in Table 2 only for conditions in which the true B constant meets the |B_true| ≥ 0.05 threshold; conditions falling below this threshold are marked with a dash (—) to avoid numerically unreliable results, and for these conditions only absolute bias and RMSE are interpreted.

thumbnail
Table 2. Bias, RMSE, and relative bias of the SL equating constants across 25 simulation conditions.

https://doi.org/10.1371/journal.pone.0353665.t002

thumbnail
Table 3. Differences in bias and RMSE relative to the reference condition (K1) across 25 simulation conditions.

https://doi.org/10.1371/journal.pone.0353665.t003

Effect of item parameter drift on bias and RMSE of equating constants

The first research question examines how the rate and magnitude of anchor-item parameter drift affect the accuracy of the A and B constants in the absence of missing data. For this purpose, conditions K1 (control), K2, K3, K4, and K5 were analyzed.

For the A constant (see Table 2), bias values ranged between −0.0043 and 0.0013, and no meaningful bias was observed in any condition. RMSE of constant A increased from 0.0512 in K1 to 0.0634 in K5; however, this increase remained limited. Relative bias values for A did not exceed the 5% threshold in any condition.

A similar pattern was observed for the B constant. Bias values were quite small (−0.0011 to 0.0060), and no statistically significant bias was detected. However, RMSE increased as the magnitude of IPD increased. The value, which was 0.0493 in K1, reached 0.0580 in K5 (IPD rate = 20%, Δb = 0.50). When the delta values in Table 3 are examined, ΔRMSE_B in K5 is 0.0087, corresponding to an increase of approximately 18% compared to K1. Relative bias (RB_B) could only be computed in conditions where the true value of B was meaningful; in those cases, values ranged from −2.25 to −4.78 and did not exceed the 5% threshold. This pattern is also visible in Fig 1. When the missing data rate is 0%, the starting points of all IPD conditions are very close, but the lines begin to diverge as the magnitude of drift increases.

thumbnail
Fig 1. RMSE of the B Constant Across Missing Data Rates and IPD Conditions.

RMSE values for the B equating constant under MCAR and MAR conditions across varying levels of missing data and IPD.

https://doi.org/10.1371/journal.pone.0353665.g001

Taken together, these findings indicate that item parameter drift alone, at least at the levels examined in this study, does not introduce substantial bias in the A and B constants, but it does increase estimation variability, particularly under moderate drift conditions (Δb = 0.50).

Effect of missing data mechanism and rate on bias and RMSE of equating constants

The second research question focuses on the effects of the missing-data mechanism and rate in the absence of IPD. For this purpose, conditions K6, K7, K8, and K9 were compared with K1. The A constant remained largely stable in the presence of missing data. Bias values ranged from −0.0125 to 0.0018, and RMSE increased only slightly from 0.0512 in K1 to a maximum of 0.0550 in K7 (MCAR 20%) (see Table 2). No clear difference was observed between MCAR and MAR conditions in terms of the A constant.

For the B constant, however, the two mechanisms show a clear divergence. Under MCAR conditions (K6–K7), bias values (0.0027 and 0.0039) remained very close to that of K1 (−0.0007), and RMSE values stayed within a narrow range (0.0490–0.0491). Under MAR conditions, the pattern changes; bias increased to 0.0094 in K8 (MAR 10%) and to 0.0251 in K9 (MAR 20%). The delta values in Table 3 further highlight this difference, with ΔBias = 0.0258 and ΔRMSE = 0.0073 in K9. Since there is no true location difference between forms in conditions without IPD, the true population value of B is zero; therefore, relative bias could not be computed due to division by zero, and interpretations were based on absolute Bias_B values. Fig 1 also visually supports this distinction. In the MCAR panel, the line for the no-IPD condition remains nearly flat as the missing rate increases, whereas in the MAR panel, a clear upward trend is observed.

A similar pattern can be seen in Fig 2, where K7 (MCAR 20%) remains close to the reference line, while K8 and K9 (MAR) show a noticeable increase.

thumbnail
Fig 2. Bias in the B Constant Across Selected MCAR and MAR Conditions.

Bias values of the B equating constant under selected missing data and IPD conditions across MCAR and MAR mechanisms.

https://doi.org/10.1371/journal.pone.0353665.g002

These findings suggest that the MCAR mechanism does not systematically affect the B constant, whereas the MAR mechanism, especially at a 20% missing rate, introduces a notable positive bias. This pattern likely reflects the systematic loss of lower-ability individuals under MAR conditions.

Combined effects of IPD and missing data

The third research question examines whether the combined effect of IPD and missing data exceeds the effect of each factor alone. For this purpose, conditions K10–K25 were analyzed.

For the A constant, combined conditions led to only a limited increase in RMSE, mainly when the missing rate was higher. The highest RMSE value (0.0661) was observed in K21 (IPD 20%, Δb = 0.50 × MCAR 20%), corresponding to a ΔRMSE of 0.0148 relative to K1. Bias did not show any meaningful variation across combined conditions.

For the B constant, the pattern is more pronounced. Under MCAR conditions (K10–K13, K18–K21), bias remained relatively low; however, RMSE increased as both IPD magnitude and missing rate increased. In K21, RMSE was 0.0613, with a ΔRMSE_B of 0.0120. According to the relative bias values of B, K20, and K21, which exceeded the threshold for meaningful bias (see Table 2; −5.92* and −9.72*, respectively).

Under MAR conditions (K14–K17, K22–K25), the pattern becomes more striking. Bias of the B constant reached 0.0274 in K17 and 0.0289 in K25. These conditions also produced the highest RMSE values; in K25, RMSE was 0.0647, with a ΔRMSE of 0.0154. Most of the conditions marked with an asterisk in Table 3 involve the combination of moderate missing rates and moderate IPD, suggesting that the joint effect of these factors is not merely additive but reflects an interaction. This is also evident in Fig 1. When the two panels are compared, the distance between lines is clearly larger under MAR than under MCAR conditions. Fig 2 shows a similar pattern, with MAR conditions (K17, K25) displaying notably higher bias than MCAR conditions (K13, K21).

Role of multiple imputation in reducing equating error

The fourth research question investigates the extent to which multiple imputation reduces equating error under missing data and IPD conditions. A direct comparison with complete data conditions (i.e., without MICE) was not possible, since MICE was applied in all conditions with missing data. Therefore, its effectiveness can only be evaluated indirectly by comparing MCAR and MAR conditions.

Under MCAR conditions, the bias values of the B constant remain close to zero across all missing rates (K6: 0.0027; K7: 0.0039; see Table 2), suggesting that MICE effectively reduced equating bias when missingness was completely random. In contrast, under MAR conditions, bias increased noticeably with missing rate (K8: 0.0094; K9: 0.0251; see Table 2), indicating that MICE does not fully eliminate bias arising from systematic missingness.

Discussion and conclusion

In this study, the effects of IPD, missing data mechanisms (MCAR, MAR), and missing data rates (10% and 20%) on the accuracy of the SL transformation constants (A and B) in IRT-based test equating were examined through simulation. When the findings were considered in light of the four research questions, it became clear that the A constant remained largely stable across all conditions. The main pattern, however, emerged in the B constant, and was primarily driven by the missing-data mechanism.

With respect to the first research question, no meaningful bias was detected in either the A or B constants under conditions that included only IPD (K1–K5). Although RMSE values increased slightly as the magnitude of IPD increased, this increase did not reach a level of practical concern. This result partially aligns with the patterns reported by Wells et al. [14] and Han et al. [30]. Those studies similarly indicated that IPD alone does not substantially distort equating constants, although estimation errors tend to accumulate as the magnitude and proportion of drift increase. In the present study, the relatively moderate levels of IPD (10% and 20%) and drift magnitudes (0.2–0.5) suggest that deterioration in equating accuracy is largely a function of drift size, with stronger effects likely to emerge at higher levels [32].

Regarding the second research question, MCAR and MAR mechanisms produced clearly different effects on the B constant. Under MCAR conditions, bias values for B remained close to zero. In contrast, under MAR conditions, particularly at the 20% missing rate, a noticeable positive bias was observed. This finding indicates that when missingness is completely random, parameter estimation is largely preserved. However, when missingness is systematically related to ability, the B constant becomes measurably distorted.

Missing data mechanisms are among the key factors influencing equating accuracy. MAR is more complex than MCAR because the probability of missingness depends on the observed variables. When appropriate methods are not employed, this complexity can lead to higher levels of error and bias in parameter estimation. Indeed, the missing-data literature emphasizes that, under MAR conditions, violations of model assumptions or the use of inadequate methods can increase estimation error [19,56]. In this respect, correctly identifying the missing-data mechanism and handling it with appropriate techniques is critical for accurate equating results. The findings of the present study confirm this pattern within the context of the 3PL model and the SL method. Under MAR conditions, the dependence of missingness on individuals’ ability levels can bias both item and ability parameter estimates, ultimately reducing equating accuracy [21].

For the third research question, the combined effect of IPD and missing data produced greater bias than either factor alone. This interaction was especially pronounced in combinations involving MAR and moderate levels of IPD, whereas it remained much more limited in conditions that included MCAR. The highest levels of error were observed in condition K25 (IPD 20%, b = 0.50 × MAR 20%), where the bias for the B constant reached 0.029 and the RMSE value was 0.065. This pattern suggests that uncertainty stemming from two sources may reinforce each other, particularly when missingness is concentrated among lower-performing examinee groups, thereby amplifying the effect of IPD. To the best of our knowledge, no study in the literature directly examines the joint effect of these two factors. Therefore, the present findings underscore the importance of considering sources threatening equating accuracy in combination rather than in isolation.

Regarding the fourth research question, the effectiveness of MI can be indirectly evaluated by comparing MCAR and MAR conditions. Under MCAR, MI largely compensated for the B constant. Under MAR, however, bias continued to increase to a notable extent, especially when moderate levels of missingness and IPD co-occurred. MI, based on Rubin’s [57] rules, is known to improve parameter estimation under ignorable missing-data conditions, yet it does not fully eliminate the effects of systematic missingness patterns. While MCAR does not systematically distort parameter estimates, MAR conditions are known to produce noticeable bias. The present study shows that this general pattern also holds in the context of equating and highlights cases in which MAR-related bias cannot be fully corrected through MICE. These findings suggest that MICE was more effective under MCAR conditions than under MAR conditions, particularly when moderate levels of missingness (20%) and IPD (Δb = 0.50) co-occurred.

Overall, the findings of this study point to three main conclusions. First, the missing-data mechanism, particularly MAR, can introduce systematic bias into equating constants. Second, estimation error increases markedly as the rate of missing data rises. Third, although the effect of IPD alone is limited, it becomes considerably stronger when combined with MAR.

These findings also have several practical implications for operational test equating. First, missing-data patterns in operational assessments should be systematically monitored and documented prior to equating, as the mechanism, not only the rate, determines the extent of equating error. Therefore, in large-scale assessments, missing-data patterns should be systematically monitored and reported. Second, although MICE reduces error under MAR conditions, it does not eliminate it entirely. In the present study, imputation was conducted using item responses only, without incorporating person-level covariates. Drawing on general MI theory, including performance-related auxiliary variables, such as total score, in the imputation model may help reduce bias under MAR conditions by better capturing the cause of missingness [58]; however, the extent to which this approach improves equating accuracy warrants empirical investigation in future studies. Third, in the presence of IPD, the quality control of the anchor item set is critically important. Identifying and removing drifting anchor items prior to equating appears essential for limiting equating error under both MCAR and MAR conditions.

Several limitations of the present study should be acknowledged. First, the MAR mechanism was operationalized by assigning differential missingness probabilities based on examinees’ latent ability (θ). Although this approach is commonly used in IRT simulation studies and reflects realistic omission behavior in testing contexts [15,21], the resulting condition is more appropriately interpreted as a near-MAR mechanism rather than a strictly observed-variable MAR process. Future studies may extend this framework by incorporating directly observed auxiliary variables or by examining fully MNAR conditions.

A second limitation concerns the imputation strategy. In the present study, MICE was implemented using item response data only, without including person-level auxiliary variables such as total score or subgroup membership. Including such variables may improve the plausibility of the MAR assumption and further reduce residual bias under missing data conditions [56,58]. Future research may also consider comparing MICE with no-imputation or complete-case approaches to evaluate the relative effectiveness of different missing data handling strategies in IRT equating contexts.

Additional design assumptions should also be considered when interpreting the generalizability of the findings. First, the guessing parameter (c) was fixed at 0.20 across all items, which simplified estimation but does not reflect the item-level variability in pseudo-guessing typically observed in operational testing programs. Second, item parameter drift was simulated exclusively in the positive direction to facilitate the detectability of equating error and isolate the effect of drift magnitude. Accordingly, the findings may be most applicable to testing conditions in which drift occurs predominantly in a single direction. Future studies should examine how item-level variation in c parameters and bidirectional drift patterns interact with missing data mechanisms to influence equating accuracy.

The present study examined the joint effects of item parameter drift and missing data on IRT-based test equating accuracy within a unified simulation framework. Three main conclusions emerge from the findings. First, the missing data mechanism, rather than the rate of missingness alone, is the primary driver of equating error, with MAR conditions introducing substantial positive bias in the B constant that cannot be fully mitigated by MICE. Second, item parameter drift alone produced relatively limited bias at the levels examined, but amplified equating error considerably when combined with MAR. Third, while MICE effectively reduces error under MCAR, its performance is limited under systematic missingness conditions. Together, these findings underscore the importance of diagnosing missing data mechanisms prior to equating, monitoring anchor item stability, and developing more robust imputation strategies for operational testing programs, particularly in psychological assessment contexts where equating errors may affect individual-level decisions.

References

  1. 1. Kolen MJ, Brennan RL. Test equating: methods and practices. New York: Springer. 1995.
  2. 2. Baker FB. The basics of item response theory. College Park (MD): ERIC Clearinghouse on Assessment and Evaluation. 2001.
  3. 3. Hori K, Fukuhara H, Yamada T. Item response theory and its applications in educational measurement Part II: Theory and practices of test equating in item response theory. WIREs Computational Stats. 2020;14(3).
  4. 4. Keller LA, Keller RR. The effect of changing content on IRT scaling methods. Appl Meas Educ. 2015;28(2):99–114.
  5. 5. Wu T, Kim SY, Westine C, Boyer M. IRT observed-score equating for rater-mediated assessments using a hierarchical rater model. J Educ Meas. 2025;62(1):145–71.
  6. 6. Crocker L, Algina J. Introduction to classical and modern test theory. New York: Harcourt. 1986.
  7. 7. Hanson BA, Beguin AA. Obtaining a common scale for item response theory item parameters using separate versus concurrent estimation in the common-item equating design. Appl Psychol Meas. 2002;26(1):3–24.
  8. 8. Bock RD, Muraki E, Pfeiffenberger W. Item pool maintenance in the presence of item parameter drift. J Educ Meas. 1988;25(4):275–85.
  9. 9. Goldstein H. Measuring changes in educational attainment over time: problems and possibilities. J Educ Meas. 1983;20(4):369–77.
  10. 10. Donoghue JR, Isham SP. A comparison of procedures to detect item parameter drift. Appl Psychol Meas. 1998;22(1):33–51.
  11. 11. Risk NM. The impact of item parameter drift in computer adaptive testing (CAT). Urbana (IL): University of Illinois at Urbana-Champaign. 2015.
  12. 12. Li X. An investigation of the item parameter drift in the examination for the certificate of proficiency in English (ECPE). Spaan Fellow Work Pap Second Foreign Lang Assess. 2008;6:1–28.
  13. 13. Michaelides MP. A Review of the Effects on IRT Item Parameter Estimates with a Focus on Misbehaving Common Items in Test Equating. Front Psychol. 2010;1:167. pmid:21833230
  14. 14. Wells CS, Subkoviak MJ, Serlin RC. The effect of item parameter drift on examinee ability estimates. Appl Psychol Meas. 2002;26(1):77–87.
  15. 15. Pohl S, Grafe L, Rose N. Dealing with omitted and not-reached items in competence tests: evaluating approaches accounting for missing responses in item response theory models. Educ Psychol Meas. 2014;74(3):423–52.
  16. 16. Tabachnick BG, Fidell LS. Using multivariate statistics. 6th ed. Boston (MA): Pearson. 2013.
  17. 17. Kang H. The prevention and handling of the missing data. Korean J Anesthesiol. 2013;64(5):402–6. pmid:23741561
  18. 18. RUBIN DB. Inference and missing data. Biometrika. 1976;63(3):581–92.
  19. 19. Little RJA, Rubin DB. Statistical analysis with missing data. New York: Wiley. 2002. https://doi.org/10.1002/9781119013563
  20. 20. Waterbury GT. Missing Data and the Rasch Model: The Effects of Missing Data Mechanisms on Item Parameter Estimation. J Appl Meas. 2019;20(2):154–66. pmid:31120433
  21. 21. Finch H. Estimation of item response theory parameters in the presence of missing data. J Educ Meas. 2008;45(3):225–45.
  22. 22. Graham JW. Missing data analysis: making it work in the real world. Annu Rev Psychol. 2009;60:549–76. pmid:18652544
  23. 23. Allison PD. Missing data techniques for structural equation modeling. J Abnorm Psychol. 2003;112(4):545–57. pmid:14674868
  24. 24. Olinsky A, Chen S, Harlow L. The comparative efficacy of imputation methods for missing data in structural equation modeling. Eur J Oper Res. 2003;151(1):53–79.
  25. 25. van Buuren S. Multivariate imputation by chained equations. R Package version 3.19.0. https://cran.r-project.org/web/packages/mice/mice.pdf
  26. 26. Asiret S, Omur Sunbul S. Effect of missing data on test equating methods under NEAT design. Int J Psychol Educ Stud. 2023;10(3):702–13.
  27. 27. Ozdemir G, Atar B. Investigation of the missing data imputation methods on characteristic curve transformation methods used in test equating. J Meas Eval Educ Psychol. 2022;13(2):105–16.
  28. 28. Sinharay S, Holland PW. The missing data assumptions of the nonequivalent groups with anchor test (neat) design and their implications for test equating. ETS Research Report Series. 2009;2009(1).
  29. 29. Chen DF. Impact of item parameter drift on IRT linking methods. Greensboro (NC): University of North Carolina. 2021.
  30. 30. Han KT, Wells C, Sireci SG. The impact of multidirectional item parameter drift on IRT scaling coefficients and proficiency estimates. Appl Meas Educ. 2012;25(2):97–117.
  31. 31. Jurich DP, DeMars CE, Goodman JT. Investigating the impact of compromised anchor items on IRT equating under the nonequivalent anchor test design. Appl Psychol Meas. 2012;36(4):291–308.
  32. 32. Juric D, Liu C. Detecting item parameter drift in small sample Rasch equating. Appl Meas Educ. 2023;36(4):326–39.
  33. 33. Kopp JP, Jones AT. Impact of item parameter drift on Rasch scale stability in small samples over multiple administrations. Appl Meas Educ. 2020;33(1):24–33.
  34. 34. Lee WC, Ban JC. A comparison of IRT linking procedures. Appl Meas Educ. 2009;23(1):23–48.
  35. 35. Yurtcu M, Guzeller CO. Investigation of equating error in tests with differential item functioning. Int J Assess Tools Educ. 2018;5(1):50–7.
  36. 36. Dooley K. In: Baum J, editor. London: Blackwell. 2002. p. 829–48.
  37. 37. Feinberg RA, Rubright JD. Conducting Simulation Studies in Psychometrics. Educational Measurement. 2016;35(2):36–49.
  38. 38. DeMars CE. Detection of item parameter drift over multiple test administrations. Appl Meas Educ. 2004;17(3):265–300.
  39. 39. Han KT, Wells CS, Sireci SG. The impact of multidirectional item parameter drift on IRT scaling coefficients and proficiency estimates. Appl Meas Educ. 2012;25(2):97–117.
  40. 40. Kim SH, Cohen AS. In: Chicago (IL), 1991.
  41. 41. Sass DA, Schmitt TA, Walker CM. Estimating non-normal latent trait distributions within item response theory using true and estimated item parameters. Appl Meas Educ. 2008;21(1):65–88.
  42. 42. Brown A, Croudace TJ. Scoring and estimating score precision using multidimensional IRT. In: Reise SP, Revicki DA, editors. Handbook of item response theory modeling: applications to typical performance assessment. New York: Routledge. 2015. p. 307–33.
  43. 43. Han KT, Guo F. Potential impact of item parameter drift due to practice and curriculum change on item calibration in computerized adaptive testing. Reston (VA): Graduate Management Admission Council. 2011.
  44. 44. Enders CK. Applied missing data analysis. New York: Guilford Press. 2010.
  45. 45. Robitzsch A. On the Treatment of Missing Item Responses in Educational Large-Scale Assessment Data: An Illustrative Simulation Study and a Case Study Using PISA 2018 Mathematics Data. Eur J Investig Health Psychol Educ. 2021;11(4):1653–87. pmid:34940395
  46. 46. Dai S, Vo TT, Kehinde OJ, He H, Xue Y, Demir C. Performance of polytomous IRT models with rating scale data: an investigation over sample size, instrument length, and missing data. Front Educ. 2021;6:721963.
  47. 47. Chalmers RP. Mirt: A multidimensional item response theory package for the R environment. J Stat Softw. 2012;48:1–29.
  48. 48. van Buuren S, Groothuis-Oudshoorn K. Mice: multivariate imputation by chained equations in R. J Stat Softw. 2011;45(3):1–67.
  49. 49. Graham JW, Olchowski AE, Gilreath TD. How many imputations are really needed? Some practical clarifications of multiple imputation theory. Prev Sci. 2007;8(3):206–13. pmid:17549635
  50. 50. White IR, Royston P, Wood AM. Multiple imputation using chained equations: Issues and guidance for practice. Stat Med. 2011;30(4):377–99. pmid:21225900
  51. 51. Battauz M. EquateIRT: An R package for IRT test equating. J Stat Softw. 2015;68(7):1–22.
  52. 52. Stocking ML, Lord FM. Developing a common metric in item response theory. Appl Psychol Meas. 1983;7(2):201–10.
  53. 53. Kolen MJ, Brennan RL. Test equating, scaling, and linking: methods and practices. 3rd ed. New York: Springer. 2014. https://doi.org/10.1007/978-1-4939-0317-7
  54. 54. Harwell M. A Strategy for Using Bias and RMSE as Outcomes in Monte Carlo Studies in Statistics. J Mod Appl Stat Methods. 2019;17(2).
  55. 55. Hoogland JJ, Boomsma A. Robustness studies in covariance structure modeling: an overview and a meta-analysis. Sociol Methods Res. 1998;26(3):329–67.
  56. 56. Enders CK. Applied missing data analysis. 2nd ed. New York: Guilford Press. 2022.
  57. 57. Rubin DB. Multiple imputation for nonresponse in surveys. New York: John Wiley & Sons. 1987.
  58. 58. Collins LM, Schafer JL, Kam CM. A comparison of inclusive and restrictive strategies in modern missing data procedures. Psychol Methods. 2001;6(4):330–51. pmid:11778676