Figures
Abstract
Background
Accurate assessment of child mortality rates is essential for policy formulation, resource allocation, and tracking progress toward Sustainable Development Goal-3 (SDG-3). The study of child mortality provides useful information to know the demographic situation of country. The death is vital event and recorded through the civil registration system. But in many developing countries, the quality of registered death is not very much reliable due to illiteracy and ignorance in population. So, the lack of accurate registration of death has forced demographers to explore the indirect techniques for estimating child mortality. In this paper, authors have utilized an indirect technique for estimating neonatal, infant and under-five mortality rate by using data on proportion of dead children and proportion of 4 + birth-order.
Methods
The method is mainly based on technique of linear line regression analysis, where the proportion of dead infants among all children born to currently married females (15–49 years) and proportion of 4 + birth order are taken as the independent variables and the neonatal mortality rate, or infant mortality rate, or under-five mortality rate are used as the dependent variable. This study is based on data collected in fifth round of National Family Health Survey (NFHS), 2019−21. After applying the inclusion criteria, a total of 512,408 currently married women aged 15–49 years were included in the analysis.
Result
For above mentioned child mortalities are calculated for India as well its major states. The actual and predicted mortality rates overlapped substantially when both predictors were used together, indicating the suitability of the proposed model. The combined-predictor models (proportion of dead children and proportion of 4 + birth-order) achieved higher R² values (0.91 for neonatal mortality rate (NMR), 0.92 for infant mortality rate (IMR), and 0.96 for under-five mortality rate (U5MR) along with the minor percentage differences between observed and estimated rates (mostly within ±10%) than single-predictor models.
Conclusion
The findings suggest that integrating both the predictors, proportion of dead children and the proportion of higher birth order children improves the accuracy of indirect estimates of NMR, IMR, and U5MR. This approach may serve as a practical tool for generating timely mortality estimates in contexts where civil registration and survey data are incomplete or infrequent.
Citation: Singh M, Singh A, Tiwari AK (2026) Indirect technique to estimate neonatal, infant, and under-five mortality rates for India and its states: A population-based cross-sectional study. PLoS One 21(9): e0353613. https://doi.org/10.1371/journal.pone.0353613
Editor: Ashish Wasudeo Khobragade, All India Institute of Medical Sciences - Raipur, INDIA
Received: February 1, 2025; Accepted: August 26, 2026; Published: September 15, 2026
Copyright: © 2026 Singh et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.
Data Availability: Third party data was obtained for this study from the DHS Program. Data may be requested from the DHS Program after creating an account and submitting a concept note. More access information can be found on the DHS Program website (https://dhsprogram.com/data/Access-Instructions.cfm). Interested researchers would be able to access these data in the same manner as the authors. The authors had no special access privileges that others would not have.
Funding: The author(s) received no specific funding for this work.
Competing interests: The authors have declared that no competing interests exist.
Abbreviations: CRVS, Civil Registration and Vital Statistics; CVPP, Cross Validation Prediction Power; DHS, Demographic and Health Survey; DLHS, District Level Household and Facility Survey; IMR, Infant Mortality Rate; IIPS, International Institute for Population Sciences; MOHFW, Ministry of Health and Family Welfare; NFHS, National Family Health Survey; NMR, Neonatal Mortality Rate; Ml, Mortality index; SDG, Sustainable Development Goal; SRS, Sample Registration System; U5MR, Under-Five Mortality Rate; UN, United Nations; UN IGME, United Nations Inter-agency Group for Child Mortality Estimation
Background
Neonatal, infant, and under-five mortality are critical indicators of a nation’s healthcare system, socioeconomic development, and overall well-being. In India, despite significant advancements in healthcare infrastructure and interventions, high mortality rates among newborns and young children persist, particularly in rural and marginalized communities. Accurate estimation of these mortality rates is essential for policy formulation, resource allocation, and monitoring progress towards Sustainable Development Goal 3 (SDG 3)—ensuring healthy lives and promoting well-being for all at all ages. Policy planners and health administrators rely on such data to design development strategies, address population health needs, and evaluate public health initiatives. But vital registration system often fails to capture all events due to limited administrative reach, illiteracy, and lack of awareness. While direct surveys is valuable for collecting detailed information but are expensive, time-consuming, and logistically challenging to conduct on a large scale. Given these limitations, indirect estimation methods offer a practical and cost-effective alternative by using readily available demographic, health, and socioeconomic indicators to infer mortality patterns and trends.
Demographers have been compelled to develop methods for predicting fundamental demographic parameters from erroneous or partial data due to the lack of reliable population counts and registration of important events. As we know, Brass focuses on methods for estimating mortality rates when data are limited or flawed [1]. His finding suggests techniques to derive reliable estimates of mortality even when data are incomplete or of poor quality. Another technique was used for estimating child mortality from information on proportion of dead children amongst children ever born to women with reference to the length of her marriage [2]. The alternate indirect techniques, such as those proposed for situations where mortality has been declining include the works of [3–5]. Islam and Alam used child mortality index (Ml) to calculate child mortality [6]. Neupert proposed a method to estimate infant mortality rates in small areas with limited data [7]. By using indirect estimation techniques based on fertility, childlessness, and total mortality data, it provides accurate estimates of infant mortality rates for various small geographic areas.
It is common knowledge that the percentage of dead among all children born to women in standard age groups can be used to estimate child mortality. It was first given by (Coal and Demeny, Brass and Coale) widely known as the Brass method [1,8]. On the other hand, the proportion of births of order four or higher (4 + birth order) has been consistently linked to elevated child mortality risks through mechanisms such as “resource dilution” where parental resources and attention are spread across more children, and “maternal depletion” where closely spaced and frequent pregnancies reduce maternal health reserves [9]. Higher birth order, defined as the fourth or subsequent birth within a family, has been associated with an increased risk of child mortality [10]. While some studies have reported a clear association between higher birth order and elevated mortality rates among children [11]. In 2003, Yadava and Tiwari applied an indirect estimation technique to calculate infant and under-five mortality rates using the proportion of dead children alone for India and Bangladesh [12]. Therefore, in this paper, we have evaluated the adequacy of this technique using recent NFHS data and extended it by incorporating the proportion of higher birth-order (4+) along with the proportion of dead children to improve estimation accuracy. The regression line analysis was conducted and tested using data from fifth round of NFHS.
Materials and methods
The data
The National Family Health Survey (NFHS), 2019–21 data-set has been used in this study. The NFHS is a large-scale, nationally representative survey conducted in India under the Ministry of Health and Family Welfare. NFHS-5 fieldwork for India was carried out by 17 Field Agencies in two phases: phase one from June 17, 2019, to January 30, 2020, and phase two from January 2, 2020, to April 30, 2021. Data were collected from 636,699 households, comprising 724,115 women and 101,839 males [13]. It provides comprehensive data on various health and demographic indicators, including fertility, maternal and child health, family planning, and nutrition at national, state and district level. By using information on the number of births and deaths within this period, along with the age at death of the deceased children, researchers may identify mortality rates. The present study used data from the Individual Recode (IR) file of NFHS-5 (2019–21). After applying the inclusion criteria, a total of 512,408 currently married women aged 15–49 years were included in the analysis. The information on proportion of dead children to currently married females (15–49 yrs.) and proportion of 4 + birth order may be calculated from the information given in the survey.
The choice of predictors
In such estimation techniques, the choice of predictor (s) is very critical and important. The inappropriate choice may leads to misleading or not useful conclusion. The basic need in the choice of the variables(s) is that of significant correlation between the dependent and independent variable(s) and independent variable is easily obtainable. In selecting the predictor variables for the present study, we focused on indicators those are both theoretically and empirically robust in their association with child mortality. The information on proportion of dead children to currently married females 15–49 years is a well-established indirect indicator of child mortality and proportion of four or higher orders birth has been consistently linked to child mortality. Therefore, proportion of dead children to currently married females (15–49 years) have been identified from [12] and along with proportion of dead children we have also considered one more predictor variable, i.e., proportion of 4 + birth order for this study [9].
Methodology
The essentially method is based on the technique of regression line. The regression line concept has been utilized to investigate the relationship between the dependent and independent variable(s).
The straight-line simple regression model with respect to population parameters β0 and β1 can be explained by:
Where X and Y are independent and dependent variables respectively.
β0 indicates the average value of Y when X=0, and
β1 the slope of the line which indicates the change in the Y for per unit change in the X
For more than one independent variables, multiple regression analysis is used. The mathematical form of multiple regression model as:
Where:
Yi is the ith value of dependent variable.
β0 is the intercept term, and
β1, β2,…. βn are regression coefficients representing the change in Y associated with a one-unit change in X1, X2,…., Xn respectively
εi is the error term
For the present study neonatal mortality rate (NMR)/ infant mortality rate (IMR)/ under five mortality rate (U5MR) was utilized as a dependent variable (Y). And the information on proportion of dead children among all children born to married mothers in the 15−49 age-group and proportion of 4 + birth orders considered as independent variable(s) namely X1 and X2 respectively [1,12]. The line of regression is drawn by taking major states of India using NFHS-5 data.
Mathematical form of fitted equations
For Neonatal mortality rates (per thousand live births),
Y = 458.41 X1 - 3.23, (R2 = 0.82) [X1 is the Proportion of dead children]
Y = 67.31 X2 + 10.28, (R2 = 0.80) [X2 is the Proportion of 4+ birth order]
Y = 267.66 X1 + 36.31 X2 + 0.875, (R2 = 0.91) [Both predictors X1 and X2 taken together]
For Infant mortality rates (per thousand live births),
Y = 599.23 X1 - 2.19, (R2 = 0.81) [X1 is the Proportion of dead children]
Y = 89.79 X2 + 15.13, (R2 = 0.83) [X2 is the Proportion of 4+ birth order]
Y = 325.62 X1 + 52.08 X2 + 3.68, (R2 = 0.92) [Both predictors X1 and X2 taken together]
For Under-five mortality rates (per thousand live births),
Y = 685.03 X1 + 0.386, (R2 = 0.86) [X1 is the Proportion of dead children]
Y = 100.53 X2 + 20.59, (R2 = 0.84) [X2 is the Proportion of 4+ birth order]
Y = 400.68 X1 + 54.13 X2 + 6.51, (R2 = 0.96) [Both predictors X1 and X2 taken together]
Thus, the proposed predictors (proportion of dead children along with the proportion of 4 + birth order) improve the value of R2 and provide better explanation of considered child mortalities estimates. The R² values indicate the proportion of variation in mortality explained by the regression models. Higher R² values represent better model fit, suggesting that the selected predictors account for a substantial proportion of the variability in mortality estimates.
Model validation of the fitted model
- 1. Cross Validation predication power
A predictive model’s estimated performance in practice or the degree to which the suggested model is population-stabilized must be determined. In this regard, Herzberg’s cross validity prediction power (CVPP) technique has been applied [14], and it is described as [15]:
where n is the number of instances, c is the correlation coefficient between the estimated and observed values of the dependent variables, and p is the number of explanatory factors in the model.
- 2. Shrinkage and Stability of R2
The model’s shrinkage is indicated by the standard adjustment applied to the coefficient of determination to account for the arbitrary effects of further sampling:
Where ρ2υ is Cross Validation predication power (CVPP) and R2 is the coefficient of determination of the model. The stability of R2 of the model is equal to (1- Shrinkage) means lower shrinkage value provides more stability of the model.
Note: All data management, data analysis, and statistical analyses were performed using Microsoft Excel and Stata Statistical Software, Version 16 (StataCorp LLC, College Station, TX, USA).
Results
Table 1 shows the proportion of dead children among all children born to married mothers in the 15–49 age groups and proportion of 4 + birth orders for India. Haryana, Gujarat, Maharashtra, Rajasthan, Punjab, Tamil Nadu, Karnataka, Orissa, West Bengal, Assam: these states exhibit lower mean numbers of children ever born compared to the national average. The Proportion of dead children is found to be high in four states, i.e., MP, UP, Bihar and Orissa. These states also have higher proportion of 4 + births.
Table 2 indicates the estimates of neonatal mortality rate (NMR) for India and its major states. Results presented in the table shows that the estimated values of NMR for most of the states are closest to observed value of NMR. The neonatal mortality rate for India is estimated 23.0 by using proposed all three regression models, but there are wide variations found among the considered states. Three states located in central and eastern part of country like MP, UP and Bihar has relatively high neonatal mortality rates. India and its 13 major states under study the difference between observed and estimated value of NMR is less than 10% for 5 (India, Gujarat, Maharashtra, Rajasthan and Karnataka), 7 (India, UP, MP, Bihar, Haryana, Gujarat and Assam) and 10 (India, UP, MP, Bihar, Haryana, Gujarat, Maharashtra, Tamil Nadu, Karnataka and Orissa) for different predictors like proportion of dead children, proportion of 4 + birth order and both proportion of dead children and 4 + birth order taken together respectively. To evaluate the suitability of the suggested approach, a cut-off difference of 10% was set between the observed and estimated values.
Table 3 shows the estimates for infant mortality rate under the proposed predictors for India and its major states. To evaluate the suitability of the suggested approach, a cut-off difference of 10% was set between the observed and estimated values for the infant mortality rates. Out of 14 observed values for India and its states, there are 8 states (India, Bihar, Haryana, Gujarat, Maharashtra, Rajasthan, Punjab and Karnataka), 7 states (India, MP, Haryana, Gujarat, Maharashtra, Karnataka and Assam) and 11 states (India, Uttar Pradesh, Madhya Pradesh, Bihar, Haryana, Gujarat, Maharashtra, Rajasthan, Karnataka, Orissa and Assam) were found to be below 10% for different predictors, i.e., proportion of dead children, proportion of 4 + birth order and both proportion of dead children and 4 + birth order taken together respectively. Utilizing this standard, we noted the estimated values and observed values showed that they were fairly near to one another.
Table 4 reveals that on an average the estimated value of U5MR for most of the states and India are closer to observed value. Again, utilizing the cut off 10% difference, there are only 11 states are found to be within the limit (India, Uttar Pradesh, Madhya Pradesh, Bihar, Haryana, Gujarat, Maharashtra, Punjab, Tamil Nadu, Karnataka and Assam), 9 states (India, Uttar Pradesh, Madhya Pradesh, Bihar, Haryana, Gujarat, Tamil Nadu, West Bengal and Assam)and 13 considered states except Punjab for proportion of dead children, proportion of 4 + birth order and both proportion of dead children and 4 + birth order taken together respectively.
The Table 2–4 also provides the 95% confidence interval for estimates of NMR, IMR and U5MR. For NMR, the observed values lie in the 95% confidence intervals for almost all states as well as India for proposed predictors, whereas the Yadava and Tiwari (2003) model and the higher-order birth alone as a predator miss the observed values in two and three states, respectively. For IMR, the regression model using both predictors demonstrate the strongest performance, with observed values consistently falling within its confidence limits and exhibiting smaller deviations from the true estimates. The proportion of dead-children model also performs reasonably well but shows notable mismatches in states such as Tamil-nadu and Orissa. The proportion of higher-order birth model proves less reliable, failing to capture observed values in approximately three states, including Bihar, Rajasthan, and West Bengal. For U5MR, the combined model yields the most consistent and accurate results, nearly always encompassing the observed values within its 95% confidence intervals and minimizing deviations. The proportion of dead-children model performs moderately well but with larger discrepancies in states such as Punjab and Orissa, while the higher-order birth model fails in around three states, notably Bihar, Rajasthan, and Maharashtra.
Table 5 presents regression models along with the coefficient of determination (R2), shrinkage, and stability of different mathematical models for estimation of neonatal mortality rate, infant mortality rate, and under-five mortality rate per 1,000 live births for India and some of its major states, based on data from the NFHS conducted between 2019−21. Among NMR, the stability values are 0.7652, 0.7442 and 0.8288 for first, second and third model respectively and shrinkage is minimum for third model. Similarly, among IMR, shrinkage is found to be 0.2454, 0.2239 and 0.1711 for 1st, 2nd and 3rd model respectively. And for U5MR, the stability is found to be lowest for combined model.
Discussion
The present study considers the proportion of dead children among all children born to married women aged 15–49 years with the proportion of 4 + birth order children provide a more accurate and robust estimation of neonatal mortality rate (NMR), infant mortality rate (IMR), and under-five mortality rate (U5MR) in India than relying on either predictor alone. This approach advances earlier demographic research that highlighted the importance of multi-variable indirect estimation techniques in data-limited settings.
A notable strength of our approach is that indirect estimates were within ±10% of observed NFHS-5 values for most states, demonstrating practical utility where direct estimates are infeasible. This confirms the relevance of indirect methods with easily available data. Importantly, our analysis also illustrates how indirect methods can fill interim data gaps. The ability to generate reliable subnational estimates using readily available variables is particularly valuable for monitoring progress toward Sustainable Development Goal (SDG)- 3.
From a policy perspective, the proposed approach offers a rapid, low-cost means to estimate child mortality indicators in settings where CRVS systems remain weak. State and district health planners can use such interim estimates to identify high-burden regions, allocate resources for maternal and child health interventions, and monitor SDG 3 progress between NFHS rounds. At the global level, our approach shares conceptual similarities with the estimation framework employed by the UN Inter-agency Group for Child Mortality Estimation (UN IGME), which also relies on indirect indicators and survey-based measures to overcome data limitations [16]. However, unlike the UN IGME, which uses complex statistical modelling and Bayesian smoothing techniques to produce harmonised cross-country estimates, our model is specifically designed for subnational applications within India. It emphasizes methodological simplicity, transparency, and ease of replication at the state level, making it a practical tool for decentralized health planning. In this sense, it complements rather than replaces global estimation efforts. In summary, the study highlights the value of strengthening traditional indirect estimation methods by integrating multiple predictors, thereby enhancing both precision and applicability. The proposed model can serve as a cost-effective tool for policymakers in resource-limited settings, enabling evidence-based decision-making and more targeted child health interventions.
Limitations
This study has certain limitations that should be acknowledged. First, the regression-based approach assumes a stable and consistent relationship between the selected predictors and child mortality outcomes across time and states, an assumption that may be affected by changes in healthcare access, fertility patterns, and socio-economic conditions. Second, while NFHS data provide nationally representative estimates, so there may be chance of occurrence of non-sampling and sampling errors, particularly in smaller states or subpopulations. Furthermore, while the regression residuals were generally small, they do suggest the presence of unexplained variation, which highlights the limitations.
Conclusion
On the basis of the above results, it is reasonable to conclude that the information on proportion of dead children among all children born to married mothers in the 15–49 age-group and proportion of 4 + birth orders provide the better estimates of NMR, IMR and U5MR. The results also reveals that the estimates of these child mortalities are more closed to the observed values compare to estimates calculated using Yadava and Tiwari. While there is some shrinkage in the R2 values across the models, the stability remains relatively high for the model that when both predictors taken together, indicating robustness in its predictive ability. The percent difference between observed and estimated values of child mortalities by proposed regression models and positions of observed values in confidence interval indicate the accuracy of technique.
On the basis of above study, we may conclude that the proposed regression has the potential to provide quick and better estimates of neonatal, infant and under-five mortality from a limited and easily available data.
Future research directions
Future studies could focus on validating the proposed indirect estimation method using alternative datasets, such as the Sample Registration System (SRS) or District Level Household and Facility Survey (DLHS), to test its robustness across different data sources and time periods. Incorporating additional predictors—such as maternal education, household wealth, and access to healthcare indicators etc. may improve the explanatory power and adaptability of the model’s. Moreover, longitudinal analyses could be used to examine how the relationship between predictors and mortality outcomes evolves over time, enabling periodic recalibration.
References
- 1.
Brass W. Methods for estimating faculty and mortality from limited and defective data. Chapel Hill, North Carolina: Laboratories for Population Statistics. 1975.
- 2. Hill K. Indirect estimation of child mortality: an assessment of methods. PLoS Med. 2013;10(5):e1001388.
- 3. Trussel J. A re-estimation of the multiplying factors for the brass technique for determining childhood survival rate. Population Studies. 1975;29.
- 4. Palloni A, Heligman L. Re-estimation of structural parameters to obtain estimates of mortality in developing countries. Popul Bull UN. 1985;(18):10–33. pmid:12314306
- 5. Preston SH, Palloni A. Fine-tuning brass-type indirect estimates of child mortality. Population Bulletin of the United Nations. 1978;10:72–91.
- 6. D’Souza S. The assessment of preventable infant and child deaths in developing countries: some applications of a new index. World Health Stat Q. 1989;42(1):16–25. pmid:2711700
- 7. Neupert RF, Menjívar RE, Castilla REF. Indirect estimation of infant mortality in small areas: Applications and methodological insights. Revista Brasileira de Estudos de População. 2012;29(1):123–40.
- 8.
Coale A, Demeny P. Regional model life table and state populations. Princeton University Press. 1966.
- 9. Conde-Agudelo A, Rosas-Bermúdez A, Kafury-Goeta AC. Birth spacing and risk of adverse perinatal outcomes: a meta-analysis. JAMA. 2006;295(15):1809–23. pmid:16622143
- 10. Hobcraft JN, McDonald JW, Rutstein SO. Demographic Determinants of Infant and Early Child Mortality: A Comparative Analysis. Population Studies. 1985;39(3):363–85.
- 11.
Rutstein S. Effect of birth intervals on mortality and health: multivariate cross-country analyses. In Conference on Optimal Birth Spacing for Central America, Antigua, Guatemala. 2003.
- 12. Yadava RC, Tiwari AK. An indirect technique for estimations of infant and child mortality: Data analysis from India and Bangladesh. Health and Population Perspectives and Issues. 2003;26(2):67–73.
- 13.
International Institute for Population Sciences (IIPS), ICF. National Family Health Survey (NFHS-5), 2019-21: India: Volume II. Mumbai: IIPS. 2021.
- 14.
Herzberg PA. The parameters of cross-validation. University of Illinois at Urbana-Champaign. 1967.
- 15. Tiwari AK, Singh BP, Patel V. Retrospective Study of Investigation of Possible Predictors for Total Fertility Rate in India. JSRR. 2020;:111–9.
- 16.
United Nations. Manual X: Indirect Techniques for Demographic Estimation. New York: United Nations. 1983.