Enhancing the accuracy of Employer Health Benefit Cost Estimates: A 2024 Methodology Update
Understanding the true cost of employer-sponsored health insurance is crucial for businesses, employees, and policymakers alike. At KFF, we are committed to providing the most accurate and reliable data possible through our annual Employer Health Benefits survey (EHBS). This article details important methodological improvements implemented in the 2024 survey, designed to minimize bias and enhance the precision of our premium estimates.
addressing Missing Data: A shift Towards Predictive Modeling
In the 2024 EHBS, we encountered instances were respondents were unable to provide their firm’s single coverage premium, or their responses contained inconsistencies. Specifically, 9.8% of responses required attention. to address this, and minimize potential non-response bias, we employed data imputation techniques.
Previously, we utilized a “hot deck” approach, matching missing values with premiums from firms with similar characteristics. However, this year we moved to a more refined method. This new approach leverages the power of machine learning to provide more accurate estimates.
Here’s a breakdown of the updated process:
* Combined Premium Estimation: When both single and family premiums were missing, we estimated the single premium based on specific firm characteristics. These include cost-sharing arrangements, deductible amounts, and firm demographics.
* Relationship-based Imputation: following single premium estimation, family premiums and individual worker contributions (when missing) were imputed using a ”hot deck” approach, but based on their relationship to the newly estimated single premium.
* Premium Ratio Imputation: If a firm did provide a family premium, the single premium was imputed using the ratio between the two.
* Overall Impact: This change affected approximately 8% of the total responses.
Introducing a Random Forest Machine Learning Model
The core of our improved imputation process is a random forest machine learning model. This model was rigorously trained using EHBS data from 2021-2023. We carefully selected the most relevant features using stepwise regression and fine-tuned the model’s parameters with a grid search algorithm.
Why is this better?
Compared to the traditional hot-decking method, the random forest model demonstrates a considerably improved ability to explain the variation in observed data:
* R-squared: 0.2071 (Random Forest) vs. 0.00018540 (Hot-Decking)
* This means the model captures a much larger proportion of the variability in premiums,leading to more reliable estimates.
Minimal Impact on Overall Premiums, significant Gains in Precision
While this methodological change represents a significant improvement in accuracy, the overall impact on the 2024 premium estimates is relatively small. We estimate the 2024 premium would be only 0.8% different if we had continued using the previous method.
Though,the real benefit lies in the increased precision of premium estimates for specific demographic subgroups.You can now have greater confidence in the data when analyzing costs for different populations.
Clarity and Further Details
We believe in complete transparency regarding our methodology. For a more detailed explanation of the 2024 survey, including data collection and processing procedures, please refer to the Survey Methodology section of the 2024 KFF Employer Health Benefits Survey report.
Data Sources & Citations:
* Bureau of Labor Statistics. Current Employment Statistics-CES (National). https://www.bls.gov/ces/publications/highlights/highlights-archive.htm (Cited 2024 Aug 1)
* Bureau of Labor Statistics,Mid-Atlantic Information Office. Consumer Price Index historical tables for, U.S. city Average (1967 = 100) of Annual Inflation. https://www.bls.gov/regions/mid-atlantic/data/consumerpriceindexhistorical1967base_us_table.htm (Cited 2