AI Models & Platforms
Google AI Model Tops CDC Evaluation of Flu Hospitalization Forecasts

The Centers for Disease Control and Prevention on September 30, 2026 published its evaluation of the 2025-2026 FluSight forecasting season, identifying Google’s Google_SAI-FluEns as the top-performing model among individual team submissions for forecasting U.S. flu hospital admissions. Google said its entry ranked first among the season’s 39 eligible models.
In Google’s announcement, Zahra Shamsi of Google Research said the forecasts were developed using Empirical Research Assistance (ERA), an AI tool the company describes as generating optimization algorithms across a range of scientific fields. Google said research on ERA was published in the journal Nature and that the tool’s underlying technology is available to trusted testers through its experimental science tools.
The Season’s Results
According to the CDC’s evaluation report, 34 teams contributed influenza hospital admission forecasts from 53 unique models during the season. Thirty-nine models met the analysis’s inclusion criteria by submitting at least 75% of forecast targets; CDC excluded national forecasts from scoring because of differences in scale and Puerto Rico forecasts because of data availability. Of the 39 included models, 33 performed better than the FluSight baseline, a reference model that projects the prior week’s admissions forward.
The FluSight ensemble, the combined forecast CDC uses to communicate its messaging, ranked seventh overall for the season on the primary performance metric and was one of 12 models that consistently outperformed the baseline in all jurisdictions. Among the individual team submissions, Google_SAI-FluEns was the top performer.
CDC characterized the 2025-2026 season as moderate under its preliminary in-season severity assessment. The highest weekly number of hospital admissions exceeded 40,000, with hospitalizations beginning to rise in mid-November 2025 and peaking nationally in the week ending December 27, 2025.
The report also documents where forecasting struggled. It states that ensemble 50% and 95% prediction intervals failed to anticipate the increases in hospitalizations observed in late December 2025 and the decreases in mid-January 2026. Around the season’s peak week, less than 25% of two-week-horizon forecast prediction intervals across jurisdictions contained the observed values, and FluSight ensemble coverage stabilized near 95% starting in February 2026.
How CDC Scored the Forecasts
CDC solicited weekly influenza hospital admission forecasts from academic, industry, and government teams from November 2025 through May 20, 2026, and first published them to its FluSight webpage on December 12, 2025. The main forecasting target was weekly influenza hospital admissions for the current week and up to three weeks in the future for the United States, each state, Puerto Rico, and Washington, D.C., using data from the National Healthcare and Safety Network. The final target data used for scoring were published July 1, 2026.
Forecasts were evaluated primarily on relative weighted interval score, a standard measure of how closely a forecast’s set of prediction intervals aligns with what was ultimately observed. The metric is calculated relative to the baseline, with a value below 1 representing performance better than the baseline, and scores were computed on log-transformed admission counts as a geometric mean of pairwise score ratios between models. Forecasts were also evaluated on 50% and 95% coverage, the frequency with which prediction intervals captured the values later observed. Models were categorized from submitted metadata as statistical, mechanistic, artificial intelligence or machine learning, or ensemble, though those categories were not used in scoring.
CDC has hosted influenza forecasting challenges annually since the 2013-2014 season, except for 2020-2021, when limited influenza activity prevented stable forecasts. An evaluation of FluSight’s forecasts of flu emergency department visits is forthcoming.
The Research Behind Google’s Entry
A preprint submitted May 15, 2026, by Sarah Martinson, Michael P. Brenner, Martyna Plomecka, Brian P. Williams, Nicholas G. Reich, and Zahra Shamsi describes the system behind the forecasts: an autonomous system in which a large language model guides a tree search that writes, tests, and refines executable forecasting software. The authors report that in a fully prospective, real-time evaluation during the 2025-2026 U.S. respiratory season, the system produced models for influenza, COVID-19, and respiratory syncytial virus using varied methodologies, and that an ensemble built from the machine-generated models performed as well as or better than CDC’s human-curated hub ensembles out-of-sample. They also report that the system produced workable forecasts for RSV under data-scarce “cold start” conditions.
In an April 29, 2026 Google Research post, the team said it began submitting weekly flu forecasts for every U.S. state and at horizons up to four weeks when the CDC challenge opened in November 2025, and that it later joined CDC’s live forecasting hubs for COVID-19 and RSV. The post said public leaderboards for flu and COVID-19 run by Reich, a University of Massachusetts Amherst biostatistics professor and a consultant on the project, showed Google performing at or near the top during its submission periods. The post names Shamsi, Martinson, Reich, Plomecka, and Williams as the leads of the epidemiological forecasting work.
Why CDC Runs Flu Forecasting
CDC’s flu forecasting program is designed to predict in advance when and where increases in flu hospitalizations may occur, according to the agency’s program overview. CDC says accurate forecasts can inform antiviral treatment allocation, preparation for surges in flu-related hospitalizations, the distribution of health care staff, hospital beds, and treatment resources, as well as messaging on mitigation measures and vaccination timing. The effort began with the 2013 “Predict the Influenza Season Challenge,” and weekly hospitalization forecasts have been based on NHSN data since the 2021-2022 season.
Google said the evaluation result validates its confidence that AI combined with human ingenuity will improve the ability to forecast diseases worldwide. CDC, FluSight partners, and stakeholders gather at the end of every forecasting season to review forecasting approaches, discuss accuracy, and plan future seasons, including the addition of new forecasting targets.












