Tag: Disease Surveillance & Epidemiology

  • XGBoost forecasted pulmonary tuberculosis better than ARIMA or Prophet

    What the study found

    The study found that an XGBoost machine-learning model produced more accurate monthly forecasts of pulmonary tuberculosis (PTB, a form of tuberculosis that affects the lungs) than seasonal ARIMA or Facebook Prophet. On unseen data from Fuzhou, XGBoost had much lower error values and better followed the observed decline and seasonal pattern.

    Why the authors say this matters

    The authors conclude that, for cities with nonlinear waning epidemics and seasonally shrinking amplitude, XGBoost may be a better forecasting tool than traditional time-series methods. They say this could support monthly PTB early-warning, resource pre-positioning, and targeted control in comparable high-density coastal urban settings.

    What the researchers tested

    The researchers used 168 monthly PTB case reports from Fuzhou covering January 2009 to December 2022, plus a 24-month prospective validation set from 2023 to 2024. They developed and tested three forecasting frameworks: seasonal ARIMA with automatic order selection, Facebook Prophet with multiplicative seasonality and change-point detection, and XGBoost using 1- to 12-month lagged incidence, calendar variables, and linear-trend covariates.

    What worked and what didn't

    All three models fit the training data closely. On unseen data, XGBoost performed best, with lower RMSE, MAE, and MSE than ARIMA or Prophet, and its residuals remained approximately white noise. Prophet slightly overestimated seasonal amplitude, while ARIMA accumulated trend extrapolation bias.

    What to keep in mind

    The study was based on one city, Fuzhou, and the findings are presented for comparable high-density coastal urban settings rather than all places. The abstract does not describe other limitations beyond the model comparison and validation design.

    • XGBoost outperformed seasonal ARIMA and Prophet on the 2023-2024 validation data.
    • The data came from 168 monthly PTB case reports in Fuzhou, China, from 2009 to 2022.
    • XGBoost better tracked the observed 5.7% annual decline and narrowing spring-summer double peaks.
    • Prophet slightly overestimated seasonal amplitude, and ARIMA showed trend extrapolation bias.
    • The authors say the approach may support PTB early-warning and resource planning.
  • CDC surveillance databases showed widespread unexplained pauses

    What the study found

    Many U.S. Centers for Disease Control and Prevention (CDC) surveillance databases that had been updated at least monthly were paused or no longer current by late 2025. The authors describe these as “unexplained pauses” in the public evidence base used for health policy.

    Why the authors say this matters

    The authors conclude that real-time federal surveillance informs clinical guidance and public health policy, so long pauses may have compromised evidence for decision making by clinicians, administrators, professional organizations, and policymakers. They also suggest federal databases should have minimum transparency standards, including current update status, a reason if paused, and the next expected update with criteria for resumption.

    What the researchers tested

    The researchers audited the CDC public data catalog on 28 October 2025 to identify database records that had previously been updated at least monthly. They then used each database's stated periodicity, plus a 30-day grace period, to classify databases as current or paused, and they checked whether pauses persisted as of 2 December 2025.

    What worked and what didn't

    Of 1,359 catalog records examined, 82 had been updated at least monthly. Forty-four of those databases were current and 38 were paused; 34 of the paused databases had no data entries within 6 months of the analysis date, while 4 had paused more recently. Among the paused databases, 33 were vaccination-related, 4 of the remaining 5 focused on respiratory diseases, and 1 addressed public health drug overdose deaths; by 2 December 2025, only 1 paused database had been updated.

    What to keep in mind

    The summary provides an audit of CDC catalog records at two points in time and does not explain why the pauses occurred. The available abstract does not describe effects on specific policies or describe any limitations beyond what is implied by the catalog-based approach.

    • The audit examined 1,359 CDC catalog records and focused on 82 that had been updated at least monthly.
    • On 28 October 2025, 44 of those 82 databases were current and 38 were paused.
    • Most paused databases had no data entries within the previous 6 months, and only 1 paused database had been updated by 2 December 2025.
    • Most paused databases were vaccination-related; others mainly concerned respiratory diseases or drug overdose deaths.
    • The authors say unexplained pauses may weaken evidence used for health decision making and public trust.
  • Protocol compares two systems for Belgian ILI surveillance

    What the study found

    The article is a study protocol, so it does not report final findings. It describes a planned comparison of two surveillance systems for influenza-like illness (ILI, a flu-like syndrome) in Belgian general practices.

    Why the authors say this matters

    The authors say the work could help identify the most suitable alternative for effective and long-term ILI surveillance. They also conclude that the protocol could serve as a basis for validating other syndromic surveillance data from extraction-based systems in primary care.

    What the researchers tested

    The researchers are carrying out an observational retrospective study covering three influenza seasons from 2021 to 2024. They are comparing the code-based COVID-19 Barometer in General Practices, which extracts data from electronic medical records, with the questionnaire-based Belgian Sentinel General Practitioners network.

    What worked and what didn't

    The protocol says both qualitative and quantitative measures will be used to assess nine attributes: data quality, ILI incidence, sensitivity, representativeness, timeliness, acceptability, simplicity, stability, and flexibility. The study will use CDC surveillance evaluation guidelines and the Simple Multi-Attribute Rating Technique, with experts scoring and weighting three alternatives; the alternative with the higher endorsement will be considered preferable.

    What to keep in mind

    No final results are reported in the abstract because this is a protocol. The abstract does not describe limitations beyond the fact that the comparison is being planned and evaluated across the specified influenza seasons.

    • The article is a protocol, not a results paper.
    • It compares a code-based electronic medical record surveillance tool with a questionnaire-based sentinel network.
    • The focus is influenza-like illness surveillance in Belgian general practices.
    • The study uses nine evaluation attributes, including timeliness, sensitivity, and data quality.
    • Experts will score and weight three alternatives using a multi-criteria decision method.
  • Review finds machine learning may strengthen U.S. infectious disease surveillance

    What the study found

    The review finds that combining big data with machine learning may improve infectious disease surveillance and control in the U.S. The abstract describes potential gains in timeliness, accuracy, and robustness.

    Why the authors say this matters

    The authors suggest this matters because U.S. infectious disease surveillance has faced delayed feedback, inefficient data infrastructure, and limited predictive capacity. The study suggests that machine learning-enabled disease control could improve the accuracy, speed, and robustness of infectious disease control in the U.S.

    What the researchers tested

    This is a narrative review, meaning the authors synthesized current literature rather than running a new experiment. They reviewed conventional public health data sources and newer digital, genomic, and non-conventional sources, along with machine learning approaches such as supervised learning, unsupervised learning, and deep learning.

    What worked and what didn't

    The review presents practical applications of machine learning for early outbreak warning, disease control, resource allocation, and precision medicine for public health. It also presents these methods in relation to detection, forecasting, and risk assessment. The abstract does not report comparative test results for specific methods.

    What to keep in mind

    The summary provided does not describe study limitations in detail. Because this is a review, the abstract does not state that the authors conducted new data collection or direct performance testing.

    • The review argues that big data and machine learning may improve U.S. infectious disease surveillance and control.
    • It highlights delayed feedback, inefficient data infrastructure, and limited predictive capacity as existing challenges.
    • The review covers electronic health records, syndromic surveillance, mobility datasets, social media data, wearable biosensing, and genomic pathogen sequencing.
    • It discusses supervised learning, unsupervised learning, and deep learning for detection, forecasting, and risk assessment.
    • The abstract mentions applications in early warning, disease control, resource allocation, and precision medicine for public health.