Setting the Standard: Recommended Practices for Data Preprocessing in Data-Driven Climate Prediction

dc.contributor.authorFurtado, Jason
dc.contributor.authorMolina, Maria
dc.contributor.authorArcodia, Marybeth
dc.contributor.authorAnderson, Weston
dc.contributor.authorBeucler, Tom
dc.contributor.authorCallahan, John
dc.contributor.authorCiasto, Laura
dc.contributor.authorGensini, Vittorio
dc.contributor.authorL’Heureux, Michelle
dc.contributor.authorPegion, Kathleen
dc.contributor.authorPérez-Carrasquilla, Jhayron
dc.contributor.authorSonnewald, Maike
dc.contributor.authorTakahashi, Ken
dc.contributor.authorXiang, Baoqiang
dc.contributor.authorZimmerman, Brian G.
dc.date.accessioned2026-09-16T20:31:03Z
dc.date.available2026-09-16T20:31:03Z
dc.date.issued2026-06-01
dc.description.abstractArtificial intelligence (AI)—and specifically machine learning (ML)—applications for climate prediction across time scales are proliferating quickly. The emergence of these methods prompts a revisit to the impact of data preprocessing, a topic familiar to the climate community, as more traditional statistical models work with relatively small sample sizes. Indeed, the skill and confidence in the forecasts produced by data-driven models are directly influenced by the quality of the datasets and how they are treated during model development, yielding the colloquialism, “garbage in, garbage out.” As such, this article establishes protocols for the proper preprocessing of input data for AI/ML models designed for climate prediction (i.e., subseasonal-to-decadal and longer time scales). The three aims are to 1) educate researchers, developers, and end users on the effects that data preprocessing has on climate prediction; 2) provide recommended practices for data preprocessing for such applications; and 3) empower end users to decipher whether the models they are using are properly designed for their objectives. Specific topics covered include creating (standardized) anomalies, dealing with nonstationarity and the spatiotemporally correlated nature of climate data, and handling of extreme values and variables with potentially complex distributions. Case studies will illustrate how using different preprocessing techniques can produce different predictions from the same model, which can create confusion and decrease confidence in the overall process. Ultimately, implementing the recommended practices set forth in this article will enhance the robustness and transparency of AI/ML in climate prediction studies.
dc.description.peer-reviewPor pares
dc.description.sponsorshipNational Science Foundation (NSF), Grant 2425735
dc.formatapplication/pdf
dc.identifier.citationFurtado, J. C., Molina, M. J., Arcodia, M. C., Anderson, W., Beucler, T., Callahan, J. A., Ciasto, L. M., Gensini, V. A., L’Heureux, M., Pegion, K., Pérez-Carrasquilla, J. S., Sonnewald, M., Takahashi, K., Xiang, B., & Zimmerman, B. G. (2026). Setting the standard: Recommended practices for data preprocessing in data-driven climate prediction.==$Bulletin of the American Meteorological Society, 107$==(6), E1386–E1401. https://doi.org/10.1175/BAMS-D-24-0292.1
dc.identifier.doihttps://doi.org/10.1175/BAMS-D-24-0292.1
dc.identifier.govdocindex-oti2018
dc.identifier.journalBulletin of the American Meteorological Society
dc.identifier.urihttps://hdl.handle.net/20.500.12816/5876
dc.language.isoeng
dc.publisherAmerican Meteorological Society
dc.rightshttp://purl.org/coar/access_right/c_f1cf
dc.subjectStatistical techniques
dc.subjectClimate prediction
dc.subjectArtificial intelligence
dc.subjectMachine learning
dc.subject.ocdehttps://purl.org/pe-repo/ocde/ford#1.05.10
dc.titleSetting the Standard: Recommended Practices for Data Preprocessing in Data-Driven Climate Prediction
dc.typehttp://purl.org/coar/resource_type/c_6501
dc.type.versionhttp://purl.org/coar/version/c_970fb48d4fbd8a85

Archivos

Bloque original

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
Furtado_et_al_2026_Bulletin of the American Meteorological Society.pdf
Tamaño:
2.22 MB
Formato:
Adobe Portable Document Format

Bloque de licencias

Mostrando 1 - 1 de 1
Cargando...
Miniatura
Nombre:
license.txt
Tamaño:
1.71 KB
Formato:
Item-specific license agreed upon to submission
Descripción:

Colecciones