0
Article ? AI-assigned paper type based on the abstract. Classification may not be perfect — flag errors using the feedback button. Tier 2 ? Original research — experimental, observational, or case-control study. Direct primary evidence. Sign in to save

Cross-study shift defines the reliability limits of FTIR polymer classification

ChemRxiv 2026
M. Maksuda Khanam, Saleena Younus, M. Khabir Uddin, Julhash U. Kazi

Summary

Scientists use AI tools to identify types of microplastic pollution from their chemical "fingerprints," but this study found these tools often perform much worse when tested on samples from a new lab than the impressive accuracy numbers usually reported would suggest. This matters because if we're relying on flawed tools to track which plastics are contaminating our water, food, and environment, we might be getting an inaccurate picture of our real-world exposure. The researchers suggest a more honest way to test and use these tools going forward, including flagging results the AI isn't confident about rather than guessing.

Automated Fourier-transform infrared (FTIR) classifiers for microplastics are often evaluated by randomly splitting spectra pooled from several studies, although deployment requires transfer to new sources. We harmonized 4,783 labelled spectra from four independent studies and compared a study-mixed random split with leave-one-study-out (LOSO) evaluation. We then examined task composition, source-associated structure, nested deep and domain-oriented models, external reference-library pooling, and selective prediction on locked domains. XGBoost achieved random-split macro-F1 of 0.972 but equal-study present-class LOSO macro-F1 of 0.672. The gap was class- and source-dependent: PE/PP transfer remained strong, whereas adding PET and PS reopened most of the shared-fold deficit despite adequate class support. Polymer class dominated the global geometry of the harmonized data, yet source study remained highly predictable, including within PE and PP. Position-aware convolutional, self-supervised and domain-adversarial configurations did not consistently outperform XGBoost, and external libraries produced non-monotonic effects, indicating that directional compatibility mattered more than library size. A maximum-probability threshold selected only from inner study holdouts identified lower-risk subsets on locked De Frond and Open Specy domains, although coverage was domain-dependent. These results support a practical trust framework for multi-source FTIR classification: validate by study, report fold- and support-aware metrics, screen external libraries for compatibility, and defer spectra outside the model's reliable operating region.

Share this paper