Environment International | |
Prioritizing cancer hazard assessments for IARC Monographs using an integrated approach of database fusion and text mining | |
Mary K. Schubauer-Berigan1  Dinesh Kumar Barupal2  Kathryn Z. Guyton3  Jiri Zavadil3  Michael Korenjak4  | |
[1] Corresponding authors at: CAM Building 17 E 102nd St Third Floor, Icahn School of Medicine at Mt Sinai, New York, NY 10029, USA (D.K. Barupal).;Department of Environmental Medicine and Public Health, Icahn School of Medicine at Mt Sinai, NY, USA;Epigenomics and Mechanisms Branch, International Agency for Research on Cancer, Lyon, France;Evidence Synthesis and Classification Branch, International Agency for Research on Cancer, Lyon, France; | |
关键词: IARC Monographs; Text mining; Hazard identification; Database fusion; Chemoinformatics; | |
DOI : | |
来源: DOAJ |
【 摘 要 】
Background: Systematic evaluation of literature data on the cancer hazards of human exposures is an essential process underlying cancer prevention strategies. The scope and volume of evidence for suspected carcinogens can range from very few to thousands of publications, requiring a complex, systematically planned, and critical procedure to nominate, prioritize and evaluate carcinogenic agents. To aid in this process, database fusion, cheminformatics and text mining techniques can be combined into an integrated approach to inform agent prioritization, selection, and grouping. Results: We have applied these techniques to agents recommended for the IARC Monographs evaluations during 2020–2024. An integration of PubMed filters to cover cancer epidemiology, key characteristics of carcinogens, chemical lists from 34 databases relevant for cancer research, chemical structure grouping and a literature data-based clustering was applied in an innovative approach to 119 agents recommended by an advisory group for future IARC Monographs evaluations. The approach also facilitated a rational grouping of these agents and aids in understanding the volume and complexity of relevant information, as well as important gaps in coverage of the available studies on cancer etiology and carcinogenesis. Conclusion: A new data-science approach has been applied to diverse agents recommended for cancer hazard assessments, and its applications for the IARC Monographs are demonstrated. The prioritization approach has been made available at www.cancer.idsl.me site for ranking cancer agents.
【 授权许可】
Unknown