PuSH - Publication Server of Helmholtz Zentrum München: Publishing neural networks in drug discovery might compromise training data privacy.

Navigation

Home

Deutsch

Research

Advanced Search

Browse by ...

... Journal

... Publication Type

... Research Data

... Publication Year

Publication overview

Support & Contact

Contact persons

Help

Data protection

Krüger, F. ; Östman, J.* ; Mervin, L.* ; Tetko, I.V. ; Engkvist, O.*

Publishing neural networks in drug discovery might compromise training data privacy.

J. Cheminformatics 17:38 (2025)

Publ. Version/Full Text

DOI

PMC

	Open Access Gold

Abstract
Metrics
Extra information

This study investigates the risks of exposing confidential chemical structures when machine learning models trained on these structures are made publicly available. We use membership inference attacks, a common method to assess privacy that is largely unexplored in the context of drug discovery, to examine neural networks for molecular property prediction in a black-box setting. Our results reveal significant privacy risks across all evaluated datasets and neural network architectures. Combining multiple attacks increases these risks. Molecules from minority classes, often the most valuable in drug discovery, are particularly vulnerable. We also found that representing molecules as graphs and using message-passing neural networks may mitigate these risks. We provide a framework to assess privacy risks of classification models and molecular representations, available at https://github.com/FabianKruger/molprivacy . Our findings highlight the need for careful consideration when sharing neural networks trained on proprietary chemical structures, informing organisations and researchers about the trade-offs between data confidentiality and model openness.

Altmetric

Additional Metrics?

[➜Log in]

Edit extra informations Login

Publication type Article: Journal article

Document type Scientific Article

Keywords Cheminformatics ; Drug Discovery ; Machine Learning ; Membership Inference Attack ; Privacy ; Qsar

e-ISSN 1758-2946

Journal Journal of Cheminformatics

Quellenangaben Volume: 17, Issue: 1, Article Number: 38

Publisher BioMed Central

Publishing Place Campus, 4 Crinan St, London N1 9xw, England

Reviewing status Peer reviewed

Institute(s) Institute of Structural Biology (STB)

Grants HORIZON EUROPE European Research Council