An Integrated Generative AI Framework for De Novo Molecular Design, Molecular Property Prediction, and Drug Discovery
Contributors
Deepak Gupta
Keywords
Proceeding
Track
Engineering, Sciences and Mathematics
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Accessing novel chemical matter is still a hard problem in drug design because of the astronomical size of chemical space and difficulty in encoding molecular structure information. In this work, a deep learning–based molecular generation framework to automatically design novel compounds from a deep learning generative model trained on ZINC database of drug-like molecules is introduced. Valid SMILES from filtered ZINC dataset to create subset dataset was sampled , after molecular validation and preprocessing the molecules were used to train the generative model. Experimental results demonstrate high generative performance with 100% validity, 99.9% uniqueness and 100% novelty. To assess molecular physicochemical properties and drug-likeness of generated molecules various molecular descriptors were calculated on generated molecules such as molecular weight (MW), logarithm of the partition coefficient (LogP), topological polar surface area (TPSA), hydrogen bond donors (HBD), hydrogen bond acceptors (HBA) and quantitative estimate of drug-likeness (QED). Generated molecules follow good physicochemical properties with 254.57 Da mean MW, 2.06 mean LogP, 43.83 Å mean TPSA, 1.06 HBD, 2.70 HBA, and 0.743 QED score which represents high drug-likeness. Also generated molecules follow Lipinski’s rule of five with 99.4% drug likeness molecules. To analyze structural diversity of our generated molecular library we conducted chemical space visualization, similarity analysis and scaffold diversity statistics. Using PCA the generated molecules lie mostly on ZINC dataset chemical space and continue into adjacent chemical space. Calculating pairwise similarity, mean tanimoto similarity of 0.12 was obtained which indicates high diversity amongst molecules. Finally scaffold stats was calculated and observed 9932 scaffold occurrences with 7791 unique scaffolds.Top of Form