Logo

Optimize development process of new drug candidates

May 7, 2024/2 min read

Query public chemical databases

Explore previously studied and novel molecular structures (SMILES) for drug candidates. Easily query public chemical databases with ChEMBL and PubChem APIs. Build the solution flow by selecting a preferred chemical database, then query by target protein and specify feature generation and modeling options for bioactivity prediction.

Explore the project

Query Public Chemical Databases

Reuse and extend coding functions

Power recipes in the flow using Python with open source chemoinformatics libraries. Create molecular descriptors and fingerprint features from SMILES chemical structures with RDKit and pre-trained transformer language models like ChemBERTA with HuggingFace.

Learn more about extensibility with dataiku

Reuse and Extend Coding Functions

Cluster and predict bioactivity

Gain deeper understanding of the molecular structure space with dimension reduction methods like t-SNE, PCA, and clustering, then train ML models to predict the bioactivity of small molecules for a given protein target.

See the solution in action

Cluster and Predict Bioactivity

Deeply explore molecular properties

Use dynamic, interactive dashboards to understand previously studied and novel molecules. Understand chemical properties and molecular descriptors with interactive visualizations, analyses, and tables. Filter to review individual molecule summaries, and screen leading candidates in-silico through the Dataiku App.

Go to the chemical space analysis dashboard

Deeply Explore Molecular Properties

Visualize development candidates

Assess novel small molecules and gain valuable insights by computing and visualizing molecular descriptors and fingerprints. Use these insights and the trained bioactivity prediction model to score new molecules, and find similar molecules in previous experiments.

Explore the molecular similarity dashboard

Visualize Development Candidates

Accelerate molecular prediction

Gain full flexibility by adjusting to specific research fields with a composable and extendable solution. Adapt and expand research to gain new insights and make progressively better molecular predictions.

Explore the dataiku flow

Accelerate Molecular Prediction

Answer key molecular research questions

The Dataiku Solution for Molecular Property Prediction helps answer a broad range of questions like: How can I reduce cost and time in the drug discovery and development process? How can I use AI and ML to research molecules and their bioactivity for a given protein target? Can I use ML to help prioritize hit-to-lead development and experimentation?

Read the documentation on molecular property prediction

Answer Key Molecular Research Questions

Optimize molecular screening before experimental work

Improve the success and efficiency of the drug development cycle by prioritizing lead development candidate experimentation. Accelerate the process of predicting target protein bioactivity in novel small molecules. Leverage AI to build a pipeline of potential drug solutions to help bring a stable compound quickly to preclinical and clinical testing.

Discover the full capabilities of dataiku

Request a demo from a Dataiku industry expert

Ready for AI success?