Seven toxicity endpoints.
One integrated risk score.
MultiEndpointTox is an AI-powered, SHAP-explainable platform that predicts hERG, hepatotoxicity, nephrotoxicity, Ames mutagenicity, skin sensitization, cytotoxicity, and reproductive toxicity from a single SMILES string — with confidence-weighted integrated risk assessment for early-stage drug discovery decisions.
01 · Platform
PLATFORM OVERVIEW
MultiEndpointTox is an AI-powered platform that predicts 7 key toxicity endpoints and provides integrated risk assessment to support early-stage decision-making in drug discovery.
What is MultiEndpointTox?
An integrated, interpretable, and validated AI platform that predicts multiple toxicity endpoints and delivers a comprehensive risk assessment in a single solution.
AI-Powered Predictions
State-of-the-art machine learning models trained on high-quality multimodal data.
Integrated Risk Assessment
Combine multiple predictions into a robust, confidence-weighted risk score.
Interpretable Insights
Explainable AI and visual analytics reveal key molecular drivers behind each prediction.
Decision Confidence
Transparent, consistent, and reliable outputs to support safer and faster decisions.
7 Toxicity Endpoints
Comprehensive coverage of key safety liabilities.
Advanced AI/ML Models
Multi-task learning, domain adaptation, and ensemble methods.
Integrated Risk Score
Population-weighted, confidence-based risk for better prioritization.
Explainable & Transparent
SHAP-based explanations and intuitive visual analytics.
Validated & Reliable
Rigorous validation across datasets, scaffolds, and external benchmarks.
Built for Discovery
Designed for early-stage screening and lead optimization.
Save Time
Reduce experimental burden with fast, accurate predictions.
Reduce Costs
Prioritize safer compounds and avoid late-stage failures.
Improve Safety
Identify liabilities early and design safer molecules.
Boost Success Rate
Make confident, data-driven decisions at every step.
Accelerate Innovation
Enable smarter discovery with trusted AI insights.
02 · Endpoints
7 KEY TOXICITY ENDPOINTS
Comprehensive prediction of seven critical toxicity liabilities to support safer compounds and better decision-making. Hover any card for its example risk profile.
Hepatotoxicity
classification
Predicts the potential of compounds to cause liver injury or dysfunction.
Decision threshold (0.717) deliberately favors sensitivity over specificity (23.8%) — false negatives are costlier than false positives in early safety screening.
Nephrotoxicity
classification
Assesses the likelihood of kidney damage or impaired renal function.
Ames Mutagenicity
classification
Evaluates the potential of compounds to cause genetic mutations.
Skin Sensitization
classification
Predicts the potential of a compound to trigger an allergic skin reaction.
Only ~31% applicability-domain coverage — most novel-compound predictions fall outside the training distribution and should be treated as exploratory.
Reproductive Toxicity
classification
Estimates the risk of adverse effects on fertility and reproductive health.
Underpowered (n=127) — scaffold AUC=0.588, not significantly better than random on novel scaffolds. Reported for information only and excluded from the integrated risk score.
hERG Cardiotoxicity
regression
Predicts the potential to inhibit hERG potassium channels (pIC50 regression), which may lead to cardiac arrhythmias.
Scaffold-split R²=0.378 vs. 0.621 on a random split — a documented 39% generalization gap. Treat novel-scaffold predictions with caution; a graph-neural-network replacement is on the roadmap.
Cytotoxicity
classification
Assesses the potential of compounds to cause cell damage or reduce cell viability.
Comprehensive Safety Coverage
Address 7 critical toxicity liabilities in one platform.
Early Risk Identification
Detect potential liabilities at the earliest stages.
Better Prioritization
Focus resources on the most promising compounds.
Informed Decisions
Integrate toxicity insights into confident decisions.
Safer Drug Discovery
Reduce late-stage failures and improve success rates.
03 · Framework
ADVANCED AI/ML MODELS
State-of-the-art machine learning with multi-task learning and domain adaptation to deliver accurate, robust, and generalizable toxicity predictions.
Multi-Task Learning
Leverage shared molecular knowledge across 7 toxicity endpoints for more accurate and data-efficient predictions.
Shared Representation — Stronger Learning, Better Accuracy
Our AI/ML Framework
INPUT
Structure · Descriptors · Fingerprints · Graph Features
OUTPUT
7 Toxicity Predictions
Domain Adaptation
Improve model generalizability across different chemical spaces, assays, and laboratories.
- Reduces dataset bias and batch effects
- Enhances performance on unseen data
- Built for real-world applicability
High Accuracy
Scaffold-split validated performance across diverse benchmarks.
AUC up to 0.92
Across Key Endpoints
Robustness
Stable predictions across chemical space and noise.
Consistent
Across Scaffolds
Data Efficiency
Learn more with less data through inductive bias and multi-task sharing.
Multi-Task
Sharing
Scalability
Handle large datasets and complex models efficiently.
Cloud & HPC
Optimized
Continuous Learning
Models improve over time with new data and feedback.
Adaptive
Framework
Uncertainty-Aware
Quantify prediction confidence to support risk-based decisions.
Confidence
Estimates
04 · Risk Assessment
INTEGRATED RISK ASSESSMENT
From single-endpoint predictions to an integrated, confidence-weighted risk score for smarter, data-driven decision-making.
1. Endpoint Predictions
AI/ML models generate probability of liability for each endpoint.
2. Weighting & Integration
Population weights + confidence adjustment + risk aggregation.
3. Integrated Risk Score
Final score between 0 and 1 indicates overall risk level.
Endpoint Risk Matrix (Example)
| hERG | Hepato | Ames | Nephro | Skin Sens | Cyto | Repro Tox | |
|---|---|---|---|---|---|---|---|
| Drug A | 0.17 | 0.91 | 0.03 | 0.18 | 0.04 | 0.20 | 0.58 |
| Drug B | 0.23 | 0.64 | 0.05 | 0.82 | 0.97 | 0.10 | 0.58 |
| Drug C | 0.23 | 0.86 | 0.00 | 0.70 | 0.01 | 0.12 | 0.58 |
| Drug D | 0.20 | 0.93 | 0.02 | 0.67 | 0.97 | 0.18 | 0.58 |
| Drug E | 0.05 | 0.95 | 0.02 | 0.68 | 0.00 | 0.09 | 0.58 |
Integrated Risk Score
95% CI: 0.48 – 0.60
Population Weighting
Endpoints are weighted based on population-level prevalence and clinical relevance.
Risk Distribution (Example Dataset)
Risk Weighting Scenarios
The API accepts a scenario parameter that reprioritizes which endpoint drives the integrated score — useful when one liability matters more for a given program.
- BaselineManuscript Table 4 weights (default).
- EqualAll endpoint weights set to 1.0.
- Cardiac priorityhERG weight raised to 3.0.1.5 → 3.0
- Hepatic priorityHepatotoxicity weight raised to 3.0.1.5 → 3.0
- Genotoxicity priorityAmes weight raised to 3.0.1.3 → 3.0
Holistic Safety View
See the big picture across all critical toxicity endpoints.
Better Prioritization
Focus on compounds with lower overall risk.
Population Awareness
Accounts for population and metabolizer variability.
Data-Driven Confidence
Confidence-weighted approach improves decision reliability.
Faster, Smarter Decisions
Reduce attrition and accelerate safer drug discovery.
05 · Explainability
INTERPRETABILITY & TRANSPARENCY
We make every prediction explainable with state-of-the-art AI techniques and intuitive visual analytics.
Build Trust
Understand why a model made a prediction.
Ensure Reliability
Detect spurious patterns and data artifacts.
Drive Discovery
Reveal key molecular drivers of toxicity.
Support Decisions
Provide scientific rationale for safer compound selection.
Regulatory Readiness
Deliver transparent and auditable predictions.
Global Feature Importance
Example: Hepatotoxicity
Local Explanation (SHAP Beeswarm)
Each dot = one compound
Local Explanation (Waterfall Plot)
How features push the prediction from the base value to the final result.
Model Explainability Summary
Molecular Highlight
Example compound — drag to rotate
06 · Validation
ROBUST VALIDATION & PERFORMANCE
Scaffold-split cross-validation, reported honestly — including where it's a harder story than a random split would tell.
Scaffold-Split GroupKFold
5-fold cross-validation with all compounds sharing a scaffold kept in the same fold.
No Leakage
SMOTE, feature selection, and scaling are fit inside each training fold only.
Temporal Split (in progress)
Methodology implemented; time-split validation on newer compounds is not yet reported.
External Benchmarks (in progress)
ClinTox and Tox21 external validation scripts exist; results are not yet published.
Train/Test Overlap Audit
Scaffold & exact-structure leakage checked between splits.
Performance Summary
Primary scaffold-split GroupKFold(k=5) AUC per classification endpoint
ROC Curves
Reconstructed from each endpoint's reported AUC (the API returns summary AUC, not raw TPR/FPR arrays) — shape is illustrative, area matches the published figure exactly.
hERG is a pIC50 regression model, so ROC/AUC doesn't apply. Scaffold-split R² = 0.38, versus R² = 0.62 on a random split — see the generalization gap chart below.
Scaffold vs. Random-Split Generalization Gap
Random splits let a model exploit scaffold-level similarity between train and test sets, inflating performance. These are the two endpoints where both numbers are published.
External Validation Status
Per the model card's TRIPOD-AI checklist: external validation is partial. Reported honestly as status, not results that don't exist yet.
| Dataset | Purpose | Status |
|---|---|---|
| ClinTox | External classification benchmark (FDA-approved vs. withdrawn-for-toxicity) | In progress |
| Tox21 | External multi-assay toxicity benchmark | In progress |
| Temporal split | Time-split validation (train on older compounds, test on newer) | Methodology implemented |
| Train/test overlap check | Scaffold & exact-structure leakage audit between splits | Implemented |
0.79 – 0.92
Classification AUC
Scaffold-split, 6 classification endpoints
0.81
Integrated Score
AUC on a withdrawn/safe compound set
R² 0.38
hERG (regression)
Scaffold-split — see generalization gap below
Excluded
Reproductive Tox
Underpowered (n=127) — informational only
07 · Workflow
END-TO-END WORKFLOW
A seamless, integrated pipeline from chemical input to actionable toxicity insights.
Input & Data Ingestion
SMILES / SDF, curated, standardized molecular data.
Molecular Representation
2D/3D descriptors, fingerprints, graph representations.
AI/ML Modeling
Multi-task learning, domain adaptation, ensemble models.
Multi-Endpoint Predictions
Uncertainty-aware predictions across 7 endpoints.
Integrated Risk Assessment
Confidence-weighted, population-adjusted risk score.
Interpretation & Explanation
Global & local SHAP explanations, mechanistic insight.
Actionable Decisions
Prioritize, de-risk, optimize, and report.
Large & Curated Toxicity Data
Diverse, high-quality datasets across endpoints.
Advanced AI/ML
State-of-the-art models with multi-task learning and domain adaptation.
Robust Validation
Rigorous internal & external validation ensures reliability.
Scalable Infrastructure
Cloud-native platform for high performance and enterprise scalability.
Secure & Compliant
Data security, privacy, and compliance by design.
Comprehensive Toxicity Profile
7-endpoint predictions in one view.
Actionable Risk Score
Confidence-weighted risk with clear interpretation.
Mechanistic Insights
Understand “why” through molecular drivers.
Data-Driven Decisions
Make informed choices faster with greater confidence.
Reports & Dashboards
Shareable, audit-ready outputs for teams and regulators.
Continuous Improvement
Models learn and improve with new data and feedback.
From Data to Discovery. From Prediction to Protection.
MultiEndpointTox empowers better science, safer compounds, and smarter decisions.
08 · Case Study
VALPROIC ACID (VPA)
Comprehensive AI-powered toxicity assessment — live prediction from the deployed MultiEndpointTox API, not a canned screenshot.
Running full 7-endpoint analysis for Valproic Acid…
Live inference + SHAP explanation on the deployed model. If the API has been idle, this includes a cold start — first requests can take up to 2–3 minutes.
CCCC(CCC)C(=O)O