# mims-harvard/TDC

Therapeutics Commons (TDC): Multimodal Foundation for Therapeutic Science

Repository: https://github.com/mims-harvard/TDC
Canonical: https://ross.abutalabs.com/products/tdc
Homepage: https://tdcommons.ai
Language: Jupyter Notebook
License: MIT
License Family: permissive
Topics: machine-learning, therapeutics, drug-discovery, datasets, biology, chemistry, biomedicine, bioinformatics, cheminformatics, deep-learning, benchmarks, artificial-intelligence, precision-medicine, medicine, biotech
Last push: 2025-07-13T17:44:10+00:00

## Health v2 (maintenance only)
Score: 46/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 31, release rhythm 35, longevity 100
- inputs: {"age_days": 2176, "days_push": 416, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1276, forks 222 (observed 2026-08-28T04:04:13.130233+00:00)

## What it is
Therapeutics Data Commons (TDC) is an open-science initiative and Python library providing AI-ready datasets, machine learning tasks, and curated benchmarks for drug discovery and therapeutic science. It includes data processing tools, systematic evaluation strategies, data splits, and molecule generation oracles, integrated via the PyTDC library.

## Use cases
- benchmark machine learning models for drug discovery
- download AI-ready datasets for molecular property prediction
- evaluate drug-target binding affinity prediction models
- find standard data splits for biomedical ML benchmarks
- generate molecules with oracles for molecular optimization
- compare model performance on therapeutic science leaderboards
- train models on single-cell and clinical datasets

## When to choose
- you need curated, benchmarked datasets for drug discovery or therapeutics ML
- you want standardized evaluation and data splits for biomedical tasks
- you are researching which AI methods work best for therapeutic applications

## When to avoid
- you need a general-purpose cheminformatics toolkit rather than ML datasets and benchmarks
- your domain is unrelated to biology, chemistry, or medicine
- you need production drug discovery software rather than research datasets

## Facets
- artifact type: dataset
- maturity: active
- function: machine-learning, benchmarking, data-science, data-generation
- domain: bioinformatics, machine-learning, healthcare, chemistry
- platform: python, cross-platform
- tags: drug-discovery, therapeutics, benchmarks, cheminformatics, biomedicine, datasets, molecule-generation, leaderboards

## Member repositories
- mims-harvard/TDC (main) score 46

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:13.130233+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T05:02:57.723984+00:00, confidence not recorded.
  - readme: https://github.com/mims-harvard/TDC (fetched 2026-08-28T04:04:13.130233+00:00, sha e0ddc7dfdbbb)
  - homepage: https://tdcommons.ai (fetched 2026-08-29T12:14:04.962808+00:00, sha fc489e249214)
- Data as of 2026-08-30T08:39:29.467469+00:00.
