# agemagician/ProtTrans

ProtTrans is providing state of the art pretrained language models for proteins. ProtTrans was trained on thousands of GPUs from Summit and hundreds of Google TPUs using Transformers Models.

Repository: https://github.com/agemagician/ProtTrans
Canonical: https://ross.abutalabs.com/products/prottrans
Language: Jupyter Notebook
License: MIT
License Family: permissive
Last push: 2025-05-22T08:47:16+00:00

## Health v2 (maintenance only)
Score: 33/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 22, release rhythm 8, longevity 100
- inputs: {"age_days": 2305, "days_push": 468, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1324, forks 169 (observed 2026-08-28T04:04:22.316927+00:00)

## What it is
ProtTrans provides state-of-the-art pre-trained Transformer language models for protein sequences, trained on thousands of GPUs and hundreds of TPUs. It includes notebooks and code for feature extraction, fine-tuning, prediction, and protein sequence generation.

## Use cases
- extract embeddings from protein sequences
- fine-tune a protein language model on downstream tasks
- predict protein secondary structure or per-residue properties
- generate novel protein sequences
- visualize transformer attention on proteins
- benchmark protein language models

## When to choose
- you need pretrained protein language models like ProtT5 or ProtBERT
- you want embeddings for protein sequences for downstream ML
- you are doing bioinformatics or Covid-19 related protein research

## When to avoid
- you need a production-ready web service rather than models and notebooks
- you work outside protein/bioinformatics domains
- you lack GPU/TPU resources for large transformer models

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, nlp, sdk
- domain: bioinformatics, deep-learning, healthcare
- platform: python, cross-platform
- tags: protein-language-models, transformers, pretrained-models, transfer-learning, embeddings, bioinformatics, protein-sequences, natural-language-processing, gpu

## Member repositories
- agemagician/ProtTrans (main) score 33

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:22.316927+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:46:41.040221+00:00, confidence not recorded.
  - readme: https://github.com/agemagician/ProtTrans (fetched 2026-08-28T04:04:22.316927+00:00, sha e42a8714fa38)
- Data as of 2026-08-30T08:39:29.467469+00:00.
