Ross ROSS = Recommend OSS · open-source software intelligence for agents

OFA-Sys/OFA

Official repository of OFA (ICML 2022). Paper: OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework observed · 2026-08-28

github.com/OFA-Sys/OFA · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

32/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1677
  • days_rel: n/a
  • days_push: 861
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2557 stars · 248 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

OFA is a unified sequence-to-sequence pretrained model supporting English and Chinese that unifies cross-modality, vision, and language tasks through a single framework. The repository provides training, finetuning, and prompt-tuning code plus checkpoints for tasks like image captioning, VQA, visual grounding, and text-to-image generation.

Use cases

  • generate image captions with a pretrained vision-language model
  • run visual question answering on images
  • perform visual grounding on referring expressions
  • generate images from text descriptions
  • finetune a unified multimodal seq2seq model
  • do prompt tuning on a pretrained multimodal model
  • classify images and text with one unified model

When to choose

  • you need a single pretrained model covering multiple vision-language tasks
  • you want strong image captioning or VQA baselines with leaderboard-level checkpoints
  • you need both English and Chinese multimodal support
  • you want to experiment with prompt tuning on multimodal models

When to avoid

  • you need a lightweight production inference server rather than research code
  • you want the latest large multimodal LLMs rather than a 2022-era seq2seq model
  • you need tasks outside the supported vision-language set
  • you require active development or frequent updates

Facets

library · maturity maintenance

machine-learning deep-learning nlp image-processing llm-training prompt-engineering artificial-intelligence machine-learning computer-vision deep-learning python multimodal vision-language pretrained-models sequence-to-sequence image-captioning visual-question-answering text-to-image prompt-tuning research-code natural-language-processing gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
OFA-Sys/OFAmain32

For agents

markdown · JSON · MCP: product_card(name="OFA-Sys/OFA")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem