Ross ROSS = Recommend OSS · open-source software intelligence for agents

jingyi0000/VLM_survey resource

Collection of AWESOME vision-language models for vision tasks observed · 2026-08-28

github.com/jingyi0000/VLM_survey observed · 2026-08-28

Health v2 · maintenance only

51/100

  • Activity 47
  • Release rhythm 35
  • Longevity 89

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1252
  • days_rel: n/a
  • days_push: 323
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

3128 stars · 234 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A curated awesome-list repository accompanying the TPAMI 2024 survey 'Vision-Language Models for Vision Tasks', cataloging papers on VLMs for image classification, object detection, segmentation, and other vision tasks. It is regularly updated with new papers, code links, and related collections.

Use cases

  • find papers on vision-language models for image classification
  • survey of CLIP-based object detection and segmentation methods
  • keep up with recent VLM research for vision tasks
  • find code implementations of multimodal vision papers
  • literature review for a thesis on vision-language models
  • discover VLM transfer learning and knowledge distillation papers

When to choose

  • you need a curated, categorized reading list of VLM papers with code links
  • you want a peer-reviewed survey (TPAMI) as a starting point for research
  • you want an actively maintained list updated through 2025

When to avoid

  • you need runnable software or a library rather than a paper collection
  • you need exhaustive coverage of every VLM paper rather than a curated selection
  • you need a license-protected dataset or models

Facets

learning-resource · maturity active

nlp computer-vision machine-learning computer-vision deep-learning tutorials cross-platform awesome-list vision-language-models survey clip multimodal knowledge-distillation transfer-learning research-papers natural-language-processing

1 source

Member repositories

RepositoryRoleHealth v2
jingyi0000/VLM_surveymain51

For agents

markdown · JSON · MCP: product_card(name="jingyi0000/VLM_survey")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem