# jingyi0000/VLM_survey

Collection of AWESOME vision-language models for vision tasks

Repository: https://github.com/jingyi0000/VLM_survey
Canonical: https://ross.abutalabs.com/products/vlm_survey
License Family: other
Topics: computer-vision, deep-learning, knowledge-distillation, survey, transfer-learning, vision-language-model, clip, multi-modal-model
Last push: 2025-10-14T09:40:40+00:00

## Health v2 (maintenance only)
Score: 51/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 47, release rhythm 35, longevity 89
- inputs: {"age_days": 1252, "days_push": 323, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3128, forks 234 (observed 2026-08-28T04:07:45.323055+00:00)

## What it is
A curated awesome-list repository accompanying the TPAMI 2024 survey 'Vision-Language Models for Vision Tasks', cataloging papers on VLMs for image classification, object detection, segmentation, and other vision tasks. It is regularly updated with new papers, code links, and related collections.

## Use cases
- find papers on vision-language models for image classification
- survey of CLIP-based object detection and segmentation methods
- keep up with recent VLM research for vision tasks
- find code implementations of multimodal vision papers
- literature review for a thesis on vision-language models
- discover VLM transfer learning and knowledge distillation papers

## When to choose
- you need a curated, categorized reading list of VLM papers with code links
- you want a peer-reviewed survey (TPAMI) as a starting point for research
- you want an actively maintained list updated through 2025

## When to avoid
- you need runnable software or a library rather than a paper collection
- you need exhaustive coverage of every VLM paper rather than a curated selection
- you need a license-protected dataset or models

## Facets
- artifact type: learning-resource
- maturity: active
- function: nlp, computer-vision, machine-learning
- domain: computer-vision, deep-learning, tutorials
- platform: cross-platform
- tags: awesome-list, vision-language-models, survey, clip, multimodal, knowledge-distillation, transfer-learning, research-papers, natural-language-processing

## Member repositories
- jingyi0000/VLM_survey (main) score 51

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:45.323055+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:26:20.440328+00:00, confidence not recorded.
  - readme: https://github.com/jingyi0000/VLM_survey (fetched 2026-08-28T04:07:45.323055+00:00, sha 219070fbdf57)
- Data as of 2026-08-30T08:39:29.467469+00:00.
