Ross ROSS = Recommend OSS · open-source software intelligence for agents

zhaochen0110/Awesome_Think_With_Images resource

Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation. observed · 2026-08-28

github.com/zhaochen0110/Awesome_Think_With_Images · homepage observed · 2026-08-28

Health v2 · maintenance only

51/100

  • Activity 71
  • Release rhythm 35
  • Longevity 32

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 460
  • days_rel: n/a
  • days_push: 177
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1503 stars · 47 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A curated awesome-list repository accompanying a survey paper on 'Thinking with Images' for large vision-language models (LVLMs). It organizes research into three stages of cognitive autonomy: tool-driven visual exploration, programmatic visual manipulation, and intrinsic visual imagination.

Use cases

  • find papers on multimodal reasoning with images
  • research how LVLMs use visual tools for reasoning
  • survey visual reasoning methods for vision-language models
  • get started with thinking-with-images research
  • find resources on programmatic visual manipulation
  • track state-of-the-art in intrinsic visual imagination
  • prepare a literature review on multimodal AI reasoning

When to choose

  • you need a comprehensive, structured paper list on visual reasoning in LVLMs
  • you are starting research on multimodal reasoning and want curated foundations
  • you want to follow the three-stage taxonomy of tool use, visual programming, and visual imagination

When to avoid

  • you need runnable code or a software library rather than a paper list
  • you want a general multimodal AI resource not focused on visual reasoning
  • you need production tooling for vision-language model deployment

Facets

learning-resource · maturity active

nlp computer-vision machine-learning artificial-intelligence computer-vision awesome-lists tutorials awesome-list survey-paper multimodal-reasoning vision-language-models research-curation thinking-with-images natural-language-processing web-server

1 source

Member repositories

RepositoryRoleHealth v2
zhaochen0110/Awesome_Think_With_Imagesmain51

For agents

markdown · JSON · MCP: product_card(name="zhaochen0110/Awesome_Think_With_Images")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem