# apple/ml-ferret

Repository: https://github.com/apple/ml-ferret
Canonical: https://ross.abutalabs.com/products/ml-ferret
Language: Python
License: NOASSERTION
License Family: other
Last push: 2024-10-09T02:40:44+00:00

## Health v2 (maintenance only)
Score: 27/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 75
- inputs: {"age_days": 1062, "days_push": 693, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 8674, forks 518 (observed 2026-08-28T04:10:24.170350+00:00)

## What it is
Apple's Ferret, an end-to-end multimodal large language model (MLLM) that accepts any-form referring and grounds anything in its responses, with fine-grained open-vocabulary referring and grounding via hybrid region representation. Includes Ferret-UI, a UI-centric MLLM for referring, grounding, and reasoning on mobile interfaces, plus training/evaluation code and checkpoints.

## Use cases
- ground objects in images with natural language referring
- build a multimodal chatbot that understands image regions
- run referring and grounding reasoning on mobile UI screenshots
- fine-tune a vision-language model with the GRIT instruction dataset
- evaluate MLLMs on Ferret-Bench multimodal benchmarks

## When to choose
- you need fine-grained region-level referring and grounding in an MLLM
- you want a UI-understanding model for app screenshots
- you're doing research on multimodal instruction tuning

## When to avoid
- you need a production-ready commercial product (research-only license)
- you just need a general-purpose VLM without grounding capabilities
- you lack GPU resources for 7B/13B model inference

## Facets
- artifact type: library
- maturity: maintenance
- function: machine-learning, deep-learning, nlp, computer-vision, llm-inference, llm-training
- domain: large-language-models, computer-vision, artificial-intelligence, image-processing
- platform: python
- tags: multimodal, vision-language-model, visual-grounding, referring-expression, research-code, ferret-ui, instruction-tuning, natural-language-processing, linux, gpu

## Member repositories
- apple/ml-ferret (main) score 27

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:10:24.170350+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:25:36.691974+00:00, confidence not recorded.
  - readme: https://github.com/apple/ml-ferret (fetched 2026-08-28T04:10:24.170350+00:00, sha 954a800ef268)
- Data as of 2026-08-30T08:39:29.467469+00:00.
