roboflow/maestro
streamline the fine-tuning process for multimodal models: PaliGemma 2, Florence-2, and Qwen2.5-VL observed · 2026-08-28
Health v2 · maintenance only
62/100
- Activity 99
- Release rhythm 8
- Longevity 72
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1013
- days_rel: 574
- days_push: 9
- n_releases_24m: 1
Adoption not part of the score
2694 stars · 222 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
maestro is a Python library from Roboflow that streamlines fine-tuning of multimodal vision-language models such as Florence-2, PaliGemma 2, and Qwen2.5-VL. It encapsulates configuration, data loading, reproducibility, and training loop setup into ready-to-use recipes with LoRA/QLoRA support.
Use cases
- fine-tune a vision-language model for object detection
- train Florence-2 on custom captioning data
- fine-tune Qwen2.5-VL to extract JSON from images
- adapt PaliGemma 2 for visual question answering
- run LoRA fine-tuning of VLMs on a free Colab GPU
- fine-tune a multimodal model without writing training loop code
When to choose
- you want a simple, recipe-driven way to fine-tune supported VLMs like Florence-2, PaliGemma 2, or Qwen2.5-VL
- you need LoRA/QLoRA fine-tuning for vision-language tasks like detection, captioning, VQA, or JSON extraction
- you prefer best-practice defaults over hand-rolling a training pipeline
When to avoid
- you need to fine-tune models outside the supported list (Florence-2, PaliGemma 2, Qwen2.5-VL, Phi-3-vision)
- you require full control over every aspect of the training loop or custom architectures
- you need text-only LLM fine-tuning with no vision component
Facets
library · maturity active
llm-training machine-learning deep-learning image-processing cli machine-learning deep-learning computer-vision artificial-intelligence large-language-models python cli fine-tuning vision-language-models lora qlora paligemma florence-2 qwen2-vl multimodal object-detection vqa image-captioning transformers gpu
2 sources
- readme: https://github.com/roboflow/maestro · fetched 2026-08-28 · 69d87507f899
- homepage: https://maestro.roboflow.com · fetched 2026-08-29 · 1e06227ac77e
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| roboflow/maestro | main | 62 |
For agents
markdown · JSON · MCP: product_card(name="roboflow/maestro")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem