function: llm-training
818 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| mlrun/mlrun MLRun is an open-source MLOps and AI orchestration framework for building, training, deploying, and monitoring machine learning and generat… | 96 | 1692 | active |
| scverse/scvi-tools scvi-tools is a Python library for deep probabilistic modeling and analysis of single-cell and spatial omics data, built on PyTorch, PyTorc… | 93 | 1681 | stable |
| JIA-Lab-research/ControlNeXt ControlNeXt is the official implementation of a controllable generation method for images and videos, built on Stable Diffusion XL, Stable … | 24 | 1646 | active |
| amazon-far/holosoma Holosoma is a Python framework for training and deploying reinforcement learning policies on humanoid robots, supporting locomotion and who… | 61 | 1616 | active |
| SonyResearch/micro_diffusion Official implementation of Sony Research's micro-budget approach to training large-scale text-to-image diffusion transformer models from sc… | 24 | 1594 | active |
| MoonshotAI/Kimi-Linear Kimi Linear is a hybrid linear attention architecture (Kimi Delta Attention, based on Gated DeltaNet) released by Moonshot AI with 48B-para… | 40 | 1588 | active |
| FoundationVision/Infinity Infinity is a bitwise autoregressive text-to-image generation model (CVPR 2025 Oral) with released training and inference code, checkpoints… | 56 | 1587 | active |
| baaivision/Emu3.5 Emu3.5 is BAAI's native multimodal foundation model that jointly predicts next states across vision and language, trained on 10T+ interleav… | 43 | 1547 | active |
| lucidrains/DALLE-pytorch A PyTorch implementation/replication of OpenAI's DALL-E, a text-to-image transformer, including a discrete VAE and optional CLIP for rankin… | 23 | 5627 | maintenance |
| ByteDance-Seed/Triton-distributed Triton-distributed is a distributed compiler built on OpenAI Triton for computation-communication overlapping on multi-GPU systems. It lets… | 63 | 1526 | active |
| Tencent/TFace TFace is a research platform from Tencent Youtu Lab for trusty face analysis, covering face recognition, face security (anti-spoofing), fac… | 57 | 1521 | active |
| microsoft/Mage Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage… | 57 | 1516 | active |
| bytedance/SALMONN SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video… | 73 | 1513 | active |
| ZFTurbo/Music-Source-Separation-Training A Python training framework for music source separation models, supporting many architectures such as MDX23C, Demucs, Band Split RoFormer, … | 83 | 1512 | active |
| wb14123/seq2seq-couplet A deep learning project that generates Chinese couplets (对联) using a seq2seq model built with TensorFlow. It includes training scripts, a w… | 32 | 5487 | maintenance |
| TencentARC/MotionCtrl MotionCtrl is the official implementation of a SIGGRAPH 2024 paper providing a unified and flexible motion controller for video generation … | 30 | 1500 | active |
| uccl-project/uccl UCCL is a high-performance GPU communication library written in C++ that provides collectives (as a drop-in NCCL/RCCL replacement), P2P tra… | 71 | 1496 | active |
| bojone/bert4keras A lightweight, clean reimplementation of BERT and other transformer models (RoBERTa, ALBERT, T5, GPT, ELECTRA, NEZHA) for Keras/tf.keras. I… | 23 | 5415 | maintenance |
| k2-fsa/icefall Icefall is a collection of speech recognition (ASR) and TTS training recipes built on the k2 and lhotse libraries, implemented in Python wi… | 64 | 1482 | active |
| Lightning-AI/lightning-thunder Lightning Thunder is a source-to-source deep learning compiler for PyTorch that optimizes models for training and inference. It provides a … | 72 | 1469 | active |
| argilla-io/argilla Argilla is an open-source collaboration tool for AI engineers and domain experts to build, annotate, and curate high-quality datasets for N… | 79 | 5085 | maintenance |
| explosion/spacy-transformers A spaCy v3 extension package that provides pipeline components for using pretrained transformer models like BERT, RoBERTa, XLNet, and GPT-2… | 71 | 1409 | stable |
| thunlp/OpenPrompt OpenPrompt is a PyTorch-based open-source framework for prompt-learning, providing a standard, flexible pipeline of templates and verbalize… | 23 | 4890 | maintenance |
| bytedance/UNO UNO is a research framework from ByteDance for subject-driven image generation with diffusion transformers, supporting both single- and mul… | 38 | 1362 | active |
| morettt/my-neuro An open-source AI desktop companion framework inspired by Neuro-sama, letting users build a customizable Live2D character with sub-second v… | 84 | 1342 | active |
| ZJU-REAL/ClawGUI ClawGUI is a unified Python framework for GUI agents covering the full lifecycle: online reinforcement learning training (ClawGUI-RL with G… | 70 | 1338 | active |
| ImprintLab/Medical-SAM-Adapter Medical SAM Adapter (MSA) is a Python framework that fine-tunes Meta's Segment Anything Model for medical image segmentation using lightwei… | 39 | 1322 | active |
| open-edge-platform/geti Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru… | 98 | 1317 | active |
| jishengpeng/WavTokenizer WavTokenizer is a state-of-the-art discrete neural audio codec that compresses speech, music, and general audio into only 40 or 75 discrete… | 27 | 1316 | active |
| DAMO-NLP-SG/VideoLLaMA2 VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro… | 25 | 1307 | active |
| robodhruv/visualnav-transformer Official code and pre-trained checkpoints for the GNM, ViNT, and NoMaD family of general-purpose goal-conditioned visual navigation policie… | 19 | 1294 | active |
| microsoft/BioGPT BioGPT is Microsoft's domain-specific generative Transformer language model pre-trained on biomedical text, with implementation code and pr… | 32 | 4488 | maintenance |
| amaiya/ktrain ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul… | 25 | 1268 | active |
| LTH14/fractalgen A PyTorch implementation of Fractal Generative Models (FractalGen), enabling pixel-by-pixel high-resolution image generation. It includes p… | 24 | 1244 | active |
| HJYao00/Mulberry Mulberry is a research implementation of an o1-like multimodal large language model (MLLM) that performs step-by-step reasoning and reflect… | 49 | 1243 | active |
| declare-lab/tango Tango is a family of latent diffusion models for text-to-audio generation, with Tango 2 improving prompt alignment via DPO-based fine-tunin… | 45 | 1239 | active |
| Aratako/Irodori-TTS Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer… | 58 | 1219 | active |
| TIGER-AI-Lab/OpenResearcher OpenResearcher is a fully open-source pipeline for synthesizing long-horizon deep research trajectories using LLM agents with retrieval and… | 54 | 1211 | active |
| NVIDIA/audio-flamingo NVIDIA's PyTorch implementation of the Audio Flamingo series of large audio-language models (AF1, AF2, AF3, and Music Flamingo) for audio u… | 50 | 1182 | active |
| mlfoundations/open_flamingo OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima… | 23 | 4118 | maintenance |
| bytedance/1d-tokenizer A research repository from ByteDance containing code and pretrained model weights for 1D visual tokenizers (TiTok, TA-TiTok, FlowTok) and i… | 29 | 1172 | active |
| nv-tlabs/LLaMA-Mesh LLaMA-Mesh is a fine-tuned large language model from NVIDIA Research that generates and understands 3D meshes by representing vertex coordi… | 28 | 1166 | active |
| NVIDIA-NeMo/Gym NeMo Gym is a Python library from NVIDIA for evaluating and improving LLM models and agents using environments. It provides infrastructure … | 80 | 1156 | active |
| context-labs/HALO HALO (Hierarchical Agent Loop Optimizer) is a desktop application that imports LLM agent traces from sources like Langfuse, Arize, JSONL, o… | 77 | 1149 | active |
| rohitgandikota/sliders Official implementation of Concept Sliders, LoRA adaptors that enable precise, plug-and-play control of attributes in diffusion models like… | 52 | 1139 | active |
| xzf-thu/Mega-ASR Mega-ASR is a foundation automatic speech recognition model trained on 2.6M samples spanning 7 atomic acoustic conditions and 54 compound r… | 59 | 1136 | active |
| aliyun/SimAI SimAI is a large-scale network simulation toolkit from Alibaba Cloud for modeling AI training and inference workloads on GPU clusters, publ… | 66 | 1133 | active |
| premieroctet/photoshot Photoshot is an open-source web application that generates custom AI avatars from user-uploaded selfies using fine-tuned text-to-image mode… | 22 | 3870 | maintenance |
| alibaba-damo-academy/RynnVLA-002 RynnVLA-002 is a unified autoregressive Vision-Language-Action and world model that generates robot actions from text and image observation… | 43 | 1119 | active |
| openai/improved-diffusion The official codebase for OpenAI's Improved Denoising Diffusion Probabilistic Models paper, providing a Python package for training and sam… | 32 | 3844 | maintenance |
| unitreerobotics/unifolm-world-model-action UnifoLM-WMA-0 is Unitree's open-source world-model-action framework for general-purpose robot learning across multiple robotic embodiments.… | 50 | 1109 | active |
| rhymes-ai/Aria Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda… | 23 | 1087 | active |
| meta-pytorch/monarch Monarch is a distributed programming framework for PyTorch built on scalable actor messaging, with actors grouped into meshes, supervision-… | 80 | 1073 | active |
| InternRobotics/InternNav InternNav is an open-source PyTorch-based toolbox for building embodied navigation foundation models, supporting vision-language navigation… | 58 | 1061 | active |
| bytedance/SandboxFusion A secure, self-hosted code sandbox service from ByteDance that runs and judges code generated by LLMs across 20+ programming languages via … | 63 | 1060 | active |
| NVIDIA/DreamDojo NVIDIA's official PyTorch codebase for DreamDojo, a generalist robot world model pretrained on 44k hours of human egocentric video and post… | 48 | 1059 | active |
| juanjuandog/FinSight-AI FinSight AI is an open-source equity research workspace for A-share companies that turns market data, filings, and financial metrics into s… | 59 | 1034 | active |
| XGenerationLab/XiYan-SQL XiYan-SQL is a multi-generator ensemble framework for converting natural language questions into SQL queries, achieving SOTA results on ben… | 59 | 1019 | active |
| Alpha-VLLM/Lumina-DiMOO Lumina-DiMOO is an open-source omni diffusion large language model that uses fully discrete diffusion to handle multimodal inputs and outpu… | 55 | 1015 | active |
| hezarai/hezar Hezar is an all-in-one Python AI library for the Persian language, covering NLP, speech recognition, OCR, and image captioning through a ta… | 78 | 1013 | active |
| wang-rui/phishguard-scaffold PhishGuard is a Python research framework that jointly performs phishing detection and dissemination control on social media using LLaMA-ba… | 47 | 1009 | active |
| PixArt-alpha/PixArt-alpha PixArt-α is a Transformer-based text-to-image diffusion model with PyTorch model definitions, pre-trained weights, and inference/training c… | 27 | 3304 | maintenance |
| FudanNLP/fastNLP fastNLP is a lightweight, modularized and extensible NLP framework in Python that reduces engineering boilerplate such as data processing l… | 23 | 3141 | maintenance |
| Conchylicultor/DeepQA DeepQA is a TensorFlow implementation of Google's 'A Neural Conversational Model', a seq2seq RNN-based deep learning chatbot. It supports t… | 32 | 2910 | maintenance |
| tensorflow/lingvo Lingvo is a TensorFlow-based framework for building neural networks, particularly sequence models, with a focus on speech recognition, mach… | 72 | 2864 | maintenance |
| illuin-tech/colpali ColPali Engine is the Python library for training and running inference with ColVision visual document retrieval models such as ColPali, Co… | 91 | 2798 | maintenance |
| microsoft/pai OpenPAI is an open-source AI platform from Microsoft that provides resource scheduling and cluster management for machine learning workload… | 66 | 2688 | maintenance |
| OFA-Sys/OFA OFA is a unified sequence-to-sequence pretrained model supporting English and Chinese that unifies cross-modality, vision, and language tas… | 32 | 2557 | maintenance |
| google-research/electra ELECTRA is a research library from Google for self-supervised pre-training of transformer text encoders using a discriminator-based objecti… | 10 | 2367 | maintenance |
| MetaGLM/FinGLM FinGLM is an open, community-driven financial LLM project centered on a dialog-based question-answering system that analyzes Chinese listed… | 27 | 2258 | maintenance |
| allenai/longformer Longformer is a pretrained transformer model family (including the LongformerEncoderDecoder/LED variant) that processes long documents up t… | 23 | 2205 | maintenance |
| archinetai/audio-diffusion-pytorch A PyTorch library for audio generation using diffusion models, supporting unconditional and text-conditional generation, diffusion autoenco… | 23 | 2096 | maintenance |
| alibaba/AliceMind AliceMind is Alibaba's collection of pre-trained encoder-decoder language models and related NLP techniques, including StructBERT, PALM, VE… | 23 | 2041 | maintenance |
| NVlabs/alpamayo NVIDIA Alpamayo 1 is an open 10B-parameter reasoning vision-language-action (VLA) model for autonomous vehicles that pairs driving trajecto… | 59 | 2005 | maintenance |
| google/sling SLING is a natural language frame semantics parser that annotates text with frame semantic graph representations using bi-directional LSTMs… | 10 | 1930 | maintenance |
| appvision-ai/fast-bert Fast-Bert is a Python deep learning library for training and deploying BERT, RoBERTa, and XLNet based models for NLP tasks, starting with m… | 23 | 1917 | maintenance |
| Ucas-HaoranWei/Vary Official ECCV 2024 implementation of Vary, a method for scaling up the vision vocabulary of large vision-language models. It provides train… | 26 | 1889 | maintenance |
| MrNothing/AI-Blocks AI-Blocks is a WYSIWYG desktop application for visually building machine learning models by dragging objects with attached scripts in a sce… | 23 | 1857 | maintenance |
| microsoft/i-Code Microsoft's i-Code is a collection of research models and frameworks for integrative, composable multimodal AI spanning vision, language, a… | 32 | 1703 | maintenance |
| google-research/pegasus PEGASUS is Google Research's implementation of transformer encoder-decoder models pre-trained with the Gap Sentences Generation objective f… | 10 | 1655 | maintenance |
| allenai/bilm-tf A TensorFlow implementation of the bidirectional language model (biLM) used to compute ELMo deep contextualized word representations. It su… | 32 | 1612 | maintenance |
| Delta-ML/delta DELTA is a deep learning based end-to-end natural language and speech processing platform built on TensorFlow and Python 3. It provides one… | 10 | 1607 | maintenance |
| nlpyang/BertSum BertSum is the official PyTorch implementation of the paper 'Fine-tune BERT for Extractive Summarization', providing preprocessing pipeline… | 32 | 1506 | maintenance |
| microsoft/NeuronBlocks NeuronBlocks is an NLP deep learning modeling toolkit from Microsoft that lets users build end-to-end neural network training and inference… | 23 | 1452 | maintenance |
| Improbable-AI/walk-these-ways A sim-to-real reinforcement learning starter kit for the Unitree Go1 quadruped robot, implementing the Walk These Ways (MoB) locomotion con… | 32 | 1438 | maintenance |
| SamLynnEvans/Transformer A PyTorch implementation of the Transformer seq2seq model designed to build language translators from parallel corpora. It accompanies a tu… | 32 | 1430 | maintenance |
| facebookresearch/diplomacy_cicero Research code and model checkpoints for Cicero and Diplodocus, AI agents that play the board game Diplomacy at human level by combining lan… | 10 | 1428 | maintenance |
| opendilab/DI-star DI-star is a large-scale distributed training platform for building StarCraft II game AI, including supervised and reinforcement learning t… | 37 | 1393 | maintenance |
| wouterkool/attention-learn-to-route A PyTorch implementation of the attention-based neural model from the ICLR 2019 paper 'Attention, Learn to Solve Routing Problems!', traine… | 32 | 1383 | maintenance |
| nvidia-cosmos/cosmos-predict2.5 NVIDIA Cosmos-Predict2.5 is a family of world foundation models (WFMs) that generate video predictions of future world states for physical … | 72 | 1355 | maintenance |
| NVlabs/prismer Official PyTorch implementation of Prismer, a data- and parameter-efficient vision-language model that ensembles pre-trained task-specific … | 30 | 1309 | maintenance |
| google-research/multilingual-t5 Code and resources for mT5, a massively multilingual text-to-text transformer pretrained on the mC4 corpus covering 101 languages. It repro… | 10 | 1294 | maintenance |
| ARM-software/ML-KWS-for-MCU TensorFlow models and training scripts for keyword spotting (wake-word detection) on Arm Cortex-M microcontrollers, accompanying the 'Hello… | 32 | 1249 | maintenance |
| XiangLi1999/Diffusion-LM Diffusion-LM is the official research code for the paper 'Diffusion-LM Improves Controllable Text Generation', implementing a diffusion-bas… | 32 | 1245 | maintenance |
| Timthony/self_drive A self-driving RC car project based on Raspberry Pi and TensorFlow/Keras. It collects camera images while a human drives the car on a taped… | 32 | 1132 | maintenance |
| Harmonai-org/sample-generator A set of tools and Jupyter notebooks for training generative diffusion models on arbitrary audio samples, built around Dance Diffusion. It … | 32 | 1117 | maintenance |
| pytorch/torchdynamo TorchDynamo is a Python-level JIT compiler that speeds up unmodified PyTorch programs by capturing Python bytecode into FX graphs. The proj… | 10 | 1078 | maintenance |
| kakaobrain/rq-vae-transformer The official PyTorch implementation of 'Autoregressive Image Generation using Residual Quantization' (CVPR 2022), implementing RQ-VAE and R… | 32 | 1030 | maintenance |
| turtlesoupy/this-word-does-not-exist A project that trains a GPT-2 variant to invent fake English words with generated definitions and example sentences, powering the thiswordd… | 72 | 1023 | maintenance |
| lucidrains/musiclm-pytorch A PyTorch library implementing MusicLM, Google's text-to-music generation model, by combining text-conditioned AudioLM with MuLan, a text-a… | 21 | 3293 | experimental |