function: llm-training
818 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| THUMNLab/AutoGL AutoGL is an autoML framework and toolkit for machine learning on graphs, built on PyTorch with PyTorch Geometric and DGL backends. It prov… | 47 | 1140 | active |
| horseee/LLM-Pruner LLM-Pruner is a PyTorch library implementing structural pruning of large language models based on gradient information, as published at Neu… | 29 | 1136 | active |
| NovaSearch-Team/RAG-Retrieval A Python library and toolkit for unified fine-tuning, inference, and distillation of RAG retrieval models, including embedding models, ColB… | 60 | 1127 | active |
| BeingBeyond/Being-H Being-H is a family of human-centric embodied foundation models, including VLA models (Being-H0.5, Being-H0) and latent world-action models… | 63 | 1126 | active |
| OpenDriveLab/UniVLA UniVLA is an open-source framework for training cross-embodiment vision-language-action (VLA) robot policies using task-centric latent acti… | 43 | 1124 | active |
| THUDM/SwissArmyTransformer SwissArmyTransformer (sat) is a PyTorch library for developing custom Transformer model variants where models like BERT, GPT, T5, GLM, and … | 23 | 1121 | active |
| FlagAI-Open/FlagAI FlagAI is a Python toolkit for training, fine-tuning, and deploying large-scale AI models across NLP, CV, and vision-language tasks. It int… | 64 | 3869 | maintenance |
| PRIME-RL/TTRL TTRL is an open-source implementation of Test-Time Reinforcement Learning, a method for training LLMs with RL on unlabeled data using major… | 51 | 1120 | active |
| datadreamer-dev/DataDreamer DataDreamer is an open-source Python library for prompting LLMs, generating synthetic datasets, and training or aligning models in reproduc… | 33 | 1117 | active |
| HITsz-TMG/Uni-MoE Uni-MoE is a family of open-source Mixture-of-Experts (MoE) based omnimodal large language models that understand and generate across text,… | 68 | 1116 | active |
| yahoo/TensorFlowOnSpark TensorFlowOnSpark is a Python library that lets existing TensorFlow programs run distributed training and inference on Apache Spark and Had… | 23 | 3845 | maintenance |
| RUC-NLPIR/ARPO ARPO (Agentic Reinforced Policy Optimization) is a reinforcement learning algorithm and training framework for LLM agents, published at ICL… | 54 | 1109 | active |
| MoonshotAI/MoonEP MoonEP is an Expert Parallelism communication library for Mixture-of-Experts training that keeps token loads perfectly balanced across rank… | 56 | 1101 | active |
| yongliang-wu/DFT DFT (Dynamic Fine-Tuning) is the official implementation of an ICLR 2026 paper that improves Supervised Fine-Tuning of LLMs by dynamically … | 61 | 1100 | active |
| BytedTsinghua-SIA/MemAgent MemAgent is a reinforcement-learning framework for training LLM agents that process arbitrarily long contexts via a memory mechanism within… | 55 | 1100 | active |
| yandex/YaLM-100B YaLM-100B is a GPT-like pretrained language model with 100 billion parameters, trained by Yandex on English and Russian text using DeepSpee… | 32 | 3757 | maintenance |
| flagos-ai/FlagGems FlagGems is a high-performance operator library for large language models written in the Triton language, providing backend-neutral GPU ker… | 89 | 1089 | active |
| mymusise/ChatGLM-Tuning A Python toolkit for fine-tuning the ChatGLM-6B large language model using LoRA (Low-Rank Adaptation) with the Alpaca dataset. It provides … | 10 | 3740 | maintenance |
| bytedance/byteps BytePS is a high-performance parameter server framework for distributed deep neural network training, supporting TensorFlow, Keras, PyTorch… | 10 | 3717 | maintenance |
| GAIR-NLP/LIMO LIMO is a research project and training framework demonstrating that large language models can achieve strong mathematical reasoning with o… | 36 | 1083 | active |
| kimiyoung/transformer-xl Official implementation of Transformer-XL, an attention-based language model architecture that extends context beyond a fixed length via se… | 32 | 3714 | maintenance |
| lasgroup/SDPO SDPO (Self-Distilled Policy Optimization) is a research library implementing a reinforcement learning framework for post-training large lan… | 56 | 1075 | active |
| open-gigaai/giga-train GigaTrain is an efficient and scalable Python training framework for large AI models, supporting distributed multi-GPU/multi-node execution… | 62 | 1072 | active |
| THUDM/GLM GLM is a general language model pretrained with an autoregressive blank-filling objective, released with pretrained checkpoints and fine-tu… | 32 | 3652 | maintenance |
| NousResearch/DisTrO DisTrO is a family of low-latency distributed optimizers that reduce inter-GPU communication requirements by three to four orders of magnit… | 44 | 1058 | active |
| open-gigaai/giga-models GigaModels is an open-source Python framework providing pipelines for training, inference, deployment, and compression of multi-modal, gene… | 62 | 1057 | active |
| XYZ-AI-Lab/axrl AxisRL is an agentic reinforcement learning post-training framework for large language models, built on SGLang for high-throughput rollout … | 55 | 1056 | active |
| X-LANCE/SLAM-LLM SLAM-LLM is a deep learning toolkit for training custom multimodal large language models focused on speech, language, audio, and music proc… | 55 | 1056 | active |
| zhuzilin/ring-flash-attention A Python library implementing RingAttention on top of FlashAttention for distributed long-context transformer training. It provides varlen … | 44 | 1049 | active |
| arcee-ai/DistillKit DistillKit is an open-source Python toolkit for knowledge distillation of large language models, supporting both online and offline distill… | 60 | 1047 | active |
| TIGER-AI-Lab/verl-tool VerlTool is a unified, extensible framework built on verl for training LLM agents with tool use via reinforcement learning. It decouples ac… | 66 | 1036 | active |
| NVlabs/DiffusionNFT DiffusionNFT is a research library implementing an online reinforcement learning paradigm for diffusion models that optimizes policy direct… | 47 | 1034 | active |
| NVIDIA-NeMo/Skills Nemo-Skills is a collection of Python pipelines for improving the skills of large language models, covering synthetic data generation, mode… | 70 | 1031 | active |
| kuleshov-group/bd3lms BD3-LMs is a research implementation of Block Discrete Denoising Diffusion Language Models that interpolate between autoregressive and diff… | 34 | 1029 | active |
| fla-org/native-sparse-attention Efficient Triton kernel implementations of Native Sparse Attention (NSA), a hardware-aligned, natively trainable sparse attention mechanism… | 49 | 1020 | active |
| tensorflow/adanet AdaNet is a lightweight TensorFlow-based AutoML framework that automatically learns high-quality neural network architectures and ensembles… | 10 | 3454 | maintenance |
| microsoft/Tutel Tutel is Microsoft's optimized Mixture-of-Experts (MoE) library for efficient training and inference of large language models, featuring dy… | 79 | 1016 | active |
| deepseek-ai/DeepSeek-Math DeepSeekMath is a 7B open language model specialized in mathematical reasoning, initialized from DeepSeek-Coder and trained on 500B math-re… | 26 | 3436 | maintenance |
| EvolvingLMMs-Lab/Otter Otter is a multi-modal vision-language model built on OpenFlamingo, instruction-tuned on the MIMIC-IT dataset with image and video understa… | 21 | 3436 | maintenance |
| kaito-project/kaito KAITO is a Kubernetes operator suite that automates LLM inference, fine-tuning, and RAG engine deployment using simplified CRD APIs. It aut… | 95 | 1009 | active |
| TRI-ML/prismatic-vlms Prismatic VLMs is a PyTorch-based codebase for training visually-conditioned language models (VLMs) with flexible vision backbones like CLI… | 25 | 1009 | active |
| AutoArk/open-audio-opd An industrial training stack for online policy distillation (OPD) of audio models, distilling compact ASR (and planned TTS) student models … | 52 | 1007 | active |
| facebookresearch/fairscale FairScale is a PyTorch extension library providing composable modules and APIs for high-performance, large-scale distributed training, incl… | 10 | 3407 | maintenance |
| MoonshotAI/checkpoint-engine Checkpoint-engine is a lightweight Python middleware for updating model weights in-place across LLM inference engines, a critical step in r… | 80 | 1005 | active |
| minimaxir/gpt-2-simple A Python package that simplifies fine-tuning OpenAI's GPT-2 text-generation model (124M/355M) on custom text and generating text from the r… | 23 | 3400 | maintenance |
| TinyLLaVA/TinyLLaVA_Factory TinyLLaVA Factory is an open-source modular PyTorch/HuggingFace codebase for training small-scale large multimodal models (LMMs) that combi… | 68 | 1004 | active |
| pjlab-sys4nlp/llama-moe LLaMA-MoE is a Python toolkit and model series for building Mixture-of-Experts (MoE) language models from dense LLaMA models via expert con… | 19 | 1001 | active |
| catalyst-team/catalyst Catalyst is a high-level PyTorch framework for deep learning research and development, focused on reproducibility, rapid experimentation, a… | 64 | 3382 | maintenance |
| argilla-io/distilabel Distilabel is a Python framework for building scalable pipelines that generate synthetic data and AI feedback, based on verified research p… | 74 | 3378 | maintenance |
| JIA-Lab-research/MGM Official PyTorch implementation of Mini-Gemini, a multimodal vision-language model framework built on LLaVA that supports dense and MoE LLM… | 25 | 3327 | maintenance |
| bytedance/lightseq LightSeq is a high-performance CUDA-based library for training and inference of sequence models like Transformer, BERT, GPT, and BART, with… | 10 | 3295 | maintenance |
| google-research/albert Official TensorFlow implementation and pretrained checkpoints of ALBERT, a lite version of BERT for self-supervised learning of language re… | 10 | 3278 | maintenance |
| JoePenna/Dreambooth-Stable-Diffusion A Jupyter Notebook-based implementation of Dreambooth fine-tuning for Stable Diffusion, adapted from XavierXiao's repo with tweaks for trai… | 32 | 3211 | maintenance |
| huawei-noah/Pretrained-Language-Model A collection of pretrained language models and optimization techniques from Huawei Noah's Ark Lab, including PanGu-α (200B-parameter Chines… | 32 | 3165 | maintenance |
| project-baize/baize-chatbot Baize is an open-source chat model built on LLaMA using LoRA parameter-efficient tuning, trained on 100k self-chat dialogs generated by Cha… | 21 | 3150 | maintenance |
| linyiLYi/bilibot A local chatbot fine-tuned from Bilibili user comments, built on Qwen1.5-32B-Chat using Apple's MLX LoRA fine-tuning. It supports text chat… | 24 | 3139 | maintenance |
| microsoft/torchscale A PyTorch library from Microsoft implementing foundation Transformer architectures such as DeepNet, Magneto, RetNet, LongNet, BitNet, and X… | 32 | 3138 | maintenance |
| dbiir/UER-py UER-py is a PyTorch framework for pre-training transformer language models (BERT, GPT-2, T5, ELMo, etc.) and fine-tuning them on downstream… | 32 | 3112 | maintenance |
| salesforce/CodeT5 Official research release of CodeT5 and CodeT5+ open code large language models from Salesforce Research for code understanding and generat… | 10 | 3093 | maintenance |
| rinongal/textual_inversion Official implementation of the Textual Inversion paper, which learns new word embeddings in a frozen text-to-image (Latent Diffusion) model… | 32 | 3055 | maintenance |
| google-research/t5x T5X is a modular, composable framework built on JAX and Flax for high-performance training, evaluation, and inference of sequence models at… | 75 | 2998 | maintenance |
| yangjianxin1/GPT2-chitchat A GPT2-based Chinese chitchat dialogue model project built on HuggingFace transformers, including training, preprocessing, and interactive … | 32 | 2996 | maintenance |
| FreedomIntelligence/LLMZoo LLM Zoo is a project providing data, models, and evaluation benchmarks for large language models, including the multilingual Phoenix and Ch… | 30 | 2935 | maintenance |
| facebookresearch/XLM PyTorch implementation of Cross-lingual Language Model Pretraining (XLM) from Facebook AI Research, covering MLM, CLM, and TLM objectives p… | 10 | 2920 | maintenance |
| Tencent/PocketFlow PocketFlow is an open-source AutoML framework from Tencent AI Lab for automatically compressing and accelerating deep learning models. Deve… | 32 | 2909 | maintenance |
| cybertronai/gradient-checkpointing A Python library that reduces GPU memory usage when training very deep neural networks via gradient checkpointing, trading computation for … | 32 | 2843 | maintenance |
| OpenPipe/OpenPipe OpenPipe is an open-source fine-tuning and model-hosting platform that turns expensive LLM prompts into cheaper fine-tuned models. It offer… | 29 | 2830 | maintenance |
| Alpha-VLLM/LLaMA2-Accessory LLaMA2-Accessory is an open-source Python toolkit for pretraining, finetuning, and deploying large language models and multimodal LLMs, inc… | 29 | 2800 | maintenance |
| PhoebusSi/Alpaca-CoT Alpaca-CoT is an instruction-tuning platform that unifies interfaces for instruction data collection, parameter-efficient fine-tuning metho… | 30 | 2791 | maintenance |
| liucongg/ChatGLM-Finetuning A Python toolkit for fine-tuning ChatGLM-6B, ChatGLM2-6B, and ChatGLM3-6B large language models using Freeze, LoRA, P-Tuning, and full-para… | 21 | 2771 | maintenance |
| kyegomez/OpenMythos OpenMythos is an open-source, theoretical PyTorch implementation of a Recurrent-Depth Transformer (RDT) architecture inspired by speculatio… | 51 | 14805 | experimental |
| JIA-Lab-research/LongLoRA LongLoRA is an efficient fine-tuning approach and codebase that extends the context length of pre-trained LLMs (Llama2 7B/13B/70B) using sh… | 27 | 2686 | maintenance |
| alibaba/pipcook Pipcook is a JavaScript application framework for machine learning and its engineering, aimed at enabling JavaScript and front-end engineer… | 67 | 2594 | maintenance |
| jcjohnson/torch-rnn torch-rnn provides high-performance, reusable RNN and LSTM modules for the Torch7 deep learning framework, used for character-level languag… | 32 | 2559 | maintenance |
| LiyuanLucasLiu/RAdam RAdam is a Python implementation of Rectified Adam, a variant of the Adam optimizer that analytically reduces the large variance of adaptiv… | 32 | 2549 | maintenance |
| automl/Auto-PyTorch Auto-PyTorch is an AutoML framework that jointly performs neural architecture search and hyperparameter optimization for PyTorch models, us… | 23 | 2541 | maintenance |
| young-geng/EasyLM EasyLM is a JAX/Flax-based framework for pre-training, finetuning, evaluating, and serving large language models like LLaMA. It scales trai… | 32 | 2514 | maintenance |
| microsoft/Graphormer Graphormer is a deep learning library from Microsoft providing a graph transformer backbone for molecular modeling tasks. It ships pre-trai… | 62 | 2474 | maintenance |
| microsoft/DialoGPT DialoGPT is a large-scale pretrained dialogue response generation model from Microsoft, based on GPT-2 and trained on 147M multi-turn Reddi… | 10 | 2420 | maintenance |
| OpenBMB/CPM-Bee CPM-Bee is a fully open-source, commercially usable 10-billion-parameter bilingual (Chinese-English) foundation large language model traine… | 71 | 2406 | maintenance |
| allenai/RL4LMs RL4LMs is a modular Python library from AllenAI for fine-tuning language models with reinforcement learning to align with human preferences… | 32 | 2394 | maintenance |
| asyml/texar Texar is a modularized Python toolkit for machine learning, especially natural language processing and text generation, built on TensorFlow… | 65 | 2389 | maintenance |
| TigerResearch/TigerBot TigerBot is a multi-language, multi-task large language model project from TigerResearch, providing pretrained and chat-tuned model weights… | 29 | 2259 | maintenance |
| x-flux XLabs AI's training scripts for fine-tuning the FLUX.1 diffusion model with LoRA, ControlNet, and IP-Adapter adapters, using DeepSpeed and … | 23 | 2229 | maintenance |
| epfLLM/meditron Meditron is a suite of open-source medical large language models (7B and 70B) adapted from Llama-2 via continued pretraining on a curated m… | 27 | 2208 | maintenance |
| lucidrains/reformer-pytorch A PyTorch implementation of the Reformer, an efficient Transformer architecture using LSH attention, reversible networks, and chunking to h… | 23 | 2191 | maintenance |
| alibaba/EasyNLP EasyNLP is a comprehensive PyTorch-based NLP toolkit from Alibaba that provides training, inference, and deployment for pre-trained languag… | 23 | 2184 | maintenance |
| ai-forever/ru-gpts A repository of Russian GPT-3 language models (ruGPT3XL/Large/Medium/Small and ruGPT2Large) with usage and fine-tuning examples. It provide… | 32 | 2087 | maintenance |
| THUDM/P-tuning-v2 P-tuning v2 is a Python implementation of deep prompt tuning, applying trainable continuous prompts at every transformer layer so prompt tu… | 32 | 2078 | maintenance |
| neulab/prompt2model Prompt2Model is a Python library that takes a natural language task description (like an LLM prompt) and automatically generates a small, s… | 21 | 2018 | maintenance |
| haitongli/knowledge-distillation-pytorch A PyTorch framework for running knowledge distillation experiments, supporting both shallow (teacher-to-small-CNN) and deep distillation on… | 32 | 2000 | maintenance |
| databricks/spark-deep-learning Deep Learning Pipelines for Apache Spark, now reduced to the HorovodRunner component for distributed deep learning training via Horovod on … | 23 | 1987 | maintenance |
| mit-han-lab/once-for-all Once-for-All (OFA) is a PyTorch library implementing the ICLR 2020 Once-for-All network, which trains a single supernet that can be special… | 23 | 1956 | maintenance |
| danieldjohnson/biaxial-rnn-music-composition A Python implementation of a biaxial recurrent neural network (LSTM-based) trained to generate classical music from MIDI data. It includes … | 32 | 1926 | maintenance |
| Palashio/libra Libra is a Python autoML library that lets users build, train, and evaluate machine learning models with one-line natural-language-style qu… | 40 | 1907 | maintenance |
| minimaxir/aitextgen aitextgen is a Python library for training and generating text with GPT-2 and GPT Neo models, built on PyTorch, Hugging Face Transformers, … | 23 | 1838 | maintenance |
| zai-org/CogView CogView is a 4-billion-parameter pretrained transformer model for text-to-image image generation, released alongside a NeurIPS 2021 paper. … | 32 | 1799 | maintenance |
| jquesnelle/yarn Reference implementation of YaRN, an efficient method for extending the context window of large language models, published as an ICLR 2024 … | 29 | 1777 | maintenance |
| cmu-db/noisepage NoisePage is a relational database management system from Carnegie Mellon University designed for autonomous, self-driving operation, with … | 10 | 1765 | maintenance |
| huggingface/transfer-learning-conv-ai A clean, commented PyTorch codebase for training a dialog/chatbot agent by transfer learning from OpenAI GPT/GPT-2 language models. It repr… | 32 | 1754 | maintenance |