Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: llm-training

818 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
mlrun/mlrun
MLRun is an open-source MLOps and AI orchestration framework for building, training, deploying, and monitoring machine learning and generat…
961692active
scverse/scvi-tools
scvi-tools is a Python library for deep probabilistic modeling and analysis of single-cell and spatial omics data, built on PyTorch, PyTorc…
931681stable
JIA-Lab-research/ControlNeXt
ControlNeXt is the official implementation of a controllable generation method for images and videos, built on Stable Diffusion XL, Stable …
241646active
amazon-far/holosoma
Holosoma is a Python framework for training and deploying reinforcement learning policies on humanoid robots, supporting locomotion and who…
611616active
SonyResearch/micro_diffusion
Official implementation of Sony Research's micro-budget approach to training large-scale text-to-image diffusion transformer models from sc…
241594active
MoonshotAI/Kimi-Linear
Kimi Linear is a hybrid linear attention architecture (Kimi Delta Attention, based on Gated DeltaNet) released by Moonshot AI with 48B-para…
401588active
FoundationVision/Infinity
Infinity is a bitwise autoregressive text-to-image generation model (CVPR 2025 Oral) with released training and inference code, checkpoints…
561587active
baaivision/Emu3.5
Emu3.5 is BAAI's native multimodal foundation model that jointly predicts next states across vision and language, trained on 10T+ interleav…
431547active
lucidrains/DALLE-pytorch
A PyTorch implementation/replication of OpenAI's DALL-E, a text-to-image transformer, including a discrete VAE and optional CLIP for rankin…
235627maintenance
ByteDance-Seed/Triton-distributed
Triton-distributed is a distributed compiler built on OpenAI Triton for computation-communication overlapping on multi-GPU systems. It lets…
631526active
Tencent/TFace
TFace is a research platform from Tencent Youtu Lab for trusty face analysis, covering face recognition, face security (anti-spoofing), fac…
571521active
microsoft/Mage
Mage is a family of lightweight 4B-parameter multimodal models from Microsoft, including Mage-VL for image and video understanding and Mage…
571516active
bytedance/SALMONN
SALMONN is a family of open-source multi-modal large language models from ByteDance and Tsinghua that unify speech, audio, music, and video…
731513active
ZFTurbo/Music-Source-Separation-Training
A Python training framework for music source separation models, supporting many architectures such as MDX23C, Demucs, Band Split RoFormer, …
831512active
wb14123/seq2seq-couplet
A deep learning project that generates Chinese couplets (对联) using a seq2seq model built with TensorFlow. It includes training scripts, a w…
325487maintenance
TencentARC/MotionCtrl
MotionCtrl is the official implementation of a SIGGRAPH 2024 paper providing a unified and flexible motion controller for video generation …
301500active
uccl-project/uccl
UCCL is a high-performance GPU communication library written in C++ that provides collectives (as a drop-in NCCL/RCCL replacement), P2P tra…
711496active
bojone/bert4keras
A lightweight, clean reimplementation of BERT and other transformer models (RoBERTa, ALBERT, T5, GPT, ELECTRA, NEZHA) for Keras/tf.keras. I…
235415maintenance
k2-fsa/icefall
Icefall is a collection of speech recognition (ASR) and TTS training recipes built on the k2 and lhotse libraries, implemented in Python wi…
641482active
Lightning-AI/lightning-thunder
Lightning Thunder is a source-to-source deep learning compiler for PyTorch that optimizes models for training and inference. It provides a …
721469active
argilla-io/argilla
Argilla is an open-source collaboration tool for AI engineers and domain experts to build, annotate, and curate high-quality datasets for N…
795085maintenance
explosion/spacy-transformers
A spaCy v3 extension package that provides pipeline components for using pretrained transformer models like BERT, RoBERTa, XLNet, and GPT-2…
711409stable
thunlp/OpenPrompt
OpenPrompt is a PyTorch-based open-source framework for prompt-learning, providing a standard, flexible pipeline of templates and verbalize…
234890maintenance
bytedance/UNO
UNO is a research framework from ByteDance for subject-driven image generation with diffusion transformers, supporting both single- and mul…
381362active
morettt/my-neuro
An open-source AI desktop companion framework inspired by Neuro-sama, letting users build a customizable Live2D character with sub-second v…
841342active
ZJU-REAL/ClawGUI
ClawGUI is a unified Python framework for GUI agents covering the full lifecycle: online reinforcement learning training (ClawGUI-RL with G…
701338active
ImprintLab/Medical-SAM-Adapter
Medical SAM Adapter (MSA) is a Python framework that fine-tunes Meta's Segment Anything Model for medical image segmentation using lightwei…
391322active
open-edge-platform/geti
Geti is an open-source, end-to-end Vision AI application from Intel that takes users from raw images to deployed computer vision models, ru…
981317active
jishengpeng/WavTokenizer
WavTokenizer is a state-of-the-art discrete neural audio codec that compresses speech, music, and general audio into only 40 or 75 discrete…
271316active
DAMO-NLP-SG/VideoLLaMA2
VideoLLaMA 2 is an open-source video large language model that adds spatial-temporal modeling and audio understanding to video-LLMs. It pro…
251307active
robodhruv/visualnav-transformer
Official code and pre-trained checkpoints for the GNM, ViNT, and NoMaD family of general-purpose goal-conditioned visual navigation policie…
191294active
microsoft/BioGPT
BioGPT is Microsoft's domain-specific generative Transformer language model pre-trained on biomedical text, with implementation code and pr…
324488maintenance
amaiya/ktrain
ktrain is a lightweight Python wrapper around TensorFlow Keras that provides pre-canned, low-code models for text, vision, graph, and tabul…
251268active
LTH14/fractalgen
A PyTorch implementation of Fractal Generative Models (FractalGen), enabling pixel-by-pixel high-resolution image generation. It includes p…
241244active
HJYao00/Mulberry
Mulberry is a research implementation of an o1-like multimodal large language model (MLLM) that performs step-by-step reasoning and reflect…
491243active
declare-lab/tango
Tango is a family of latent diffusion models for text-to-audio generation, with Tango 2 improving prompt alignment via DPO-based fine-tunin…
451239active
Aratako/Irodori-TTS
Irodori-TTS is a Flow Matching-based text-to-speech model with training and inference code, built on a Rectified Flow Diffusion Transformer…
581219active
TIGER-AI-Lab/OpenResearcher
OpenResearcher is a fully open-source pipeline for synthesizing long-horizon deep research trajectories using LLM agents with retrieval and…
541211active
NVIDIA/audio-flamingo
NVIDIA's PyTorch implementation of the Audio Flamingo series of large audio-language models (AF1, AF2, AF3, and Music Flamingo) for audio u…
501182active
mlfoundations/open_flamingo
OpenFlamingo is an open-source PyTorch implementation of DeepMind's Flamingo, a large multimodal vision-language model that interleaves ima…
234118maintenance
bytedance/1d-tokenizer
A research repository from ByteDance containing code and pretrained model weights for 1D visual tokenizers (TiTok, TA-TiTok, FlowTok) and i…
291172active
nv-tlabs/LLaMA-Mesh
LLaMA-Mesh is a fine-tuned large language model from NVIDIA Research that generates and understands 3D meshes by representing vertex coordi…
281166active
NVIDIA-NeMo/Gym
NeMo Gym is a Python library from NVIDIA for evaluating and improving LLM models and agents using environments. It provides infrastructure …
801156active
context-labs/HALO
HALO (Hierarchical Agent Loop Optimizer) is a desktop application that imports LLM agent traces from sources like Langfuse, Arize, JSONL, o…
771149active
rohitgandikota/sliders
Official implementation of Concept Sliders, LoRA adaptors that enable precise, plug-and-play control of attributes in diffusion models like…
521139active
xzf-thu/Mega-ASR
Mega-ASR is a foundation automatic speech recognition model trained on 2.6M samples spanning 7 atomic acoustic conditions and 54 compound r…
591136active
aliyun/SimAI
SimAI is a large-scale network simulation toolkit from Alibaba Cloud for modeling AI training and inference workloads on GPU clusters, publ…
661133active
premieroctet/photoshot
Photoshot is an open-source web application that generates custom AI avatars from user-uploaded selfies using fine-tuned text-to-image mode…
223870maintenance
alibaba-damo-academy/RynnVLA-002
RynnVLA-002 is a unified autoregressive Vision-Language-Action and world model that generates robot actions from text and image observation…
431119active
openai/improved-diffusion
The official codebase for OpenAI's Improved Denoising Diffusion Probabilistic Models paper, providing a Python package for training and sam…
323844maintenance
unitreerobotics/unifolm-world-model-action
UnifoLM-WMA-0 is Unitree's open-source world-model-action framework for general-purpose robot learning across multiple robotic embodiments.…
501109active
rhymes-ai/Aria
Aria is an open multimodal native Mixture-of-Experts (MoE) model with 25.3B total parameters (3.9B activated per token) and a 64K multimoda…
231087active
meta-pytorch/monarch
Monarch is a distributed programming framework for PyTorch built on scalable actor messaging, with actors grouped into meshes, supervision-…
801073active
InternRobotics/InternNav
InternNav is an open-source PyTorch-based toolbox for building embodied navigation foundation models, supporting vision-language navigation…
581061active
bytedance/SandboxFusion
A secure, self-hosted code sandbox service from ByteDance that runs and judges code generated by LLMs across 20+ programming languages via …
631060active
NVIDIA/DreamDojo
NVIDIA's official PyTorch codebase for DreamDojo, a generalist robot world model pretrained on 44k hours of human egocentric video and post…
481059active
juanjuandog/FinSight-AI
FinSight AI is an open-source equity research workspace for A-share companies that turns market data, filings, and financial metrics into s…
591034active
XGenerationLab/XiYan-SQL
XiYan-SQL is a multi-generator ensemble framework for converting natural language questions into SQL queries, achieving SOTA results on ben…
591019active
Alpha-VLLM/Lumina-DiMOO
Lumina-DiMOO is an open-source omni diffusion large language model that uses fully discrete diffusion to handle multimodal inputs and outpu…
551015active
hezarai/hezar
Hezar is an all-in-one Python AI library for the Persian language, covering NLP, speech recognition, OCR, and image captioning through a ta…
781013active
wang-rui/phishguard-scaffold
PhishGuard is a Python research framework that jointly performs phishing detection and dissemination control on social media using LLaMA-ba…
471009active
PixArt-alpha/PixArt-alpha
PixArt-α is a Transformer-based text-to-image diffusion model with PyTorch model definitions, pre-trained weights, and inference/training c…
273304maintenance
FudanNLP/fastNLP
fastNLP is a lightweight, modularized and extensible NLP framework in Python that reduces engineering boilerplate such as data processing l…
233141maintenance
Conchylicultor/DeepQA
DeepQA is a TensorFlow implementation of Google's 'A Neural Conversational Model', a seq2seq RNN-based deep learning chatbot. It supports t…
322910maintenance
tensorflow/lingvo
Lingvo is a TensorFlow-based framework for building neural networks, particularly sequence models, with a focus on speech recognition, mach…
722864maintenance
illuin-tech/colpali
ColPali Engine is the Python library for training and running inference with ColVision visual document retrieval models such as ColPali, Co…
912798maintenance
microsoft/pai
OpenPAI is an open-source AI platform from Microsoft that provides resource scheduling and cluster management for machine learning workload…
662688maintenance
OFA-Sys/OFA
OFA is a unified sequence-to-sequence pretrained model supporting English and Chinese that unifies cross-modality, vision, and language tas…
322557maintenance
google-research/electra
ELECTRA is a research library from Google for self-supervised pre-training of transformer text encoders using a discriminator-based objecti…
102367maintenance
MetaGLM/FinGLM
FinGLM is an open, community-driven financial LLM project centered on a dialog-based question-answering system that analyzes Chinese listed…
272258maintenance
allenai/longformer
Longformer is a pretrained transformer model family (including the LongformerEncoderDecoder/LED variant) that processes long documents up t…
232205maintenance
archinetai/audio-diffusion-pytorch
A PyTorch library for audio generation using diffusion models, supporting unconditional and text-conditional generation, diffusion autoenco…
232096maintenance
alibaba/AliceMind
AliceMind is Alibaba's collection of pre-trained encoder-decoder language models and related NLP techniques, including StructBERT, PALM, VE…
232041maintenance
NVlabs/alpamayo
NVIDIA Alpamayo 1 is an open 10B-parameter reasoning vision-language-action (VLA) model for autonomous vehicles that pairs driving trajecto…
592005maintenance
google/sling
SLING is a natural language frame semantics parser that annotates text with frame semantic graph representations using bi-directional LSTMs…
101930maintenance
appvision-ai/fast-bert
Fast-Bert is a Python deep learning library for training and deploying BERT, RoBERTa, and XLNet based models for NLP tasks, starting with m…
231917maintenance
Ucas-HaoranWei/Vary
Official ECCV 2024 implementation of Vary, a method for scaling up the vision vocabulary of large vision-language models. It provides train…
261889maintenance
MrNothing/AI-Blocks
AI-Blocks is a WYSIWYG desktop application for visually building machine learning models by dragging objects with attached scripts in a sce…
231857maintenance
microsoft/i-Code
Microsoft's i-Code is a collection of research models and frameworks for integrative, composable multimodal AI spanning vision, language, a…
321703maintenance
google-research/pegasus
PEGASUS is Google Research's implementation of transformer encoder-decoder models pre-trained with the Gap Sentences Generation objective f…
101655maintenance
allenai/bilm-tf
A TensorFlow implementation of the bidirectional language model (biLM) used to compute ELMo deep contextualized word representations. It su…
321612maintenance
Delta-ML/delta
DELTA is a deep learning based end-to-end natural language and speech processing platform built on TensorFlow and Python 3. It provides one…
101607maintenance
nlpyang/BertSum
BertSum is the official PyTorch implementation of the paper 'Fine-tune BERT for Extractive Summarization', providing preprocessing pipeline…
321506maintenance
microsoft/NeuronBlocks
NeuronBlocks is an NLP deep learning modeling toolkit from Microsoft that lets users build end-to-end neural network training and inference…
231452maintenance
Improbable-AI/walk-these-ways
A sim-to-real reinforcement learning starter kit for the Unitree Go1 quadruped robot, implementing the Walk These Ways (MoB) locomotion con…
321438maintenance
SamLynnEvans/Transformer
A PyTorch implementation of the Transformer seq2seq model designed to build language translators from parallel corpora. It accompanies a tu…
321430maintenance
facebookresearch/diplomacy_cicero
Research code and model checkpoints for Cicero and Diplodocus, AI agents that play the board game Diplomacy at human level by combining lan…
101428maintenance
opendilab/DI-star
DI-star is a large-scale distributed training platform for building StarCraft II game AI, including supervised and reinforcement learning t…
371393maintenance
wouterkool/attention-learn-to-route
A PyTorch implementation of the attention-based neural model from the ICLR 2019 paper 'Attention, Learn to Solve Routing Problems!', traine…
321383maintenance
nvidia-cosmos/cosmos-predict2.5
NVIDIA Cosmos-Predict2.5 is a family of world foundation models (WFMs) that generate video predictions of future world states for physical …
721355maintenance
NVlabs/prismer
Official PyTorch implementation of Prismer, a data- and parameter-efficient vision-language model that ensembles pre-trained task-specific …
301309maintenance
google-research/multilingual-t5
Code and resources for mT5, a massively multilingual text-to-text transformer pretrained on the mC4 corpus covering 101 languages. It repro…
101294maintenance
ARM-software/ML-KWS-for-MCU
TensorFlow models and training scripts for keyword spotting (wake-word detection) on Arm Cortex-M microcontrollers, accompanying the 'Hello…
321249maintenance
XiangLi1999/Diffusion-LM
Diffusion-LM is the official research code for the paper 'Diffusion-LM Improves Controllable Text Generation', implementing a diffusion-bas…
321245maintenance
Timthony/self_drive
A self-driving RC car project based on Raspberry Pi and TensorFlow/Keras. It collects camera images while a human drives the car on a taped…
321132maintenance
Harmonai-org/sample-generator
A set of tools and Jupyter notebooks for training generative diffusion models on arbitrary audio samples, built around Dance Diffusion. It …
321117maintenance
pytorch/torchdynamo
TorchDynamo is a Python-level JIT compiler that speeds up unmodified PyTorch programs by capturing Python bytecode into FX graphs. The proj…
101078maintenance
kakaobrain/rq-vae-transformer
The official PyTorch implementation of 'Autoregressive Image Generation using Residual Quantization' (CVPR 2022), implementing RQ-VAE and R…
321030maintenance
turtlesoupy/this-word-does-not-exist
A project that trains a GPT-2 variant to invent fake English words with generated definitions and example sentences, powering the thiswordd…
721023maintenance
lucidrains/musiclm-pytorch
A PyTorch library implementing MusicLM, Google's text-to-music generation model, by combining text-conditioned AudioLM with MuLan, a text-a…
213293experimental

← prev page 8 / 9 next →