Ross ROSS = Recommend OSS · open-source software intelligence for agents

resource: data-generation

175 resources, primary matches first, then adoption-weighted; health v2 shown.

ResourceHealth v2StarsMaturity
felixonmars/fcitx5-pinyin-zhwiki
A tool and generated dictionary that builds a Pinyin dictionary for fcitx5-pinyin and Rime input methods from Chinese Wikipedia data. It co…
631087active
eduosi/district
A dataset of Chinese administrative divisions (provinces, cities, districts/counties) with names, pinyin, pinyin initials, administrative c…
511082active
cov-lineages/pango-designation
The authoritative repository maintaining Pango lineage designations for SARS-CoV-2, hosting the lineage description list, sequence designat…
941076active
KhronosGroup/glTF-Sample-Models
A collection of sample glTF 1.0 and 2.0 3D models maintained by Khronos Group for testing and demonstrating the glTF format, including PBR …
103615maintenance
marbl/CHM13
The Telomere-to-Telomere (T2T) consortium's complete telomere-to-telomere assembly of the CHM13hTERT human cell line, including the T2T-CHM…
731054stable
EBazarov/nsfw_data_source_urls
A curated collection of text files listing over 1.5 million image URLs across 159 categories, intended for downloading and building NSFW im…
323577maintenance
OpenSourceAP/CrossSection
Code and data accompanying Chen and Zimmermann's paper on open source cross-sectional asset pricing. It reproduces dozens of stock-level pr…
491034active
Rivens7/Livelist
A curated IPTV live stream source list focused on IPv6, auto-updated and synced from popular Chinese IPTV projects like fanmingming/live an…
601022active
CLUEbenchmark/CLUECorpus2020
CLUECorpus2020 is a 100GB cleaned Chinese text corpus (35 billion characters) derived from Common Crawl for pre-training language models, p…
621018stable
ignis-sec/Pwdb-Public
A curated dataset of password statistics and wordlists extracted from over one billion leaked credentials, including filtered lists like a …
323286maintenance
wecatch/china_regions
A dataset of China's administrative divisions (province, city, county, town, village) scraped from the national statistics bureau standard,…
233155maintenance
thunlp/UltraChat
UltraChat is a large-scale dataset of 1.57M informative, diverse multi-round dialogue instructions generated with LLMs, plus UltraLM, a ser…
302890maintenance
sbousseaden/EVTX-ATTACK-SAMPLES
A dataset of ~200 Windows EVTX event log samples mapped to MITRE ATT&CK tactics and techniques, covering attack and post-exploitation behav…
322614maintenance
mahavivo/english-wordlists
A collection of curated English wordlists for Chinese learners, including CET-4/6, Taiwan high school 7000 words, TOEFL, GRE, and COCA 2000…
322457maintenance
npm/npm-expansions
A community-curated list of humorous three-word expansions of the acronym 'npm', used on the npmjs.com website header and published as an n…
102308maintenance
openai/gpt-2-output-dataset
A dataset from OpenAI containing 250K WebText documents alongside 250K samples each from four GPT-2 model sizes, both untruncated and Top-K…
102027maintenance
mercari/engineer-vocabulary-list
A curated list of 100 Japanese/English vocabulary words for engineers, collected from real engineering meetings at Mercari and Merpay. Dist…
321974maintenance
google-deepmind/mathematics_dataset
A Python library from DeepMind that generates synthetic mathematical question-and-answer pairs at roughly school-level difficulty across to…
321965maintenance
wy-luke/All-Chinese-Character-Set
A curated dataset of text files containing complete Chinese character sets, including 3500/7000 common hanzi lists, all characters, and ful…
691939maintenance
airyland/china-area-data
A JavaScript npm package providing China's administrative division data (provinces, cities, districts/counties) sourced from the National B…
321926maintenance
bridgetkromhout/devops-against-humanity
DevOps Against Humanity is a community-written expansion deck for Cards Against Humanity, distributed as card text in Google Docs, CSVs, an…
321848maintenance
EleutherAI/the-pile
The Pile is a large, diverse, open-source language modeling dataset composed of many smaller text sources combined together. This repositor…
321673maintenance
teknium1/GPTeacher
GPTeacher is a collection of modular instruction-tuning datasets generated by GPT-4, including General-Instruct, Roleplay-Instruct, Code-In…
301667maintenance
AgaMiko/data-augmentation-review
A curated list of data augmentation resources covering libraries, GitHub repos, and papers across images, NLP, audio, time series, graphs, …
231639maintenance
hermitdave/FrequencyWords
A repository of preprocessed word frequency lists for many languages, generated from OpenSubtitles corpora, plus the C# generator code. Eac…
321535maintenance
sahil280114/codealpaca
Code Alpaca is a 20K instruction-following dataset and training code for fine-tuning LLaMA models on code generation tasks, based on Stanfo…
301514maintenance
PolyAI-LDN/conversational-datasets
A collection of tools and scripts from PolyAI for generating large, reproducible datasets for conversational response selection, including …
321401maintenance
KaiDMML/FakeNewsNet
FakeNewsNet is a research dataset for fake news detection, with minimal CSV files of news samples from PolitiFact and GossipCop plus Python…
321346maintenance
zxcvbn001/password_brute_dictionary
A collection of Chinese-oriented password brute-force dictionaries derived from over 42 million leaked passwords across five platforms. It …
321325maintenance
dansinker/tacofancy
A community-driven repository of taco recipes organized in an object-oriented structure (base layers, mixins, condiments, seasonings, shell…
321301maintenance
google-deepmind/rc-data
A question answering corpus from DeepMind accompanying the paper 'Teaching Machines to Read and Comprehend' (Hermann et al., NIPS 2015). It…
101296maintenance
jayleicn/animeGAN
A simple PyTorch implementation of DCGAN focused on generating anime face images, including a pretrained model and Jupyter notebook demos. …
231277maintenance
hikariming/chat-dataset-baseline
A curated Chinese conversational dataset plus fine-tuning scripts for training chat models like ChatGLM, built on top of LLaMA-Factory. It …
391191maintenance
rogerzhu/MNWeeklyCategory
A categorized collection of over 10,000 curated links from 327+ issues of the Chinese 'Manong Weekly' (码农周刊) newsletter, organized into mar…
321161maintenance
google-deepmind/funsearch
FunSearch is the official repository accompanying DeepMind's Nature 2023 paper on discovering new mathematical results via program search w…
271110maintenance
kaonashi-passwords/Kaonashi
Kaonashi is a collection of password wordlists, hashcat rules, and masks derived from large-scale analysis of billions of real leaked passw…
321099maintenance
k8gege/PasswordDic
A collection of password dictionaries (wordlists) including top weak passwords from 2011-2019, SSH/VPS server passwords, admin panel passwo…
321053maintenance
1eez/103976
A dataset of 103,976 English words with Chinese translations, parts of speech, and multiple definitions, distributed as SQL, CSV, and Excel…
321053maintenance
allenai/natural-instructions
A community-built dataset of 1600+ NLP tasks with natural language instructions, from AllenAI. It supports instruction-tuning research like…
231046maintenance
unrealcv/synthetic-computer-vision
A curated list of synthetic image datasets, 3D model repositories, and rendering tools for computer vision research. It tracks publications…
321025maintenance
ohmybahgosh/RockYou2021.txt
RockYou2021.txt is a massive compiled password wordlist of roughly 82 billion unique entries (6-20 ASCII characters), aggregated from multi…
321012maintenance
mshumer/gpt-llm-trainer
A set of Jupyter/Colab notebooks that automate the LLM fine-tuning pipeline: given a task description, they generate a synthetic dataset wi…
374177experimental
first20hours/google-10000-english
A dataset of the 10,000 most common English words ranked by frequency, derived from n-gram frequency analysis of Google's Trillion Word Cor…
324464abandoned
feeddd/feeds
A community-maintained collection of free RSS/Atom/JSON feeds for WeChat official accounts, generated via Hamibot Android automation script…
322098abandoned
tintinweb/smart-contract-sanctuary
A large git-based dataset of Etherscan-verified Solidity smart contracts, organized as an index repository with per-chain submodules (Ether…
661593abandoned
FOSSASIA-Web/fossasia-communities
A repository of JSON API files describing FOSSASIA and other open source communities across Asia, used to power the FOSSASIA.net community …
321585abandoned
Intel-bigdata/HiBench
HiBench is a big data benchmark suite with 29 workloads across micro, machine learning, SQL, graph, websearch, and streaming categories for…
101484abandoned
AlanChen4/Summer-2024-SWE-Internships
A curated, auto-updated list of Summer 2024 software engineering internship postings, maintained by a Python job monitor and linked to the …
101362abandoned
robbiebarrat/rapping-neural-network
A generative art project: a recurrent neural network trained on Kanye West's discography that writes rap lyrics word by word with rhymes an…
231062abandoned
fashiontec/openwash-format
OpenWash is a specification repository from Fashiontec defining a format and specs, presumably for open data exchange in the fashion/washin…
321028abandoned
Hammer1/cozeworkflows
A curated collection of 200+ ready-to-import workflow JSON files for the Coze (扣子) AI bot-building platform, maintained by an AI blogger an…
335507active
arman-bd/guppylm
GuppyLM is a ~9M parameter language model trained from scratch to talk like a small fish, built as an educational project demonstrating the…
493435active
alex000kim/nsfw_data_scraper
A collection of shell scripts that automatically aggregate tens of thousands of images across five categories (porn, hentai, sexy, neutral,…
3212588maintenance
BannyLon/DifyAIA
An open-source collection of ready-to-import Dify workflow DSL examples (YAML files) with accompanying Flask helper services, created by a …
632642active
OpenCoder-llm/OpenCoder-llm
OpenCoder is a fully open and reproducible family of code large language models (1.5B and 8B base and chat variants) trained on 2.5 trillio…
222111active
lensesio/fast-data-dev
A Docker image that packages a complete Apache Kafka development environment, including a Kafka broker, Zookeeper/KRaft, Confluent Schema R…
562079active
Thinklab-SJTU/Bench2Drive
Bench2Drive is a closed-loop benchmark and dataset for end-to-end autonomous driving, built on CARLA with an RL-based expert driver (Think2…
681926active
legalize-dev/legalize-es
A git-versioned dataset of Spanish legislation in Markdown, where each law is a file and each reform is a commit dated with its official pu…
591873active
SocialAI-tianji/Tianji
Tianji is an open-source Chinese-language LLM application and tutorial project focused on social nuance ('renqing shigu') scenarios, coveri…
351822active
trickest/wordlists
A regularly updated collection of real-world infosec wordlists maintained by Trickest, including technology-specific path lists (WordPress,…
771791active
ubisoft/ubisoft-laforge-animation-dataset
The Ubisoft La Forge Animation Dataset (LAFAN1) is a motion capture dataset of 5 subjects, 77 sequences, and ~4.6 hours of character animat…
321568stable
microsoft/DNS-Challenge
Microsoft's repository for the Deep Noise Suppression (DNS) Challenge, containing datasets, scripts, and baseline models for training speec…
321465active
ikatsov/tensor-house
TensorHouse is a collection of reference Jupyter notebooks and demo AI/ML applications covering enterprise use cases such as marketing, pri…
231452active
pinchbench/skill
PinchBench is a benchmarking system that evaluates LLM models as OpenClaw coding agents using 53 real-world tasks like scheduling, coding, …
721325active
mims-harvard/TDC
Therapeutics Data Commons (TDC) is an open-science initiative and Python library providing AI-ready datasets, machine learning tasks, and c…
461276active
pdebench/PDEBench
PDEBench is a benchmark suite for scientific machine learning consisting of code and large ready-to-use datasets of time-dependent partial …
561191active
foospidy/payloads
A curated collection of web attack payloads (XSS, SQLi, CRLF, open redirect, password lists, and more) aggregated from many well-known secu…
323979maintenance
OpenGVLab/ScaleCUA
ScaleCUA is an open-source computer use agent project providing a large-scale cross-platform GUI operation dataset, trained models, and an …
381133active
Sentdex/pygta5
A reboot of Sentdex's project using Python and deep learning to play Grand Theft Auto 5, featuring an AI driver called Charles trained via …
323906maintenance
MinorJerry/WebVoyager
WebVoyager is the official code and dataset for a research paper on an end-to-end web agent powered by large multimodal models that complet…
261122active
ShareGPT4Omni/ShareGPT4Video
ShareGPT4Video is the official implementation of a NeurIPS 2024 paper providing a large-scale video-text dataset (40K GPT4V-generated capti…
241093active
X-EraAI/ActPhysCause-Challenge
ActPhysCause Challenge is a benchmark and dataset for action-conditioned physical and causal world modeling, hosted as Track 0 of the LoViF…
541004active
THUDM/AgentTuning
AgentTuning is a research project from Tsinghua University that instruction-tunes LLMs on multi-task agent interaction trajectories to impr…
271504maintenance
stylegan-human/StyleGAN-Human
StyleGAN-Human is a research project providing a large-scale annotated human image dataset (230K+ samples) and StyleGAN-based models for un…
341189maintenance
openai/gpt-3
The official OpenAI repository accompanying the GPT-3 paper 'Language Models are Few-Shot Learners', containing sample generations, synthet…
1015718abandoned

← prev page 2 / 2