resource: data-generation
175 resources, primary matches first, then adoption-weighted; health v2 shown.
| Resource | Health v2 | Stars | Maturity |
|---|---|---|---|
| felixonmars/fcitx5-pinyin-zhwiki A tool and generated dictionary that builds a Pinyin dictionary for fcitx5-pinyin and Rime input methods from Chinese Wikipedia data. It co… | 63 | 1087 | active |
| eduosi/district A dataset of Chinese administrative divisions (provinces, cities, districts/counties) with names, pinyin, pinyin initials, administrative c… | 51 | 1082 | active |
| cov-lineages/pango-designation The authoritative repository maintaining Pango lineage designations for SARS-CoV-2, hosting the lineage description list, sequence designat… | 94 | 1076 | active |
| KhronosGroup/glTF-Sample-Models A collection of sample glTF 1.0 and 2.0 3D models maintained by Khronos Group for testing and demonstrating the glTF format, including PBR … | 10 | 3615 | maintenance |
| marbl/CHM13 The Telomere-to-Telomere (T2T) consortium's complete telomere-to-telomere assembly of the CHM13hTERT human cell line, including the T2T-CHM… | 73 | 1054 | stable |
| EBazarov/nsfw_data_source_urls A curated collection of text files listing over 1.5 million image URLs across 159 categories, intended for downloading and building NSFW im… | 32 | 3577 | maintenance |
| OpenSourceAP/CrossSection Code and data accompanying Chen and Zimmermann's paper on open source cross-sectional asset pricing. It reproduces dozens of stock-level pr… | 49 | 1034 | active |
| Rivens7/Livelist A curated IPTV live stream source list focused on IPv6, auto-updated and synced from popular Chinese IPTV projects like fanmingming/live an… | 60 | 1022 | active |
| CLUEbenchmark/CLUECorpus2020 CLUECorpus2020 is a 100GB cleaned Chinese text corpus (35 billion characters) derived from Common Crawl for pre-training language models, p… | 62 | 1018 | stable |
| ignis-sec/Pwdb-Public A curated dataset of password statistics and wordlists extracted from over one billion leaked credentials, including filtered lists like a … | 32 | 3286 | maintenance |
| wecatch/china_regions A dataset of China's administrative divisions (province, city, county, town, village) scraped from the national statistics bureau standard,… | 23 | 3155 | maintenance |
| thunlp/UltraChat UltraChat is a large-scale dataset of 1.57M informative, diverse multi-round dialogue instructions generated with LLMs, plus UltraLM, a ser… | 30 | 2890 | maintenance |
| sbousseaden/EVTX-ATTACK-SAMPLES A dataset of ~200 Windows EVTX event log samples mapped to MITRE ATT&CK tactics and techniques, covering attack and post-exploitation behav… | 32 | 2614 | maintenance |
| mahavivo/english-wordlists A collection of curated English wordlists for Chinese learners, including CET-4/6, Taiwan high school 7000 words, TOEFL, GRE, and COCA 2000… | 32 | 2457 | maintenance |
| npm/npm-expansions A community-curated list of humorous three-word expansions of the acronym 'npm', used on the npmjs.com website header and published as an n… | 10 | 2308 | maintenance |
| openai/gpt-2-output-dataset A dataset from OpenAI containing 250K WebText documents alongside 250K samples each from four GPT-2 model sizes, both untruncated and Top-K… | 10 | 2027 | maintenance |
| mercari/engineer-vocabulary-list A curated list of 100 Japanese/English vocabulary words for engineers, collected from real engineering meetings at Mercari and Merpay. Dist… | 32 | 1974 | maintenance |
| google-deepmind/mathematics_dataset A Python library from DeepMind that generates synthetic mathematical question-and-answer pairs at roughly school-level difficulty across to… | 32 | 1965 | maintenance |
| wy-luke/All-Chinese-Character-Set A curated dataset of text files containing complete Chinese character sets, including 3500/7000 common hanzi lists, all characters, and ful… | 69 | 1939 | maintenance |
| airyland/china-area-data A JavaScript npm package providing China's administrative division data (provinces, cities, districts/counties) sourced from the National B… | 32 | 1926 | maintenance |
| bridgetkromhout/devops-against-humanity DevOps Against Humanity is a community-written expansion deck for Cards Against Humanity, distributed as card text in Google Docs, CSVs, an… | 32 | 1848 | maintenance |
| EleutherAI/the-pile The Pile is a large, diverse, open-source language modeling dataset composed of many smaller text sources combined together. This repositor… | 32 | 1673 | maintenance |
| teknium1/GPTeacher GPTeacher is a collection of modular instruction-tuning datasets generated by GPT-4, including General-Instruct, Roleplay-Instruct, Code-In… | 30 | 1667 | maintenance |
| AgaMiko/data-augmentation-review A curated list of data augmentation resources covering libraries, GitHub repos, and papers across images, NLP, audio, time series, graphs, … | 23 | 1639 | maintenance |
| hermitdave/FrequencyWords A repository of preprocessed word frequency lists for many languages, generated from OpenSubtitles corpora, plus the C# generator code. Eac… | 32 | 1535 | maintenance |
| sahil280114/codealpaca Code Alpaca is a 20K instruction-following dataset and training code for fine-tuning LLaMA models on code generation tasks, based on Stanfo… | 30 | 1514 | maintenance |
| PolyAI-LDN/conversational-datasets A collection of tools and scripts from PolyAI for generating large, reproducible datasets for conversational response selection, including … | 32 | 1401 | maintenance |
| KaiDMML/FakeNewsNet FakeNewsNet is a research dataset for fake news detection, with minimal CSV files of news samples from PolitiFact and GossipCop plus Python… | 32 | 1346 | maintenance |
| zxcvbn001/password_brute_dictionary A collection of Chinese-oriented password brute-force dictionaries derived from over 42 million leaked passwords across five platforms. It … | 32 | 1325 | maintenance |
| dansinker/tacofancy A community-driven repository of taco recipes organized in an object-oriented structure (base layers, mixins, condiments, seasonings, shell… | 32 | 1301 | maintenance |
| google-deepmind/rc-data A question answering corpus from DeepMind accompanying the paper 'Teaching Machines to Read and Comprehend' (Hermann et al., NIPS 2015). It… | 10 | 1296 | maintenance |
| jayleicn/animeGAN A simple PyTorch implementation of DCGAN focused on generating anime face images, including a pretrained model and Jupyter notebook demos. … | 23 | 1277 | maintenance |
| hikariming/chat-dataset-baseline A curated Chinese conversational dataset plus fine-tuning scripts for training chat models like ChatGLM, built on top of LLaMA-Factory. It … | 39 | 1191 | maintenance |
| rogerzhu/MNWeeklyCategory A categorized collection of over 10,000 curated links from 327+ issues of the Chinese 'Manong Weekly' (码农周刊) newsletter, organized into mar… | 32 | 1161 | maintenance |
| google-deepmind/funsearch FunSearch is the official repository accompanying DeepMind's Nature 2023 paper on discovering new mathematical results via program search w… | 27 | 1110 | maintenance |
| kaonashi-passwords/Kaonashi Kaonashi is a collection of password wordlists, hashcat rules, and masks derived from large-scale analysis of billions of real leaked passw… | 32 | 1099 | maintenance |
| k8gege/PasswordDic A collection of password dictionaries (wordlists) including top weak passwords from 2011-2019, SSH/VPS server passwords, admin panel passwo… | 32 | 1053 | maintenance |
| 1eez/103976 A dataset of 103,976 English words with Chinese translations, parts of speech, and multiple definitions, distributed as SQL, CSV, and Excel… | 32 | 1053 | maintenance |
| allenai/natural-instructions A community-built dataset of 1600+ NLP tasks with natural language instructions, from AllenAI. It supports instruction-tuning research like… | 23 | 1046 | maintenance |
| unrealcv/synthetic-computer-vision A curated list of synthetic image datasets, 3D model repositories, and rendering tools for computer vision research. It tracks publications… | 32 | 1025 | maintenance |
| ohmybahgosh/RockYou2021.txt RockYou2021.txt is a massive compiled password wordlist of roughly 82 billion unique entries (6-20 ASCII characters), aggregated from multi… | 32 | 1012 | maintenance |
| mshumer/gpt-llm-trainer A set of Jupyter/Colab notebooks that automate the LLM fine-tuning pipeline: given a task description, they generate a synthetic dataset wi… | 37 | 4177 | experimental |
| first20hours/google-10000-english A dataset of the 10,000 most common English words ranked by frequency, derived from n-gram frequency analysis of Google's Trillion Word Cor… | 32 | 4464 | abandoned |
| feeddd/feeds A community-maintained collection of free RSS/Atom/JSON feeds for WeChat official accounts, generated via Hamibot Android automation script… | 32 | 2098 | abandoned |
| tintinweb/smart-contract-sanctuary A large git-based dataset of Etherscan-verified Solidity smart contracts, organized as an index repository with per-chain submodules (Ether… | 66 | 1593 | abandoned |
| FOSSASIA-Web/fossasia-communities A repository of JSON API files describing FOSSASIA and other open source communities across Asia, used to power the FOSSASIA.net community … | 32 | 1585 | abandoned |
| Intel-bigdata/HiBench HiBench is a big data benchmark suite with 29 workloads across micro, machine learning, SQL, graph, websearch, and streaming categories for… | 10 | 1484 | abandoned |
| AlanChen4/Summer-2024-SWE-Internships A curated, auto-updated list of Summer 2024 software engineering internship postings, maintained by a Python job monitor and linked to the … | 10 | 1362 | abandoned |
| robbiebarrat/rapping-neural-network A generative art project: a recurrent neural network trained on Kanye West's discography that writes rap lyrics word by word with rhymes an… | 23 | 1062 | abandoned |
| fashiontec/openwash-format OpenWash is a specification repository from Fashiontec defining a format and specs, presumably for open data exchange in the fashion/washin… | 32 | 1028 | abandoned |
| Hammer1/cozeworkflows A curated collection of 200+ ready-to-import workflow JSON files for the Coze (扣子) AI bot-building platform, maintained by an AI blogger an… | 33 | 5507 | active |
| arman-bd/guppylm GuppyLM is a ~9M parameter language model trained from scratch to talk like a small fish, built as an educational project demonstrating the… | 49 | 3435 | active |
| alex000kim/nsfw_data_scraper A collection of shell scripts that automatically aggregate tens of thousands of images across five categories (porn, hentai, sexy, neutral,… | 32 | 12588 | maintenance |
| BannyLon/DifyAIA An open-source collection of ready-to-import Dify workflow DSL examples (YAML files) with accompanying Flask helper services, created by a … | 63 | 2642 | active |
| OpenCoder-llm/OpenCoder-llm OpenCoder is a fully open and reproducible family of code large language models (1.5B and 8B base and chat variants) trained on 2.5 trillio… | 22 | 2111 | active |
| lensesio/fast-data-dev A Docker image that packages a complete Apache Kafka development environment, including a Kafka broker, Zookeeper/KRaft, Confluent Schema R… | 56 | 2079 | active |
| Thinklab-SJTU/Bench2Drive Bench2Drive is a closed-loop benchmark and dataset for end-to-end autonomous driving, built on CARLA with an RL-based expert driver (Think2… | 68 | 1926 | active |
| legalize-dev/legalize-es A git-versioned dataset of Spanish legislation in Markdown, where each law is a file and each reform is a commit dated with its official pu… | 59 | 1873 | active |
| SocialAI-tianji/Tianji Tianji is an open-source Chinese-language LLM application and tutorial project focused on social nuance ('renqing shigu') scenarios, coveri… | 35 | 1822 | active |
| trickest/wordlists A regularly updated collection of real-world infosec wordlists maintained by Trickest, including technology-specific path lists (WordPress,… | 77 | 1791 | active |
| ubisoft/ubisoft-laforge-animation-dataset The Ubisoft La Forge Animation Dataset (LAFAN1) is a motion capture dataset of 5 subjects, 77 sequences, and ~4.6 hours of character animat… | 32 | 1568 | stable |
| microsoft/DNS-Challenge Microsoft's repository for the Deep Noise Suppression (DNS) Challenge, containing datasets, scripts, and baseline models for training speec… | 32 | 1465 | active |
| ikatsov/tensor-house TensorHouse is a collection of reference Jupyter notebooks and demo AI/ML applications covering enterprise use cases such as marketing, pri… | 23 | 1452 | active |
| pinchbench/skill PinchBench is a benchmarking system that evaluates LLM models as OpenClaw coding agents using 53 real-world tasks like scheduling, coding, … | 72 | 1325 | active |
| mims-harvard/TDC Therapeutics Data Commons (TDC) is an open-science initiative and Python library providing AI-ready datasets, machine learning tasks, and c… | 46 | 1276 | active |
| pdebench/PDEBench PDEBench is a benchmark suite for scientific machine learning consisting of code and large ready-to-use datasets of time-dependent partial … | 56 | 1191 | active |
| foospidy/payloads A curated collection of web attack payloads (XSS, SQLi, CRLF, open redirect, password lists, and more) aggregated from many well-known secu… | 32 | 3979 | maintenance |
| OpenGVLab/ScaleCUA ScaleCUA is an open-source computer use agent project providing a large-scale cross-platform GUI operation dataset, trained models, and an … | 38 | 1133 | active |
| Sentdex/pygta5 A reboot of Sentdex's project using Python and deep learning to play Grand Theft Auto 5, featuring an AI driver called Charles trained via … | 32 | 3906 | maintenance |
| MinorJerry/WebVoyager WebVoyager is the official code and dataset for a research paper on an end-to-end web agent powered by large multimodal models that complet… | 26 | 1122 | active |
| ShareGPT4Omni/ShareGPT4Video ShareGPT4Video is the official implementation of a NeurIPS 2024 paper providing a large-scale video-text dataset (40K GPT4V-generated capti… | 24 | 1093 | active |
| X-EraAI/ActPhysCause-Challenge ActPhysCause Challenge is a benchmark and dataset for action-conditioned physical and causal world modeling, hosted as Track 0 of the LoViF… | 54 | 1004 | active |
| THUDM/AgentTuning AgentTuning is a research project from Tsinghua University that instruction-tunes LLMs on multi-task agent interaction trajectories to impr… | 27 | 1504 | maintenance |
| stylegan-human/StyleGAN-Human StyleGAN-Human is a research project providing a large-scale annotated human image dataset (230K+ samples) and StyleGAN-based models for un… | 34 | 1189 | maintenance |
| openai/gpt-3 The official OpenAI repository accompanying the GPT-3 paper 'Language Models are Few-Shot Learners', containing sample generations, synthet… | 10 | 15718 | abandoned |
← prev page 2 / 2