resource: data-generation
175 resources, primary matches first, then adoption-weighted; health v2 shown.
| Resource | Health v2 | Stars | Maturity |
|---|---|---|---|
| PicoTrex/Awesome-Nano-Banana-images A curated gallery of creative image generation and editing examples (with prompts) produced by Google's Nano Banana / Nano Banana Pro (Gemi… | 43 | 23574 | active |
| hasaneyldrm/exercises-dataset A curated dataset of 1,324 fitness exercises, each with an animation GIF, 180×180 thumbnail, muscle-group/equipment metadata, and step-by-s… | 56 | 20947 | active |
| zhaoolee/ChineseBQB A large open-source collection of over 5,800 Chinese meme/sticker images hosted on GitHub, with an online search and sharing tool. It also … | 72 | 16049 | active |
| Nano Banana Pro Prompts A large curated open-source library of 10,000+ prompts for Google's Nano Banana Pro (Gemini) AI image generation model, with preview images… | 61 | 13282 | active |
| YanG-1989/m3u A community-maintained collection of IPTV live-stream M3U playlists and subscription URLs, aggregated from public sources and updated autom… | 76 | 11377 | active |
| ethereum-lists/chains A community-maintained dataset of metadata for EVM-based blockchain networks, keyed by CAIP-2 chain identifiers. It provides chain IDs, RPC… | 77 | 9822 | active |
| YouMind-OpenLab/awesome-gpt-image-2 A large curated awesome-list of 2000+ prompts for OpenAI's GPT Image 2 image generation model, with preview images and translations in 16 l… | 59 | 9494 | active |
| minimaxir/big-list-of-naughty-strings A curated list of strings with a high probability of causing issues when used as user input, provided as newline-delimited text and JSON fi… | 32 | 47712 | maintenance |
| jackvale/rectg A curated index of 500+ Chinese-language Telegram channels, groups, and bots, organized into 22 topic categories with automated scraping pl… | 70 | 9182 | active |
| snehasishroy/leetcode-companywise-interview-questions A curated dataset of LeetCode questions categorized by company (Google, Amazon, Meta, Microsoft, etc.) and recency, with difficulty, accept… | 76 | 7549 | active |
| vbskycn/iptv A continuously updated collection of free IPTV live TV stream sources (M3U and TXT playlists) for Chinese and international channels, auto-… | 77 | 7547 | active |
| xiangyuecn/AreaCity-JsSpider-StatsGov A dataset and tooling project providing China's province/city/district/town (3-4 level) administrative division data with pinyin, coordinat… | 70 | 6842 | active |
| tatsu-lab/stanford_alpaca Stanford Alpaca is the code and 52K instruction-following dataset used to fine-tune LLaMA 7B into the Alpaca instruction-following model. I… | 30 | 30247 | maintenance |
| metowolf/vCards A curated dataset of Chinese business and service contact vCards (with logos) that users import into iOS/macOS/Android address books via Ca… | 98 | 6386 | active |
| mledoze/countries An open dataset of world countries per ISO 3166-1, distributed in JSON, CSV, XML, and YAML formats with rich attributes like codes, currenc… | 65 | 6255 | active |
| aoaostar/legado A curated collection of book sources, subscription feeds, themes, and layout configs for the Legado (阅读) Android reading app, served via a … | 76 | 6098 | active |
| agenda-tech-brasil/agenda-tech-brasil A community-maintained curated list of technology events happening in Brazil, organized by month and year, with a companion website and eve… | 77 | 5701 | active |
| disposable-email-domains/disposable-email-domains A community-maintained list of disposable and temporary email address domains, distributed as a plain-text blocklist and a PyPI package. It… | 77 | 5454 | active |
| umpirsky/country-list A dataset of all countries with names and ISO 3166-1 codes, available in all languages and many data formats (JSON, YAML, XML, CSV, SQL, PH… | 67 | 5253 | stable |
| dariusk/corpora A collection of small, curated JSON corpora (word lists, names, places, and other categorical data) intended for creative coding, bot creat… | 61 | 5108 | stable |
| CollegesChat/university-information A crowdsourced dataset collecting undocumented details about Chinese universities that affect student quality of life, such as dorm facilit… | 73 | 5076 | active |
| togethercomputer/RedPajama-Data RedPajama-Data provides code and pipelines for building RedPajama-V2, an open dataset with over 30 trillion tokens of web text for training… | 68 | 4980 | active |
| apple/password-manager-resources A collaborative dataset of website-specific 'quirks' for password managers, maintained by Apple and the community. It includes password rul… | 77 | 4794 | active |
| wpzzz/blocked-sites-in-south-korea A dataset tracking websites blocked by the South Korean government, maintained via Python scripts. It provides a regularly updated list of … | 32 | 4541 | active |
| NVlabs/ffhq-dataset Flickr-Faces-HQ (FFHQ) is a dataset of 70,000 high-quality 1024x1024 PNG images of human faces, crawled from Flickr and aligned with dlib. … | 32 | 4182 | stable |
| Meroser/IPTV A curated IPTV playlist repository providing deeply customized M3U channel lists with high-definition TV logos and perfectly matched EPG (e… | 63 | 4142 | active |
| arkadiyt/bounty-targets-data A dataset repository containing hourly-updated dumps of bug bounty program scopes from platforms like HackerOne, Bugcrowd, Intigriti, YesWe… | 77 | 3915 | active |
| songguoxs/gpt4o-image-prompts A curated collection of thousands of image-generation prompts for models like Nano Banana Pro, GPT-4o/GPT-5, and Grok, stored as Markdown c… | 47 | 3791 | active |
| hampusborgos/country-flags A collection of accurate SVG and PNG renders of all countries' flags, organized by ISO-3166 country codes and available as an npm module (s… | 64 | 3784 | active |
| gaoyifan/china-operator-ip A daily-updated dataset of IPv4/IPv6 CIDR lists for Chinese network operators (China Telecom, China Mobile, China Unicom, CERNET, CSTNet, e… | 77 | 3613 | active |
| LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words A multilingual list of profane and obscene words maintained by Shutterstock, used to filter autocomplete suggestions and recommendations. I… | 32 | 3431 | active |
| six2dez/OneListForAll A curated collection of wordlists for web fuzzing, aggregating ~36 source repositories into categorized short/long lists and combined 'onel… | 67 | 3231 | active |
| mumuy/data_location An open dataset of Chinese administrative division codes (GB/T 2260) covering provinces, cities, districts/counties, and towns/townships, d… | 70 | 3177 | active |
| openmaptiles/openmaptiles OpenMapTiles is an open vector tile schema and tooling for generating zoomable map tiles from OpenStreetMap and other open data sources. It… | 73 | 3144 | active |
| uiwjs/province-city-china A dataset and npm package containing complete, regularly updated China administrative division codes (GB/T 2260) for provinces, cities, cou… | 58 | 3039 | active |
| meodai/color-names A curated open dataset of over 31,000 unique color names mapped to hex values, maintained by a community with automated quality checks. It … | 95 | 2982 | active |
| gfriends/gfriends A community-maintained repository of 100,000+ adult film actor avatar images for media servers, with automated collection and processing of… | 67 | 2857 | active |
| WebBreacher/WhatsMyName WhatsMyName is a community-maintained JSON dataset of 700+ websites with rules for checking whether a username exists on each site. It powe… | 76 | 2803 | active |
| iamcal/emoji-data A dataset of easy-to-parse emoji metadata (emoji.) plus spritesheet-style images for use on the web, supporting Emoji version 16.0. Distrib… | 68 | 2798 | active |
| skishore/makemeahanzi Make Me a Hanzi is a free, open-source dataset of dictionary and graphical data for over 9000 common simplified and traditional Chinese cha… | 64 | 2637 | stable |
| lerocha/chinook-database Chinook is a sample database representing a digital media store, available as SQL scripts for SQL Server, Oracle, MySQL, PostgreSQL, SQLite… | 43 | 2589 | active |
| adyliu/china_area A dataset of China's five-level administrative divisions (province, city, county, town, village) for 2024, sourced from the National Bureau… | 23 | 2504 | active |
| nusquama/n8nworkflows.xyz A versioned archive of over 9,600 n8n workflow templates scraped from the official n8n.io/workflows site, stored as importable JSON files w… | 61 | 2490 | active |
| a312863063/generators-with-stylegan2 A collection of pretrained StyleGAN2-based face generators producing various styles of synthetic human faces (influencer, celebrity, model,… | 35 | 2489 | stable |
| itmeo/webgradients A curated collection of 180 free CSS3 gradients for backgrounds and UI, distributed as a CSS file plus Sketch, Photoshop, and Figma formats… | 75 | 2472 | stable |
| berzerk0/Probable-Wordlists A collection of password wordlists sorted by probability of use, derived from real-world data breaches. It is a data resource (not code) in… | 23 | 9334 | maintenance |
| open-thoughts/open-thoughts OpenThoughts is a community project curating fully open datasets for training reasoning models, including OpenThoughts3-1.2M and OpenThinke… | 45 | 2323 | active |
| ldrolez/free-midi-chords A free collection of over 13,000 MIDI files containing chords and chord progressions in all keys, organized by key, chord type, and mood ta… | 74 | 2302 | active |
| qwerttvv/Beijing-IPTV A curated collection of M3U IPTV channel playlists for Beijing Unicom and Beijing Mobile, with built-in EPG references and multiple mirror … | 67 | 2274 | active |
| zhaoolee/ins A curated, ad-free database of inspiring websites for internet professionals, stored as CSV and rendered as a browsable catalog. GitHub Act… | 77 | 2254 | active |
| ise-uiuc/magicoder Magicoder is a family of fully open-source code LLMs (under 7B parameters) trained with OSS-Instruct, a method that seeds LLMs with open-so… | 27 | 2094 | stable |
| apple/ml-hypersim Hypersim is a photorealistic synthetic dataset of 74,619 images across 461 indoor scenes with dense per-pixel semantic instance segmentatio… | 60 | 2043 | stable |
| ruankaodaren/ruankao A free, regularly updated question bank for China's national software qualification exams (软考), covering all advanced, intermediate, and be… | 55 | 1941 | active |
| gaitco/quran-database A complete, structured Quran database distributed as MySQL, PostgreSQL, and SQLite dumps containing the full Arabic text, multiple translat… | 77 | 1904 | active |
| YouMind-OpenLab/awesome-seedance-2-prompts A curated awesome-list of 2000+ video generation prompts for ByteDance's Seedance 2.0 model, covering cinematic, anime, UGC, advertising, a… | 61 | 1898 | active |
| bytedance/GiantMIDI-Piano GiantMIDI-Piano is a classical piano MIDI dataset containing 10,855 transcribed MIDI files from 2,786 composers, generated by transcribing … | 10 | 1895 | stable |
| KyleBing/english-vocabulary A curated English vocabulary dataset of 54,356 words with Chinese translations, organized by level (middle school, high school, CET4, CET6,… | 76 | 1884 | active |
| pjy612/SteamManifestCache A community-maintained repository caching Steam depot manifest files, organized by AppId branches and manifest filename tags, with addition… | 63 | 1870 | active |
| hkust-nlp/ceval C-Eval is a comprehensive Chinese evaluation benchmark for foundation models, consisting of 13,948 multiple-choice questions across 52 disc… | 44 | 1867 | stable |
| apple/pico-banana-400k A large-scale dataset of ~400K text-image-edit triplets for text-guided image editing research, built from Open Images using Nano-Banana ed… | 42 | 1845 | active |
| jeanphorn/wordlist A curated collection of password wordlists, username lists, and default credentials (SSH, RDP, FTP, databases, IoT, IP cameras) for authori… | 68 | 1801 | active |
| duyet/bruteforce-database A curated collection of wordlists (11+ million entries) for password cracking, username enumeration, subdomain discovery, and web path brut… | 74 | 1726 | active |
| R6410418/Jackrong-llm-finetuning-guide An open-source educational knowledge base covering LLM fine-tuning, dataset distillation, reinforcement learning workflows (SFT, GRPO, GSPO… | 56 | 1665 | active |
| LaQuay/TDTChannels A curated, open-source catalog of free and legal over-the-air Spanish and international TV and radio channels that stream online, published… | 77 | 1652 | active |
| Hipo/university-domains-list A curated JSON dataset of world universities with their domain names, countries, and web pages, plus a small Python tooling layer and a fre… | 77 | 1649 | active |
| tencent-ailab/persona-hub PersonaHub is Tencent AI Lab's collection of 1 billion diverse personas plus code for persona-driven synthetic data creation with LLMs. It … | 27 | 1642 | active |
| Trust Wallet A community-maintained repository of metadata for thousands of crypto tokens across 100+ blockchains, including logos, info. files, dApp in… | 10 | 1608 | active |
| alecjacobson/common-3d-test-models A repository collecting common 3D test models (e.g., Armadillo, Stanford Bunny, Happy Buddha) in their original formats alongside ~10MB OBJ… | 32 | 1605 | active |
| stefangabos/world_countries A constantly updated dataset of world countries, territories, and their ISO 3166-1 alpha-2, alpha-3, and numeric codes, plus ISO 3166-2 sub… | 81 | 1592 | active |
| SexyBeast233/SecDictionary A curated collection of wordlists and dictionaries for security testing, built from real-world penetration testing experience. It includes … | 75 | 1581 | active |
| google-research/FLAN Google Research's repository for generating the FLAN instruction tuning dataset collections, including the original Flan 2021 and the expan… | 73 | 1567 | stable |
| wasiahmad/Awesome-LLM-Synthetic-Data A curated awesome-list of papers, tools, and blogs about synthetic data generation with large language models. It organizes resources by su… | 34 | 1551 | active |
| EricGuo5513/HumanML3D HumanML3D is a large 3D human motion-language dataset with 14,616 motion clips and 44,970 text descriptions, built from HumanAct12 and AMAS… | 32 | 1522 | stable |
| op7418/guizang-s-prompt A curated collection of AI prompts written by Guizang, covering image, text, and video generation use cases. Prompts are stored as markdown… | 43 | 1444 | active |
| initstring/passphrase-wordlist A large wordlist of over 20 million passphrase phrases paired with two hashcat rule files that generate 1,000+ permutations per phrase for … | 37 | 1440 | active |
| disposable/disposable A regularly updated dataset of disposable/temporary email address domains (like 10MinuteMail and GuerrillaMail), provided as plain text lis… | 75 | 1438 | active |
| cjh0613/tencent-sensitive-words An offline dataset of sensitive words extracted from Tencent's software, published as a wordlist repository. It is intended for content mod… | 10 | 1423 | active |
| carlospolop/Auto_Wordlists A repository of automatically generated security wordlists for web fuzzing, reconnaissance, and payload testing, refreshed on a schedule (D… | 76 | 1403 | active |
| monperrus/crawler-user-agents A curated JSON dataset of regular-expression patterns matching HTTP user-agents used by bots, crawlers, spiders, and scrapers. It is distri… | 75 | 1400 | active |
| zapret-info/z-i A community-maintained register of internet addresses (IPs and subnets) filtered/blocked in the Russian Federation. It provides regularly u… | 52 | 1387 | active |
| insidetrust/statistically-likely-usernames A collection of wordlists for generating statistically likely usernames for use in password attacks, username enumeration, and authorized s… | 62 | 1386 | stable |
| KanjiVG/kanjivg KanjiVG is an open dataset of SVG vector graphics for Japanese kanji, encoding stroke shapes, stroke order, radicals, and component informa… | 69 | 1364 | active |
| projectdiscovery/public-bugbounty-programs A community-curated dataset of public bug bounty and responsible disclosure programs, maintained as YAML with a JSON schema and distributed… | 76 | 1341 | active |
| Open-Source-O1/Open-O1 Open O1 is an open-source effort to replicate the reasoning capabilities of OpenAI's proprietary O1 model by curating chain-of-thought SFT … | 22 | 1340 | active |
| manami-project/anime-offline-database A JSON-based offline anime metadata dataset aggregating entries from multiple providers like MyAnimeList, AniDB, AniList, Kitsu, and more, … | 10 | 1325 | active |
| kelvins/municipios-brasileiros A dataset of all 5,570 Brazilian municipalities with IBGE codes, names, UF/state, coordinates, SIAFI codes, DDD, and timezone, distributed … | 57 | 1268 | active |
| msikma/pokesprite A database of Pokémon box sprites and inventory item sprites from the core series games, including custom shiny variants, with metadata fil… | 32 | 1265 | active |
| awesome-assistants/awesome-assistants A curated list of 240+ AI assistant definitions (system prompts) maintained in assistants.yml and exported to JSON, CSV, TSV, and HTML. It … | 27 | 1259 | active |
| ZGQ-inc/source A large curated collection of importable sources for Chinese reading and media apps, including ~28,000 book sources for Legado (开源阅读), RSS … | 41 | 1257 | active |
| joshuafuller/ATAK-Maps A curated collection of 40 MOBAC-format XML map source files for the Android Team Awareness Kit (ATAK), covering satellite, topographic, na… | 98 | 1249 | active |
| Instruction-Tuning-with-GPT-4/GPT-4-LLM A research dataset release of GPT-4-generated instruction-following data for fine-tuning large language models, including English and Chine… | 30 | 4334 | maintenance |
| wainshine/Chinese-Names-Corpus A curated corpus of Chinese, Japanese, and translated English personal names, plus surnames, kinship terms, and idioms, totaling millions o… | 46 | 4327 | maintenance |
| Activision/caldera A large OpenUSD scene dataset containing converted geometry from the Call of Duty: Warzone Caldera map, released by Activision for academic… | 22 | 1232 | stable |
| Profluent-AI/OpenCRISPR OpenCRISPR is a collection of AI-designed gene editing systems released by Profluent Bio, including OpenCRISPR-1, a Cas9-like protein and g… | 44 | 1204 | active |
| codemayq/chinese-chatbot-corpus A curated collection and unified processing pipeline for publicly available Chinese chit-chat conversation corpora, aggregating eight sourc… | 32 | 4194 | maintenance |
| wesbos/burner-email-providers A community-maintained list of disposable (burner) email provider domains, useful for filtering throwaway addresses from signup forms and m… | 74 | 1189 | active |
| samayo/country-json A collection of world country reference data published as simple JSON files, covering attributes like names, capitals, currencies, calling … | 51 | 1148 | stable |
| Lascorbe/CocoaConferences A community-maintained list of conferences for Apple-ecosystem developers (iOS, macOS, watchOS, tvOS, visionOS), published as a website at … | 76 | 1134 | active |
| unicode-org/cldr The Unicode Common Locale Data Repository (CLDR) is the largest standard repository of locale data supporting software internationalization… | 87 | 1133 | active |
| evansiroky/timezone-boundary-builder A tool and dataset project that builds the world's timezone boundaries from OpenStreetMap data, released as shapefiles and GeoJSON. Each bo… | 89 | 1114 | active |
page 1 / 2 next →