Ross ROSS = Recommend OSS · open-source software intelligence for agents

resource: data-generation

175 resources, primary matches first, then adoption-weighted; health v2 shown.

ResourceHealth v2StarsMaturity
PicoTrex/Awesome-Nano-Banana-images
A curated gallery of creative image generation and editing examples (with prompts) produced by Google's Nano Banana / Nano Banana Pro (Gemi…
4323574active
hasaneyldrm/exercises-dataset
A curated dataset of 1,324 fitness exercises, each with an animation GIF, 180×180 thumbnail, muscle-group/equipment metadata, and step-by-s…
5620947active
zhaoolee/ChineseBQB
A large open-source collection of over 5,800 Chinese meme/sticker images hosted on GitHub, with an online search and sharing tool. It also …
7216049active
Nano Banana Pro Prompts
A large curated open-source library of 10,000+ prompts for Google's Nano Banana Pro (Gemini) AI image generation model, with preview images…
6113282active
YanG-1989/m3u
A community-maintained collection of IPTV live-stream M3U playlists and subscription URLs, aggregated from public sources and updated autom…
7611377active
ethereum-lists/chains
A community-maintained dataset of metadata for EVM-based blockchain networks, keyed by CAIP-2 chain identifiers. It provides chain IDs, RPC…
779822active
YouMind-OpenLab/awesome-gpt-image-2
A large curated awesome-list of 2000+ prompts for OpenAI's GPT Image 2 image generation model, with preview images and translations in 16 l…
599494active
minimaxir/big-list-of-naughty-strings
A curated list of strings with a high probability of causing issues when used as user input, provided as newline-delimited text and JSON fi…
3247712maintenance
jackvale/rectg
A curated index of 500+ Chinese-language Telegram channels, groups, and bots, organized into 22 topic categories with automated scraping pl…
709182active
snehasishroy/leetcode-companywise-interview-questions
A curated dataset of LeetCode questions categorized by company (Google, Amazon, Meta, Microsoft, etc.) and recency, with difficulty, accept…
767549active
vbskycn/iptv
A continuously updated collection of free IPTV live TV stream sources (M3U and TXT playlists) for Chinese and international channels, auto-…
777547active
xiangyuecn/AreaCity-JsSpider-StatsGov
A dataset and tooling project providing China's province/city/district/town (3-4 level) administrative division data with pinyin, coordinat…
706842active
tatsu-lab/stanford_alpaca
Stanford Alpaca is the code and 52K instruction-following dataset used to fine-tune LLaMA 7B into the Alpaca instruction-following model. I…
3030247maintenance
metowolf/vCards
A curated dataset of Chinese business and service contact vCards (with logos) that users import into iOS/macOS/Android address books via Ca…
986386active
mledoze/countries
An open dataset of world countries per ISO 3166-1, distributed in JSON, CSV, XML, and YAML formats with rich attributes like codes, currenc…
656255active
aoaostar/legado
A curated collection of book sources, subscription feeds, themes, and layout configs for the Legado (阅读) Android reading app, served via a …
766098active
agenda-tech-brasil/agenda-tech-brasil
A community-maintained curated list of technology events happening in Brazil, organized by month and year, with a companion website and eve…
775701active
disposable-email-domains/disposable-email-domains
A community-maintained list of disposable and temporary email address domains, distributed as a plain-text blocklist and a PyPI package. It…
775454active
umpirsky/country-list
A dataset of all countries with names and ISO 3166-1 codes, available in all languages and many data formats (JSON, YAML, XML, CSV, SQL, PH…
675253stable
dariusk/corpora
A collection of small, curated JSON corpora (word lists, names, places, and other categorical data) intended for creative coding, bot creat…
615108stable
CollegesChat/university-information
A crowdsourced dataset collecting undocumented details about Chinese universities that affect student quality of life, such as dorm facilit…
735076active
togethercomputer/RedPajama-Data
RedPajama-Data provides code and pipelines for building RedPajama-V2, an open dataset with over 30 trillion tokens of web text for training…
684980active
apple/password-manager-resources
A collaborative dataset of website-specific 'quirks' for password managers, maintained by Apple and the community. It includes password rul…
774794active
wpzzz/blocked-sites-in-south-korea
A dataset tracking websites blocked by the South Korean government, maintained via Python scripts. It provides a regularly updated list of …
324541active
NVlabs/ffhq-dataset
Flickr-Faces-HQ (FFHQ) is a dataset of 70,000 high-quality 1024x1024 PNG images of human faces, crawled from Flickr and aligned with dlib. …
324182stable
Meroser/IPTV
A curated IPTV playlist repository providing deeply customized M3U channel lists with high-definition TV logos and perfectly matched EPG (e…
634142active
arkadiyt/bounty-targets-data
A dataset repository containing hourly-updated dumps of bug bounty program scopes from platforms like HackerOne, Bugcrowd, Intigriti, YesWe…
773915active
songguoxs/gpt4o-image-prompts
A curated collection of thousands of image-generation prompts for models like Nano Banana Pro, GPT-4o/GPT-5, and Grok, stored as Markdown c…
473791active
hampusborgos/country-flags
A collection of accurate SVG and PNG renders of all countries' flags, organized by ISO-3166 country codes and available as an npm module (s…
643784active
gaoyifan/china-operator-ip
A daily-updated dataset of IPv4/IPv6 CIDR lists for Chinese network operators (China Telecom, China Mobile, China Unicom, CERNET, CSTNet, e…
773613active
LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words
A multilingual list of profane and obscene words maintained by Shutterstock, used to filter autocomplete suggestions and recommendations. I…
323431active
six2dez/OneListForAll
A curated collection of wordlists for web fuzzing, aggregating ~36 source repositories into categorized short/long lists and combined 'onel…
673231active
mumuy/data_location
An open dataset of Chinese administrative division codes (GB/T 2260) covering provinces, cities, districts/counties, and towns/townships, d…
703177active
openmaptiles/openmaptiles
OpenMapTiles is an open vector tile schema and tooling for generating zoomable map tiles from OpenStreetMap and other open data sources. It…
733144active
uiwjs/province-city-china
A dataset and npm package containing complete, regularly updated China administrative division codes (GB/T 2260) for provinces, cities, cou…
583039active
meodai/color-names
A curated open dataset of over 31,000 unique color names mapped to hex values, maintained by a community with automated quality checks. It …
952982active
gfriends/gfriends
A community-maintained repository of 100,000+ adult film actor avatar images for media servers, with automated collection and processing of…
672857active
WebBreacher/WhatsMyName
WhatsMyName is a community-maintained JSON dataset of 700+ websites with rules for checking whether a username exists on each site. It powe…
762803active
iamcal/emoji-data
A dataset of easy-to-parse emoji metadata (emoji.) plus spritesheet-style images for use on the web, supporting Emoji version 16.0. Distrib…
682798active
skishore/makemeahanzi
Make Me a Hanzi is a free, open-source dataset of dictionary and graphical data for over 9000 common simplified and traditional Chinese cha…
642637stable
lerocha/chinook-database
Chinook is a sample database representing a digital media store, available as SQL scripts for SQL Server, Oracle, MySQL, PostgreSQL, SQLite…
432589active
adyliu/china_area
A dataset of China's five-level administrative divisions (province, city, county, town, village) for 2024, sourced from the National Bureau…
232504active
nusquama/n8nworkflows.xyz
A versioned archive of over 9,600 n8n workflow templates scraped from the official n8n.io/workflows site, stored as importable JSON files w…
612490active
a312863063/generators-with-stylegan2
A collection of pretrained StyleGAN2-based face generators producing various styles of synthetic human faces (influencer, celebrity, model,…
352489stable
itmeo/webgradients
A curated collection of 180 free CSS3 gradients for backgrounds and UI, distributed as a CSS file plus Sketch, Photoshop, and Figma formats…
752472stable
berzerk0/Probable-Wordlists
A collection of password wordlists sorted by probability of use, derived from real-world data breaches. It is a data resource (not code) in…
239334maintenance
open-thoughts/open-thoughts
OpenThoughts is a community project curating fully open datasets for training reasoning models, including OpenThoughts3-1.2M and OpenThinke…
452323active
ldrolez/free-midi-chords
A free collection of over 13,000 MIDI files containing chords and chord progressions in all keys, organized by key, chord type, and mood ta…
742302active
qwerttvv/Beijing-IPTV
A curated collection of M3U IPTV channel playlists for Beijing Unicom and Beijing Mobile, with built-in EPG references and multiple mirror …
672274active
zhaoolee/ins
A curated, ad-free database of inspiring websites for internet professionals, stored as CSV and rendered as a browsable catalog. GitHub Act…
772254active
ise-uiuc/magicoder
Magicoder is a family of fully open-source code LLMs (under 7B parameters) trained with OSS-Instruct, a method that seeds LLMs with open-so…
272094stable
apple/ml-hypersim
Hypersim is a photorealistic synthetic dataset of 74,619 images across 461 indoor scenes with dense per-pixel semantic instance segmentatio…
602043stable
ruankaodaren/ruankao
A free, regularly updated question bank for China's national software qualification exams (软考), covering all advanced, intermediate, and be…
551941active
gaitco/quran-database
A complete, structured Quran database distributed as MySQL, PostgreSQL, and SQLite dumps containing the full Arabic text, multiple translat…
771904active
YouMind-OpenLab/awesome-seedance-2-prompts
A curated awesome-list of 2000+ video generation prompts for ByteDance's Seedance 2.0 model, covering cinematic, anime, UGC, advertising, a…
611898active
bytedance/GiantMIDI-Piano
GiantMIDI-Piano is a classical piano MIDI dataset containing 10,855 transcribed MIDI files from 2,786 composers, generated by transcribing …
101895stable
KyleBing/english-vocabulary
A curated English vocabulary dataset of 54,356 words with Chinese translations, organized by level (middle school, high school, CET4, CET6,…
761884active
pjy612/SteamManifestCache
A community-maintained repository caching Steam depot manifest files, organized by AppId branches and manifest filename tags, with addition…
631870active
hkust-nlp/ceval
C-Eval is a comprehensive Chinese evaluation benchmark for foundation models, consisting of 13,948 multiple-choice questions across 52 disc…
441867stable
apple/pico-banana-400k
A large-scale dataset of ~400K text-image-edit triplets for text-guided image editing research, built from Open Images using Nano-Banana ed…
421845active
jeanphorn/wordlist
A curated collection of password wordlists, username lists, and default credentials (SSH, RDP, FTP, databases, IoT, IP cameras) for authori…
681801active
duyet/bruteforce-database
A curated collection of wordlists (11+ million entries) for password cracking, username enumeration, subdomain discovery, and web path brut…
741726active
R6410418/Jackrong-llm-finetuning-guide
An open-source educational knowledge base covering LLM fine-tuning, dataset distillation, reinforcement learning workflows (SFT, GRPO, GSPO…
561665active
LaQuay/TDTChannels
A curated, open-source catalog of free and legal over-the-air Spanish and international TV and radio channels that stream online, published…
771652active
Hipo/university-domains-list
A curated JSON dataset of world universities with their domain names, countries, and web pages, plus a small Python tooling layer and a fre…
771649active
tencent-ailab/persona-hub
PersonaHub is Tencent AI Lab's collection of 1 billion diverse personas plus code for persona-driven synthetic data creation with LLMs. It …
271642active
Trust Wallet
A community-maintained repository of metadata for thousands of crypto tokens across 100+ blockchains, including logos, info. files, dApp in…
101608active
alecjacobson/common-3d-test-models
A repository collecting common 3D test models (e.g., Armadillo, Stanford Bunny, Happy Buddha) in their original formats alongside ~10MB OBJ…
321605active
stefangabos/world_countries
A constantly updated dataset of world countries, territories, and their ISO 3166-1 alpha-2, alpha-3, and numeric codes, plus ISO 3166-2 sub…
811592active
SexyBeast233/SecDictionary
A curated collection of wordlists and dictionaries for security testing, built from real-world penetration testing experience. It includes …
751581active
google-research/FLAN
Google Research's repository for generating the FLAN instruction tuning dataset collections, including the original Flan 2021 and the expan…
731567stable
wasiahmad/Awesome-LLM-Synthetic-Data
A curated awesome-list of papers, tools, and blogs about synthetic data generation with large language models. It organizes resources by su…
341551active
EricGuo5513/HumanML3D
HumanML3D is a large 3D human motion-language dataset with 14,616 motion clips and 44,970 text descriptions, built from HumanAct12 and AMAS…
321522stable
op7418/guizang-s-prompt
A curated collection of AI prompts written by Guizang, covering image, text, and video generation use cases. Prompts are stored as markdown…
431444active
initstring/passphrase-wordlist
A large wordlist of over 20 million passphrase phrases paired with two hashcat rule files that generate 1,000+ permutations per phrase for …
371440active
disposable/disposable
A regularly updated dataset of disposable/temporary email address domains (like 10MinuteMail and GuerrillaMail), provided as plain text lis…
751438active
cjh0613/tencent-sensitive-words
An offline dataset of sensitive words extracted from Tencent's software, published as a wordlist repository. It is intended for content mod…
101423active
carlospolop/Auto_Wordlists
A repository of automatically generated security wordlists for web fuzzing, reconnaissance, and payload testing, refreshed on a schedule (D…
761403active
monperrus/crawler-user-agents
A curated JSON dataset of regular-expression patterns matching HTTP user-agents used by bots, crawlers, spiders, and scrapers. It is distri…
751400active
zapret-info/z-i
A community-maintained register of internet addresses (IPs and subnets) filtered/blocked in the Russian Federation. It provides regularly u…
521387active
insidetrust/statistically-likely-usernames
A collection of wordlists for generating statistically likely usernames for use in password attacks, username enumeration, and authorized s…
621386stable
KanjiVG/kanjivg
KanjiVG is an open dataset of SVG vector graphics for Japanese kanji, encoding stroke shapes, stroke order, radicals, and component informa…
691364active
projectdiscovery/public-bugbounty-programs
A community-curated dataset of public bug bounty and responsible disclosure programs, maintained as YAML with a JSON schema and distributed…
761341active
Open-Source-O1/Open-O1
Open O1 is an open-source effort to replicate the reasoning capabilities of OpenAI's proprietary O1 model by curating chain-of-thought SFT …
221340active
manami-project/anime-offline-database
A JSON-based offline anime metadata dataset aggregating entries from multiple providers like MyAnimeList, AniDB, AniList, Kitsu, and more, …
101325active
kelvins/municipios-brasileiros
A dataset of all 5,570 Brazilian municipalities with IBGE codes, names, UF/state, coordinates, SIAFI codes, DDD, and timezone, distributed …
571268active
msikma/pokesprite
A database of Pokémon box sprites and inventory item sprites from the core series games, including custom shiny variants, with metadata fil…
321265active
awesome-assistants/awesome-assistants
A curated list of 240+ AI assistant definitions (system prompts) maintained in assistants.yml and exported to JSON, CSV, TSV, and HTML. It …
271259active
ZGQ-inc/source
A large curated collection of importable sources for Chinese reading and media apps, including ~28,000 book sources for Legado (开源阅读), RSS …
411257active
joshuafuller/ATAK-Maps
A curated collection of 40 MOBAC-format XML map source files for the Android Team Awareness Kit (ATAK), covering satellite, topographic, na…
981249active
Instruction-Tuning-with-GPT-4/GPT-4-LLM
A research dataset release of GPT-4-generated instruction-following data for fine-tuning large language models, including English and Chine…
304334maintenance
wainshine/Chinese-Names-Corpus
A curated corpus of Chinese, Japanese, and translated English personal names, plus surnames, kinship terms, and idioms, totaling millions o…
464327maintenance
Activision/caldera
A large OpenUSD scene dataset containing converted geometry from the Call of Duty: Warzone Caldera map, released by Activision for academic…
221232stable
Profluent-AI/OpenCRISPR
OpenCRISPR is a collection of AI-designed gene editing systems released by Profluent Bio, including OpenCRISPR-1, a Cas9-like protein and g…
441204active
codemayq/chinese-chatbot-corpus
A curated collection and unified processing pipeline for publicly available Chinese chit-chat conversation corpora, aggregating eight sourc…
324194maintenance
wesbos/burner-email-providers
A community-maintained list of disposable (burner) email provider domains, useful for filtering throwaway addresses from signup forms and m…
741189active
samayo/country-json
A collection of world country reference data published as simple JSON files, covering attributes like names, capitals, currencies, calling …
511148stable
Lascorbe/CocoaConferences
A community-maintained list of conferences for Apple-ecosystem developers (iOS, macOS, watchOS, tvOS, visionOS), published as a website at …
761134active
unicode-org/cldr
The Unicode Common Locale Data Repository (CLDR) is the largest standard repository of locale data supporting software internationalization…
871133active
evansiroky/timezone-boundary-builder
A tool and dataset project that builds the world's timezone boundaries from OpenStreetMap data, released as shapefiles and GeoJSON. Each bo…
891114active

page 1 / 2 next →