# dariusk/corpora

A collection of small corpuses of interesting data for the creation of bots and similar stuff.

Repository: https://github.com/dariusk/corpora
Canonical: https://ross.abutalabs.com/products/corpora
Language: JavaScript
License Family: other
Topics: language, words, bots, corpus
Last push: 2026-01-19T09:43:12+00:00

## Health v2 (maintenance only)
Score: 61/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 63, release rhythm 35, longevity 100
- inputs: {"age_days": 4575, "days_push": 226, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 5108, forks 1292 (observed 2026-08-28T04:09:09.877283+00:00)

## What it is
A collection of small, curated JSON corpora (word lists, names, places, and other categorical data) intended for creative coding, bot creation, and rapid prototyping. It is language-neutral static data released under a CC0 license, designed for teaching and quick experiments rather than exhaustive reference data.

## Use cases
- get a list of common adjectives or nouns for a twitter bot
- find sample word data for a generative art project
- teach a workshop on making bots without scraping data first
- prototype an idea quickly with small curated datasets
- seed a random name or word generator
- find small JSON datasets for creative writing tools

## When to choose
- you need small, curated word or category lists for prototyping or creative projects
- you want public-domain (CC0) data usable in any language
- you are teaching and need ready-made interesting data for students

## When to avoid
- you need exhaustive dictionaries or complete linguistic databases with metadata
- you need an API with query capabilities rather than static files
- you need large-scale or production-grade datasets

## Facets
- artifact type: dataset
- maturity: stable
- function: data-generation, nlp, developer-tools
- domain: developer-tools, education
- platform: cross-platform
- tags: corpus, -data, creative-coding, bots, word-lists, cc0, prototyping, natural-language-processing

## Member repositories
- dariusk/corpora (main) score 61

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:09:09.877283+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:02:09.924756+00:00, confidence not recorded.
  - readme: https://github.com/dariusk/corpora (fetched 2026-08-28T04:09:09.877283+00:00, sha 5b415c8580a1)
- Data as of 2026-08-30T08:39:29.467469+00:00.
