ConardLi/easy-dataset
A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval observed · 2026-08-28
Health v2 · maintenance only
71/100
- Activity 80
- Release rhythm 78
- Longevity 39
Flags: no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 8.5
- age_days: 547
- days_rel: 146
- days_push: 124
- n_releases_24m: 31
Adoption not part of the score
14833 stars · 1524 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Easy Dataset is a self-hosted web application for building high-quality structured datasets for LLM fine-tuning, RAG, and model evaluation. It converts domain documents (PDF, Markdown, DOCX, EPUB, etc.) into question-answer datasets via intelligent segmentation, AI-assisted generation, labeling, export, and evaluation workflows.
Use cases
- create fine-tuning datasets from domain documents
- generate QA pairs from PDFs for LLM training
- build evaluation test sets for vertical domain models
- convert datasets between fine-tuning formats
- manage and label large batches of generated questions
- evaluate RAG recall and post-fine-tune model performance
- construct COT reasoning data for fine-tuning
When to choose
- you need to turn domain documents into structured training or eval datasets
- you want a GUI-driven pipeline covering parsing, chunking, generation, labeling, and export
- you need dataset formats for common fine-tuning frameworks and RAG evaluation
When to avoid
- you only need a small one-off script to transform existing JSONL data
- you require a fully automated headless data pipeline without a UI
- you need a permissively licensed library to embed in closed-source products (AGPL-3.0)
Facets
application · maturity active
data-generation etl rag llm-training prompt-engineering pdf nlp large-language-models machine-learning data-science artificial-intelligence self-hosted cross-platform fine-tuning-datasets dataset-creation document-parsing text-chunking qa-generation model-evaluation data-labeling agpl retrieval-augmented-generation web-server docker nodejs
4 sources
- readme: https://github.com/ConardLi/easy-dataset · fetched 2026-08-28 · 05d9dce95684
- homepage: https://docs.easy-dataset.com · fetched 2026-08-29 · c2233544658e
- site_page: https://docs.easy-dataset.com/ji-chu-gong-neng/quickstart · fetched 2026-08-29 · 50d8514d9283
- site_page: https://docs.easy-dataset.com/ji-chu-gong-neng/publish-your-docs · fetched 2026-08-29 · 78585e972f97
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| ConardLi/easy-dataset | main | 71 |
For agents
markdown · JSON · MCP: product_card(name="ConardLi/easy-dataset")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem