DataArcTech/DataArc-SynData-Toolkit
Synthetic Data Generation Platform By DataArcTech observed · 2026-08-28
Health v2 · maintenance only
57/100
- Activity 90
- Release rhythm 35
- Longevity 20
Flags: no_releases no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 285
- days_rel: n/a
- days_push: 64
- n_releases_24m: 0
Adoption not part of the score
1777 stars · 72 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
DataArc SynData Toolkit is a Python-based synthetic data generation platform for creating customized LLM training data from local corpora, Huggingface sources, or model distillation via simple configuration files. It offers both CLI and GUI interfaces and integrates end-to-end post-training workflows (SFT/GRPO via verl) with model evaluation using DeepEval.
Use cases
- generate synthetic training data for llm fine-tuning
- distill data from a teacher model to train a smaller model
- create training datasets for low-resource languages
- crawl and filter datasets from huggingface for synthesis
- run sft or grpo post-training on generated data
- evaluate fine-tuned models automatically
- generate synthetic data without writing code
When to choose
- you need an end-to-end pipeline from data synthesis to training and evaluation
- you want zero-code or config-driven synthetic data generation
- you need multilingual or low-resource-language training data
- you want flexible model provider support including local and OpenAI APIs
When to avoid
- you need a permissively licensed tool - the repo has no license, so usage rights are unclear
- you only need simple tabular synthetic data rather than LLM training corpora
- you require a mature, long-established project with extensive community validation
Facets
framework · maturity active
data-generation llm-training machine-learning cli gui etl machine-learning large-language-models artificial-intelligence python cli cross-platform synthetic-data model-distillation sft grpo post-training deepeval huggingface low-resource-languages zero-code data-engineering
1 source
- readme: https://github.com/DataArcTech/DataArc-SynData-Toolkit · fetched 2026-08-28 · 4c58609bb82c
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| DataArcTech/DataArc-SynData-Toolkit | main | 57 |
For agents
markdown · JSON · MCP: product_card(name="DataArcTech/DataArc-SynData-Toolkit")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem