HongyuanLuke/frequencylaw
Official repository for textual frequency law observed · 2026-08-28
Health v2 · maintenance only
49/100
- Activity 77
- Release rhythm 35
- Longevity 11
Flags: no_releases young no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 160
- days_rel: n/a
- days_push: 141
- n_releases_24m: 0
Adoption not part of the score
1435 stars · 38 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Official code repository for the paper 'Textual Frequency Law on Large Language Models', implementing TFL, TFD, and CTFT methods for studying how textual frequency affects LLM performance. It provides end-to-end scripts for frequency calculation, paired dataset construction, LoRA fine-tuning, and evaluation on GSM8K math reasoning and FLORES-200 translation tasks.
Use cases
- reproduce textual frequency law experiments on LLMs
- build high/low frequency paired datasets for math reasoning
- fine-tune LLMs with curriculum frequency-based training
- evaluate LLM performance on GSM8K and FLORES-200
- calculate textual frequency statistics for training data
- study how token frequency impacts LLM reasoning and translation
When to choose
- you want to reproduce or extend the paper's frequency-based LLM training methods
- you need frequency-sorted or paired datasets for reasoning/translation experiments
- you're researching how textual frequency affects model optimization
When to avoid
- you need a production-ready LLM training framework
- you want a general-purpose fine-tuning tool unrelated to frequency research
- you need a maintained library with a license and long-term support
Facets
library · maturity active
machine-learning llm-training nlp data-generation benchmarking large-language-models machine-learning python llm-research fine-tuning lora mathematical-reasoning machine-translation academic-paper dataset-construction gsm8k flores-200 natural-language-processing research
1 source
- readme: https://github.com/HongyuanLuke/frequencylaw · fetched 2026-08-28 · 630ad9afa898
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| HongyuanLuke/frequencylaw | main | 49 |
For agents
markdown · JSON · MCP: product_card(name="HongyuanLuke/frequencylaw")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem