Ross ROSS = Recommend OSS · open-source software intelligence for agents

wenge-research/YAYI2 resource

YAYI 2 是中科闻歌研发的新一代开源大语言模型,采用了超过 2 万亿 Tokens 的高质量、多语言语料进行预训练。(Repo for YaYi 2 Chinese LLMs) observed · 2026-08-28

github.com/wenge-research/YAYI2 · Python · Apache-2.0 (permissive) · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 0
  • Release rhythm 35
  • Longevity 70

Flags: no_releases archived

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 992
  • days_rel: n/a
  • days_push: 878
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2796 stars · 18 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

YAYI 2 is a family of open-source multilingual large language models (30B Base and Chat variants) developed by Wenge Research, pretrained on over 2 trillion tokens with a focus on Chinese. The repository provides model weights, inference code, fine-tuning scripts (full-parameter and LoRA), and a 500GB pretraining dataset.

Use cases

  • run a chinese open-source llm locally
  • download pretrained 30b language model weights
  • finetune a chinese llm with lora
  • chat with a chinese gpt-style model
  • get multilingual llm pretraining data
  • evaluate llm on c-eval and mmlu benchmarks

When to choose

  • you need a strong open-source Chinese-language LLM with permissive Apache-2.0 code licensing
  • you want to fine-tune a 30B model with provided LoRA or full-parameter training scripts
  • you need multilingual pretraining data for LLM research

When to avoid

  • you need a small model that runs on consumer hardware
  • you require the Chat variant, which was not fully released at last update
  • you need a commercially permissive model license (model weights use a custom YAYI license, data is CC BY-NC 4.0)

Facets

learning-resource · maturity maintenance

llm-inference llm-training chatbot large-language-models artificial-intelligence python chinese-llm pretrained-language-model 30b multilingual huggingface model-weights lora-finetuning natural-language-processing gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
wenge-research/YAYI2main10

For agents

markdown · JSON · MCP: product_card(name="wenge-research/YAYI2")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem