thu-coai/Safety-Prompts resource
Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts,用于评估和提升大模型的安全性。 observed · 2026-08-28
Health v2 · maintenance only
30/100
- Activity 0
- Release rhythm 35
- Longevity 88
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1233
- days_rel: n/a
- days_push: 918
- n_releases_24m: 0
Adoption not part of the score
1214 stars · 89 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A dataset of 100k Chinese safety prompts with ChatGPT responses covering typical unsafe scenarios and instruction attacks, for evaluating and improving LLM safety. It accompanies the paper 'Safety Assessment of Chinese Large Language Models' from Tsinghua's thu-coai group.
Use cases
- evaluate the safety of a Chinese LLM
- fine-tune a model to refuse unsafe requests
- test robustness against prompt injection and jailbreak attacks
- align model outputs with human values
- build a safety benchmark for Chinese chatbots
- research instruction attack categories like goal hijacking and prompt leaking
When to choose
- you need large-scale Chinese-language safety training or evaluation data
- you want to test LLM robustness against Chinese instruction attacks
- you are aligning a Chinese chatbot with safety norms
When to avoid
- you need English-language safety data
- you want multiple-choice safety benchmarking - use the authors' SafetyBench instead
- you need a ready-made safety detector - use the authors' ShieldLM instead
Facets
dataset · maturity stable
prompt-engineering machine-learning llm-training security large-language-models security artificial-intelligence python cross-platform llm-safety chinese-language safety-evaluation instruction-attacks alignment benchmark fine-tuning-data natural-language-processing
1 source
- readme: https://github.com/thu-coai/Safety-Prompts · fetched 2026-08-28 · c5a662737869
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| thu-coai/Safety-Prompts | main | 30 |
For agents
markdown · JSON · MCP: product_card(name="thu-coai/Safety-Prompts")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem