# thu-coai/Safety-Prompts

Chinese safety prompts for evaluating and improving the safety of LLMs. 中文安全prompts，用于评估和提升大模型的安全性。

Repository: https://github.com/thu-coai/Safety-Prompts
Canonical: https://ross.abutalabs.com/products/safety-prompts
Homepage: http://coai.cs.tsinghua.edu.cn/leaderboard/
License: Apache-2.0
License Family: permissive
Topics: attack-defense, chatgpt, instruction, llm, prompt, prompt-engineering, safety, chinese-language
Last push: 2024-02-27T08:42:12+00:00

## Health v2 (maintenance only)
Score: 30/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 88
- inputs: {"age_days": 1233, "days_push": 918, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1214, forks 89 (observed 2026-08-28T04:04:00.781817+00:00)

## What it is
A dataset of 100k Chinese safety prompts with ChatGPT responses covering typical unsafe scenarios and instruction attacks, for evaluating and improving LLM safety. It accompanies the paper 'Safety Assessment of Chinese Large Language Models' from Tsinghua's thu-coai group.

## Use cases
- evaluate the safety of a Chinese LLM
- fine-tune a model to refuse unsafe requests
- test robustness against prompt injection and jailbreak attacks
- align model outputs with human values
- build a safety benchmark for Chinese chatbots
- research instruction attack categories like goal hijacking and prompt leaking

## When to choose
- you need large-scale Chinese-language safety training or evaluation data
- you want to test LLM robustness against Chinese instruction attacks
- you are aligning a Chinese chatbot with safety norms

## When to avoid
- you need English-language safety data
- you want multiple-choice safety benchmarking - use the authors' SafetyBench instead
- you need a ready-made safety detector - use the authors' ShieldLM instead

## Facets
- artifact type: dataset
- maturity: stable
- function: prompt-engineering, machine-learning, llm-training, security
- domain: large-language-models, security, artificial-intelligence
- platform: python, cross-platform
- tags: llm-safety, chinese-language, safety-evaluation, instruction-attacks, alignment, benchmark, fine-tuning-data, natural-language-processing

## Member repositories
- thu-coai/Safety-Prompts (main) score 30

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:00.781817+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:17:31.295691+00:00, confidence not recorded.
  - readme: https://github.com/thu-coai/Safety-Prompts (fetched 2026-08-28T04:04:00.781817+00:00, sha c5a662737869)
- Data as of 2026-08-30T08:39:29.467469+00:00.
