# syv-ai/qwen38-27b-rtx3090

Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

Repository: https://github.com/syv-ai/qwen38-27b-rtx3090
Canonical: https://ross.abutalabs.com/products/qwen38-27b-rtx3090
Language: Python
License: Apache-2.0
License Family: permissive
Topics: kv-cache, llm-inference, local-llm, quantization, qwen, qwen3, rtx-3090, speculative-decoding, vllm
Last push: 2026-09-02T19:15:30+00:00

## Health v2 (maintenance only)
Score: 57/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 100, release rhythm 35, longevity 1
- inputs: {"age_days": 18, "days_push": 0, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, young
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1068, forks 137 (observed 2026-09-03T02:15:18.380663+00:00)

## Summary
No AI-extracted summary yet.

## Member repositories
- syv-ai/qwen38-27b-rtx3090 (main) score 57

## Provenance
- Observed fields: from GitHub, fetched 2026-09-03T02:15:18.380663+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Data as of 2026-08-30T08:39:29.467469+00:00.
