Ross ROSS = Recommend OSS · open-source software intelligence for agents

noonghunna/club-3090 resource

Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards. observed · 2026-08-28

github.com/noonghunna/club-3090 · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

79/100

  • Activity 99
  • Release rhythm 93
  • Longevity 9

Flags: young

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 0
  • age_days: 127
  • days_rel: 51
  • days_push: 7
  • n_releases_24m: 32

Full methodology

Adoption not part of the score

2101 stars · 127 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A community-maintained collection of recipes, configs, and scripts for serving large language models locally on consumer NVIDIA GPUs (RTX 3090/4090/5090). It is multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic, shipping curated launch scripts, hardware-aware pickers, and benchmarks for single- and dual-GPU setups.

Use cases

  • serve llms locally on rtx 3090
  • run qwen or gemma on two consumer gpus
  • find working vllm configs for 24gb cards
  • self-host an llm backend in a homelab
  • benchmark llama.cpp vs vllm on consumer gpus
  • set up local llm with open webui and image generation

When to choose

  • You own one or two RTX 3090/4090/5090 cards and want proven serving configs
  • You want a scripted, hardware-aware setup for local LLM inference on Linux/macOS or WSL2
  • You want curated per-model defaults with VRAM budgeting and benchmarks

When to avoid

  • You need production multi-node or datacenter GPU serving
  • You run native Windows without WSL2 (tooling is unsupported there)
  • You need a general-purpose inference library rather than curated configs and scripts

Facets

infra-config · maturity active

llm-inference configuration-management deployment benchmarking developer-tools large-language-models self-hosted gpu-computing developer-tools hardware self-hosted cli consumer-gpus rtx-3090 rtx-4090 rtx-5090 vllm llama-cpp local-llm homelab model-serving recipes qwen gemma linux macos gpu docker

1 source

Member repositories

RepositoryRoleHealth v2
noonghunna/club-3090main79

For agents

markdown · JSON · MCP: product_card(name="noonghunna/club-3090")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem