intel/ipex-llm
Accelerate local LLM inference and finetuning (LLaMA, Mistral, ChatGLM, Qwen, DeepSeek, Mixtral, Gemma, Phi, MiniCPM, Qwen-VL, MiniCPM-V, etc.) on Intel XPU (e.g., local PC with iGPU and NPU, discrete GPU such as Arc, Flex and Max); seamlessly integrate with llama.cpp, Ollama, HuggingFace, LangChain, LlamaIndex, vLLM, DeepSpeed, Axolotl, etc. observed · 2026-08-28
Health v2 · maintenance only
10/100
- Activity 64
- Release rhythm 40
- Longevity 100
Flags: archived
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 1
- age_days: 3656
- days_rel: 511
- days_push: 217
- n_releases_24m: 2
Adoption not part of the score
8859 stars · 1430 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
IPEX-LLM is a PyTorch LLM acceleration library for Intel hardware (iGPU, NPU, Arc/Flex/Max GPUs, and CPU), offering low-bit quantization (FP8/FP6/FP4/INT4) and integrations with llama.cpp, Ollama, vLLM, HuggingFace, LangChain, and more. The project has been archived by Intel and no longer receives maintenance, bug fixes, or security updates.
Use cases
- run local LLM inference on Intel Arc GPU
- run DeepSeek or Qwen models on Intel iGPU or NPU
- finetune LLaMA or Mistral on Intel GPUs
- accelerate Ollama or llama.cpp on Intel hardware
- quantize LLMs to INT4 or FP4 for local inference
- run 671B models on one or two Arc A770 GPUs
When to choose
- you are locked into Intel XPU hardware and need an existing, tested LLM acceleration path
- you need historical reference for Intel GPU LLM optimization techniques
- you maintain a legacy deployment already built on ipex-llm
When to avoid
- starting a new project, since the repo is archived with known security issues
- you need ongoing support, bug fixes, or new model support
- you are on NVIDIA/AMD hardware where better-maintained alternatives exist
- security-sensitive or multi-tenant environments
Facets
library · maturity abandoned
llm-inference llm-training gpu-computing machine-learning large-language-models machine-learning deep-learning gpu-computing python windows intel-xpu low-bit-quantization ollama llama-cpp vllm finetuning archived linux gpu
1 source
- readme: https://github.com/intel/ipex-llm · fetched 2026-08-28 · 9873c75e977d
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| intel/ipex-llm | main | 10 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem