Run local LLMs on Intel NPU hardware — blazing fast, fully offline, infinitely extensible.
go install github.com/dsainvg/NERVI@latest
Runs LLMs directly on Intel AI Boost NPU via OpenVINO GenAI — no GPU required, no cloud, no latency.
Upload and share extensions — tools and menu integrations that extend NERVI's capabilities.
Models run locally. Your data never leaves the machine. Quantized INT4/FP16 models fit in NPU memory.
Built in Go — fast startup, single binary, interactive TUI. Works anywhere, no runtime dependencies.