NERVI

Run local LLMs on Intel NPU hardware — blazing fast, fully offline, infinitely extensible.

Quick install go install github.com/dsainvg/NERVI@latest
⚡

NPU-Native Inference

Runs LLMs directly on Intel AI Boost NPU via OpenVINO GenAI — no GPU required, no cloud, no latency.

🧩

Extension Marketplace

Upload and share extensions — tools and menu integrations that extend NERVI's capabilities.

🔒

Fully Offline

Models run locally. Your data never leaves the machine. Quantized INT4/FP16 models fit in NPU memory.

🛠️

CLI First

Built in Go — fast startup, single binary, interactive TUI. Works anywhere, no runtime dependencies.