Welcome to NERVI

NERVI is a premium local command-line interface (CLI) LLM client designed to run high-performance AI inference on the Intel NPU (Neural Processing Unit).

NERVI acts as a bridge between the user and a background Flask-based REST service (npuserver), which leverages Intel's OpenVINO GenAI library to optimize and execute models locally.


🏗️ Architecture Overview

The system is split into two primary layers:

1. Frontend CLI Client (Go)

The Go client compiles into a single executable nervi.exe. It manages:

  • Terminal UI Rendering: The interactive chat interface, status lines, and spinner.
  • Subprocess Spawning: Auto-starting and stopping the background Flask server.
  • REST Requests: Communicating configuration changes, model downloads, and inference prompt queues.
  • Extension Hooks: Managing standard I/O IPC brokers for third-party extensions.

2. Backend Model Server (npuserver)

The backend is a Flask python application installed inside a managed virtual environment. It handles:

  • Registry Discoveries: Cataloging local Hugging Face downloads, compiled slots, and remote registries.
  • Model Download & Compilation: Downloading weight files from Hugging Face Hub and compiling them into NPU execution blobs.
  • LLM Pipelines: Running streaming and blocking chat completions using OpenVINO.

📂 System Directory Structure

All NERVI assets, virtual environments, caches, and logs are isolated inside standard configuration directories:

  • Windows: %AppData%\Roaming\nervi\ (maps to C:\Users\<username>\AppData\Roaming\nervi\)
  • Unix: ~/.config/nervi/

Directory Contents

  • venv/: A self-contained Python virtual environment.
  • extensions/: Folder containing installed custom CLI extensions, each in its own sub-folder with a manifest.json.
  • logs/: Logs generated by the background server and extension subprocesses.
  • server.pid: Tracks the Process ID of the running background server.
  • .nervi_ready: A sentinel file indicating Python virtual environment setup is complete.
  • config.json: Client configurations, mapping active port settings and user model nicknames.