Model Pulling & NPU Compilation

Model compilation is the process of translating standard Hugging Face model weights (e.g. OpenVINO FP16/INT4 structures) into hardware-optimized binary blobs that execute directly on the Intel NPU.


📋 Registry States

When you request a list of available models using nervi pull --list, the server queries the local file cache and remote lists to classify each model into one of four states:

  1. available: Discovered on remote registries but not yet downloaded locally.
  2. downloaded: Model weights are present in the Hugging Face cache directory, but no NPU compilation blob has been created yet.
  3. compiled: Pre-compiled NPU execution blobs exist locally. Ready for instant loading.
  4. active / loaded: Currently residing in the NPU's active execution memory and serving prompts.

⚡ NPU Compilation Mechanics

When compilation is triggered (by nervi pull <model> or when running nervi load on a "downloaded" model):

  1. Slot Allocation: The server allocates a folder in ~/.cache/npuserver/compiled/.
  2. Configuration: Configures a manifest JSON mapping max prompt length and speed optimization profiles.
  3. OpenVINO Pipeline: Instantiates ov_genai.LLMPipeline targeting "NPU". This compiles the weights, generating .blob files.
  4. Sentinel Touch: Writes a compiled.ok file to signify successful compilation.

📶 Handling Network & Download Interruptions

Resuming Incomplete Downloads

If your internet connection drops while downloading a model, running nervi pull <model> again will automatically attempt to resume downloading from the exact file chunk where it was interrupted.

Offline Compilation

If model weights are fully downloaded, you can trigger compilation offline. The server calls snapshot_download(local_files_only=True) and compiles the local weights without requiring an internet connection.


⏳ Client Timeout Exemption

Downloading and compiling can take 5 to 30 minutes depending on connection bandwidth and NPU speed. All client REST requests targeting the v1/models/download and v1/models/load endpoints ignore default client timeout limits (Timeout: 0).