Complete API Reference

This page contains the complete documentation for all available endpoints of the npuserver.


Base Status Endpoint

GET /

Retrieve the basic status of the API server.

Response:

{
  "status": "running",
  "active_model": "none",
  "description": "npuserver local LLM server active."
}

Health Diagnostics

GET /health or GET /currentmodel

Retrieve runtime telemetry detailing the operational state of the hardware and the currently loaded model.

Response (Loaded State):

{
  "status": "ok",
  "state": "loaded",
  "device": "NPU",
  "active_model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4"
}

Model Registry

GET /v1/models

Retrieve a fully populated JSON list mapping all discovered models and their current availability state.

Response:

{
  "object": "list",
  "data": [
    {
      "id": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
      "object": "model",
      "status": "active"
    }
  ]
}

Model Lifecycle Management

POST /v1/models/load

Compile and activate a specific model onto the NPU.

Request:

{
  "model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
  "max_prompt_len": 2048
}

POST /unload

Unload the currently active model, freeing up NPU memory.

POST /v1/models/delete

Delete a model's compiled cache files and/or downloaded weights.

Request:

{
  "model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
  "compiled_only": true
}

Model Download and Compilation

POST /v1/models/download

Download a model from Hugging Face and compile it for the NPU. This call blocks until compilation is complete.

Request:

{
  "model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
  "max_prompt_len": 2048,
  "allow_download": true
}

Chat Completions (Blocking)

POST /v1/chat/completions

Execute a chat completion request.

Request:

{
  "model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
  "messages": [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain quantum computing in one sentence."}
  ],
  "max_tokens": 150,
  "temperature": 0.7,
  "stream": false
}

Response includes:

  • Standard OpenAI-compatible response structure
  • reasoning_content field for reasoning models
  • metrics object with NPU performance data (TTFT, throughput, etc.)

Chat Completions (Streaming)

POST /v1/chat/completions (with "stream": true)

Execute a streaming chat completion using Server-Sent Events (SSE).

The response streams:

  1. Early chunks containing reasoning content (from <think> tokens)
  2. Transitional chunks with actual content
  3. Final summary chunk with full NPU engine performance metrics
  4. Stream termination: data: [DONE]