Complete API Reference
This page contains the complete documentation for all available endpoints of the npuserver.
Base Status Endpoint
GET /
Retrieve the basic status of the API server.
Response:
{
"status": "running",
"active_model": "none",
"description": "npuserver local LLM server active."
}
Health Diagnostics
GET /health or GET /currentmodel
Retrieve runtime telemetry detailing the operational state of the hardware and the currently loaded model.
Response (Loaded State):
{
"status": "ok",
"state": "loaded",
"device": "NPU",
"active_model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4"
}
Model Registry
GET /v1/models
Retrieve a fully populated JSON list mapping all discovered models and their current availability state.
Response:
{
"object": "list",
"data": [
{
"id": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
"object": "model",
"status": "active"
}
]
}
Model Lifecycle Management
POST /v1/models/load
Compile and activate a specific model onto the NPU.
Request:
{
"model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
"max_prompt_len": 2048
}
POST /unload
Unload the currently active model, freeing up NPU memory.
POST /v1/models/delete
Delete a model's compiled cache files and/or downloaded weights.
Request:
{
"model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
"compiled_only": true
}
Model Download and Compilation
POST /v1/models/download
Download a model from Hugging Face and compile it for the NPU. This call blocks until compilation is complete.
Request:
{
"model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
"max_prompt_len": 2048,
"allow_download": true
}
Chat Completions (Blocking)
POST /v1/chat/completions
Execute a chat completion request.
Request:
{
"model": "Qwen/Qwen2.5-7B-Instruct-OpenVINO-INT4",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Explain quantum computing in one sentence."}
],
"max_tokens": 150,
"temperature": 0.7,
"stream": false
}
Response includes:
- Standard OpenAI-compatible response structure
reasoning_contentfield for reasoning modelsmetricsobject with NPU performance data (TTFT, throughput, etc.)
Chat Completions (Streaming)
POST /v1/chat/completions (with "stream": true)
Execute a streaming chat completion using Server-Sent Events (SSE).
The response streams:
- Early chunks containing reasoning content (from
<think>tokens) - Transitional chunks with actual content
- Final summary chunk with full NPU engine performance metrics
- Stream termination:
data: [DONE]