A drop-in, OpenAI-compatible interface for chat completions and model listing. Authenticate with a Bearer API key and send requests to our base URL — that is all there is to it.
The API follows the OpenAI Chat Completions shape, so existing SDKs and clients work with minimal changes. Requests are distributed across health-checked backends with automatic failover, and every call is metered against your credit balance.
One endpoint
Route any model through a single base URL.
Key auth
Bearer API keys, hashed at rest.
Streaming
Server-sent events for token streams.
Create an API key in your dashboard under API Keys. The plaintext key is shown once at creation time; only a hash is stored. Send it on every request in the Authorization header:
Authorization: Bearer sk-neko-xxxxxxxxKeep your key secret. Requests with a missing, malformed or revoked key return a 401 authentication_error.
All endpoints are rooted at a single base URL. Replace your-domain.com with your deployment host.
https://your-domain.com/v1| Method | Path | Description |
|---|---|---|
| POST | /v1/chat/completions | Create a chat completion. |
| GET | /v1/models | List available models and pricing. |
/v1/chat/completionsGenerates a model response from a list of chat messages. Supports the standard OpenAI parameters.
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | required | The model id to run. Accepts the model id or display name. |
| messages | array | required | Conversation history. Each item has a `role` (system, user, assistant) and `content`. |
| stream | boolean | optional | When true, tokens are returned as a server-sent event stream. |
| temperature | number | optional | Sampling temperature between 0 and 2. |
| max_tokens | number | optional | Maximum number of tokens to generate. |
| top_p | number | optional | Nucleus sampling probability mass. |
curl https://your-domain.com/v1/chat/completions \
-H "Authorization: Bearer sk-neko-xxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "Explain load balancing in one sentence." }
]
}'A successful response mirrors the OpenAI shape, including a usage object with prompt, completion, cache-read and cache-write token counts.
Set stream: true to receive tokens incrementally as server-sent events. The stream ends with a data: [DONE] sentinel and a final usage chunk.
curl -N https://your-domain.com/v1/chat/completions \
-H "Authorization: Bearer sk-neko-xxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"stream": true,
"messages": [{ "role": "user", "content": "Count to five." }]
}'With an SDK you can consume the async iterator directly:
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.NEKO_API_KEY, // "sk-neko-..."
baseURL: "https://your-domain.com/v1",
});
const stream = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: "Say hello." }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}/v1/responsesAn OpenAI-compatible Responses endpoint. Accepts input (a string or an array of input items) plus optional instructions, and returns a response object with an output array and output_text. Function tools are supported. Requests are translated to chat completions and billed identically.
curl https://your-domain.com/v1/responses \
-H "Authorization: Bearer sk-neko-xxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"instructions": "You are a helpful assistant.",
"input": "Explain load balancing in one sentence."
}'Set stream: true to receive the Responses event sequence (response.created, response.output_text.delta, … response.completed).
curl -N https://your-domain.com/v1/responses \
-H "Authorization: Bearer sk-neko-xxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"input": "Count to five.",
"stream": true
}'/v1/modelsReturns every enabled model along with its context window and per-token pricing.
curl https://your-domain.com/v1/models \
-H "Authorization: Bearer sk-neko-xxxxxxxx"{
"object": "list",
"data": [
{
"id": "gpt-4o-mini",
"object": "model",
"created": 1730000000,
"owned_by": "neko-ai",
"display_name": "GPT-4o mini",
"context_window": 128000,
"pricing": {
"input": "0.15",
"output": "0.6",
"cache_read": "0.075",
"cache_write": "0",
"unit": "credits per 1M tokens"
}
}
]
}Errors are returned as JSON with a top-level error object.
{
"error": {
"message": "Model 'foo' does not exist or is disabled.",
"type": "invalid_request_error",
"code": "model_not_found"
}
}| Status | Type | Meaning |
|---|---|---|
| 400 | invalid_request_error | Missing or malformed request body (e.g. no `model`). |
| 401 | authentication_error | Missing, malformed or revoked API key. |
| 402 | insufficient_quota | Your credit balance is too low to reserve this request. |
| 404 | model_not_found | The requested model does not exist or is disabled. |
| 429 | rate_limit_error | Too many requests — slow down and retry with backoff. |
| 503 | api_error | No capacity available right now. Retry shortly. |