API reference

Build with the Neko-AI API

A drop-in, OpenAI-compatible interface for chat completions and model listing. Authenticate with a Bearer API key and send requests to our base URL — that is all there is to it.

Overview

The API follows the OpenAI Chat Completions shape, so existing SDKs and clients work with minimal changes. Requests are distributed across health-checked backends with automatic failover, and every call is metered against your credit balance.

One endpoint

Route any model through a single base URL.

Key auth

Bearer API keys, hashed at rest.

Streaming

Server-sent events for token streams.

Authentication

Create an API key in your dashboard under API Keys. The plaintext key is shown once at creation time; only a hash is stored. Send it on every request in the Authorization header:

Authorization header
Authorization: Bearer sk-neko-xxxxxxxx

Keep your key secret. Requests with a missing, malformed or revoked key return a 401 authentication_error.

Base URL

All endpoints are rooted at a single base URL. Replace your-domain.com with your deployment host.

Base URL
https://your-domain.com/v1
MethodPathDescription
POST/v1/chat/completionsCreate a chat completion.
GET/v1/modelsList available models and pricing.

Chat completions

POST/v1/chat/completions

Generates a model response from a list of chat messages. Supports the standard OpenAI parameters.

ParameterTypeRequiredDescription
modelstringrequiredThe model id to run. Accepts the model id or display name.
messagesarrayrequiredConversation history. Each item has a `role` (system, user, assistant) and `content`.
streambooleanoptionalWhen true, tokens are returned as a server-sent event stream.
temperaturenumberoptionalSampling temperature between 0 and 2.
max_tokensnumberoptionalMaximum number of tokens to generate.
top_pnumberoptionalNucleus sampling probability mass.
curl — non-streaming
curl https://your-domain.com/v1/chat/completions \
  -H "Authorization: Bearer sk-neko-xxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [
      { "role": "system", "content": "You are a helpful assistant." },
      { "role": "user", "content": "Explain load balancing in one sentence." }
    ]
  }'

A successful response mirrors the OpenAI shape, including a usage object with prompt, completion, cache-read and cache-write token counts.

Streaming

Set stream: true to receive tokens incrementally as server-sent events. The stream ends with a data: [DONE] sentinel and a final usage chunk.

curl — streaming
curl -N https://your-domain.com/v1/chat/completions \
  -H "Authorization: Bearer sk-neko-xxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "stream": true,
    "messages": [{ "role": "user", "content": "Count to five." }]
  }'

With an SDK you can consume the async iterator directly:

JavaScript / TypeScript
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.NEKO_API_KEY, // "sk-neko-..."
  baseURL: "https://your-domain.com/v1",
});

const stream = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Say hello." }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Responses API

POST/v1/responses

An OpenAI-compatible Responses endpoint. Accepts input (a string or an array of input items) plus optional instructions, and returns a response object with an output array and output_text. Function tools are supported. Requests are translated to chat completions and billed identically.

curl — non-streaming
curl https://your-domain.com/v1/responses \
  -H "Authorization: Bearer sk-neko-xxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "instructions": "You are a helpful assistant.",
    "input": "Explain load balancing in one sentence."
  }'

Set stream: true to receive the Responses event sequence (response.created, response.output_text.delta, … response.completed).

curl — streaming
curl -N https://your-domain.com/v1/responses \
  -H "Authorization: Bearer sk-neko-xxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "input": "Count to five.",
    "stream": true
  }'

List models

GET/v1/models

Returns every enabled model along with its context window and per-token pricing.

curl
curl https://your-domain.com/v1/models \
  -H "Authorization: Bearer sk-neko-xxxxxxxx"
200 OK
{
  "object": "list",
  "data": [
    {
      "id": "gpt-4o-mini",
      "object": "model",
      "created": 1730000000,
      "owned_by": "neko-ai",
      "display_name": "GPT-4o mini",
      "context_window": 128000,
      "pricing": {
        "input": "0.15",
        "output": "0.6",
        "cache_read": "0.075",
        "cache_write": "0",
        "unit": "credits per 1M tokens"
      }
    }
  ]
}

Error codes

Errors are returned as JSON with a top-level error object.

Error shape
{
  "error": {
    "message": "Model 'foo' does not exist or is disabled.",
    "type": "invalid_request_error",
    "code": "model_not_found"
  }
}
StatusTypeMeaning
400invalid_request_errorMissing or malformed request body (e.g. no `model`).
401authentication_errorMissing, malformed or revoked API key.
402insufficient_quotaYour credit balance is too low to reserve this request.
404model_not_foundThe requested model does not exist or is disabled.
429rate_limit_errorToo many requests — slow down and retry with backoff.
503api_errorNo capacity available right now. Retry shortly.

Ready to make your first call?

Create an account, generate an API key and send a request — your first responses are only minutes away.