OpenAI-compatible inference, on tap

Ship AI features
without the infra.

Neko-AI hosts frontier models behind a single OpenAI-compatible endpoint. Pooled provider accounts, automatic failover, credit billing and per-token pricing — so you can focus on your product.

No credit card required · Pay-as-you-go credits

Live model pricing

Every enabled model, straight from our routing layer. Prices are in credits per 1M tokens.

Full pricing
ModelContextInputOutputCache readCache write
DeepSeek V4.1 Flash1M0.00300000.01270000.00008000.0000000

Everything you need to run inference

A production-grade gateway between your app and the frontier — reliability, billing and observability included.

OpenAI-compatible API

Drop-in replacement for the OpenAI Chat Completions API. Point your existing SDK at our base URL and ship in minutes.

Load-balanced account pools

Every model is backed by a pool of upstream provider accounts with round-robin, weighted or least-latency routing and automatic failover.

Credit billing

Prepaid credits are reserved per request, settled against real token usage, and refunded automatically when a request fails.

Transparent per-token pricing

Input, output and cache tokens are priced individually. No subscriptions, no minimums — pay exactly for what you use.

Scoped API keys

Issue multiple keys, see per-key usage and revoke instantly. Plaintext keys are shown once and stored hashed.

Built-in playground

Test any model straight from the dashboard with a live, streamed chat — no code required to get a feel for output.

Two lines to migrate

Same API you already know

Change your base URL and API key, and keep using the OpenAI SDKs. Chat completions, streaming, and model listing all work exactly as you expect.

  • Point the OpenAI SDK at our base URL
  • Authenticate with a Bearer API key
  • Stream responses with stream: true
POST /v1/chat/completions
curl https://your-domain.com/v1/chat/completions \
  -H "Authorization: Bearer sk-neko-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gpt-4o-mini",
    "messages": [
      { "role": "user", "content": "Say hello in one line." }
    ],
    "stream": true
  }'

99.9%

Uptime

< 400ms

Avg. first token

3×

Failover attempts

Pooled

Provider accounts

Start building in under a minute

Create an account, generate an API key and send your first request. Top up credits only when you are ready to scale.