Neko-AI hosts frontier models behind a single OpenAI-compatible endpoint. Pooled provider accounts, automatic failover, credit billing and per-token pricing — so you can focus on your product.
No credit card required · Pay-as-you-go credits
Every enabled model, straight from our routing layer. Prices are in credits per 1M tokens.
| Model | Context | Input | Output | Cache read | Cache write |
|---|---|---|---|---|---|
| DeepSeek V4.1 Flash | 1M | 0.0030000 | 0.0127000 | 0.0000800 | 0.0000000 |
A production-grade gateway between your app and the frontier — reliability, billing and observability included.
Drop-in replacement for the OpenAI Chat Completions API. Point your existing SDK at our base URL and ship in minutes.
Every model is backed by a pool of upstream provider accounts with round-robin, weighted or least-latency routing and automatic failover.
Prepaid credits are reserved per request, settled against real token usage, and refunded automatically when a request fails.
Input, output and cache tokens are priced individually. No subscriptions, no minimums — pay exactly for what you use.
Issue multiple keys, see per-key usage and revoke instantly. Plaintext keys are shown once and stored hashed.
Test any model straight from the dashboard with a live, streamed chat — no code required to get a feel for output.
Change your base URL and API key, and keep using the OpenAI SDKs. Chat completions, streaming, and model listing all work exactly as you expect.
curl https://your-domain.com/v1/chat/completions \
-H "Authorization: Bearer sk-neko-..." \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini",
"messages": [
{ "role": "user", "content": "Say hello in one line." }
],
"stream": true
}'99.9%
Uptime
< 400ms
Avg. first token
3×
Failover attempts
Pooled
Provider accounts