AZ Labs
smart_toy

Free preview · no card required

Model Gateway

One API. Every model. Free preview.

Get a keykey8 models6 developers

Catalogue

All models

Every model the gateway exposes, with live health from the request log.

8 models available

NVIDIA

No data

nemotron-3-ultra

Nemotron 3 Ultra

550B mixture-of-experts flagship with a 1M token context window. Best general-purpose choice for long documents and complex reasoning.

Latency
24h
Context
1M
ToolsStreaming

DeepSeek

No data

deepseek-v4-flash

DeepSeek V4 Flash

Fast coding and reasoning model with strong SWE-bench performance. Good default for code generation and refactoring.

Latency
24h
Context
262K
ToolsStreamingReasoning

Cohere

No data

north-mini-code

North Mini Code

30B mixture-of-experts model tuned for code. Low latency with a 256K context window.

Latency
24h
Context
256K
ToolsStreaming

InclusionAI

No data

ling-3.0-flash

Ling 3.0 Flash

Lightweight high-throughput model for classification, extraction, and short-form generation.

Latency
24h
Context
262K
Streaming

NVIDIA

No data

nemotron-3-nano-30b

Nemotron 3 Nano 30B

Compact 30B mixture-of-experts model. Efficient choice for summarisation, routing, and agent scaffolding.

Latency
24h
Context
256K
Streaming

Thinking Machines

No data

inkling

Inkling

Reasoning model from Thinking Machines for long-horizon tasks. 256K context.

Latency
24h
Context
256K
StreamingReasoning

Meta

No data

muse-glimmer-30b

Muse Glimmer 30B

Meta image-and-text model. NIM free tier.

Latency
24h
Context
128K
Streaming

NVIDIA

No data

nemotron-3.5-lightning-30b

Nemotron 3.5 Lightning 30B

Fast NVIDIA MoE for high-throughput reasoning. 256K context.

Latency
24h
Context
256K
Streaming

Catalogue updated 18 Aug, 08:54. Health, latency, and request volume come straight from the gateway request log. One OpenAI-compatible endpoint, every model — grab a free key.