NVIDIA
No datanemotron-3-ultra
Nemotron 3 Ultra
550B mixture-of-experts flagship with a 1M token context window. Best general-purpose choice for long documents and complex reasoning.
- Latency
- —
- 24h
- —
- Context
- 1M
Free preview · no card required
One API. Every model. Free preview.
Catalogue
Every model the gateway exposes, with live health from the request log.
8 models available
NVIDIA
No datanemotron-3-ultra
550B mixture-of-experts flagship with a 1M token context window. Best general-purpose choice for long documents and complex reasoning.
DeepSeek
No datadeepseek-v4-flash
Fast coding and reasoning model with strong SWE-bench performance. Good default for code generation and refactoring.
Cohere
No datanorth-mini-code
30B mixture-of-experts model tuned for code. Low latency with a 256K context window.
InclusionAI
No dataling-3.0-flash
Lightweight high-throughput model for classification, extraction, and short-form generation.
NVIDIA
No datanemotron-3-nano-30b
Compact 30B mixture-of-experts model. Efficient choice for summarisation, routing, and agent scaffolding.
Thinking Machines
No datainkling
Reasoning model from Thinking Machines for long-horizon tasks. 256K context.
Meta
No datamuse-glimmer-30b
Meta image-and-text model. NIM free tier.
NVIDIA
No datanemotron-3.5-lightning-30b
Fast NVIDIA MoE for high-throughput reasoning. 256K context.
Catalogue updated 18 Aug, 10:46. Health, latency, and request volume come straight from the gateway request log. One OpenAI-compatible endpoint, every model — grab a free key.