01 / SERVERLESS
Serverless inference
Per-token pricing, cold starts under 200 ms, and autoscaling you never have to think about. Zero to 40K requests per second, back to zero.
SYS.00 INFERENCE CLOUD //
Heliox is the inference layer for production AI. One API across 40+ open and frontier models, 38 ms median time-to-first-token, and throughput that scales from side project to nine-figure request volume — without a single GPU on your books.
TRUSTED IN PRODUCTION AT
01 PLATFORM +
Six primitives, one control plane. Compose them from the dashboard, the CLI, or four lines of SDK — then stop thinking about them.
01 / SERVERLESS
Per-token pricing, cold starts under 200 ms, and autoscaling you never have to think about. Zero to 40K requests per second, back to zero.
02 / DEDICATED
Reserved GPU pools with hard latency floors for the traffic you can predict. Burst into serverless for the traffic you can't.
03 / TUNING
LoRA and full-parameter jobs from a single CLI command. Your data never trains anyone else's model, and the weights stay yours.
04 / EVALS
Regression-test model behaviour in CI and block bad outputs before a user ever sees them. Break the build, not the brand.
05 / OBSERVE
Trace every request from token one — latency, cost, and answer quality in a single pane. Export everywhere your team already looks.
06 / BATCH
Queue millions of async jobs at half price, and cache repeated prompt prefixes at the edge so you stop paying for the same tokens twice.
02 DEVELOPERS +
No console safari, no sales call. Export a key, pick a model, stream. The demo on the right is the entire integration.
Point your existing OpenAI-style client at Heliox. Change one line, keep your whole codebase.
Server-sent tokens, tool calls, and structured JSON output — first-class on day one, every model.
Python, TypeScript, and Go. Typed, retried, instrumented. Under 300 KB each, zero transitive drama.
03 MODELS +
Four first-party models tuned on our own silicon, plus 40+ open-weights models served at cost. Swap between them with one string.
| Model | Built for | Context | p50 latency | Input / Output |
|---|---|---|---|---|
| heliox-xenon-2 | Deep reasoning, agents, long-horizon plans | 256K | 310 ms | $4.50 / $13.50 |
| heliox-argon-3 | The production workhorse — chat, RAG, extraction | 128K | 120 ms | $0.80 / $2.40 |
| heliox-neon-mini | Real-time UX, routing, classification | 64K | 45 ms | $0.10 / $0.30 |
| heliox-embed-1 | Search, retrieval, clustering at scale | 8K | 12 ms | $0.02 / — |
▸ EVERY MODEL: streaming · tool calls · JSON mode · batch at 50% off · zero-retention option
04 NETWORK +
Our own scheduler on our own clusters across 26 regions, routed to whichever accelerator answers your request fastest. You see one endpoint; we sweat the rest.
“We moved 100% of production traffic to Heliox over one weekend. Latency halved, the bill dropped 38%, and nobody on my team has thought about GPUs since.”
05 PRICING +
No platform fee until you're earning one back. Usage is metered per token, invoiced monthly, and visible in real time down to the request.
$0 / month
Everything you need to find out if it works.
Most deployed
$249 / month + usage
For traffic that pages someone when it fails.
Custom
For workloads with their own compliance team.
06 IGNITION
A key lands in your inbox in seconds. Five million free tokens a month, forever, on every serverless model.
Your API key is being minted. Check your inbox — then run the demo above against production.
NO CREDIT CARD · 5M FREE TOKENS / MO · CANCEL ANYTIME