SYS.00 INFERENCE CLOUD //

Run frontier models at wire speed.

Heliox is the inference layer for production AI. One API across 40+ open and frontier models, 38 ms median time-to-first-token, and throughput that scales from side project to nine-figure request volume — without a single GPU on your books.

38 msp50 first token
99.99%uptime SLA
4.2Ttokens / week
26regions

TRUSTED IN PRODUCTION AT

VELLUM OKTANT parsec·bio Dune&Flint northcell_ ARRAKIA

01 PLATFORM +

Everything between your prompt and the GPU.

Six primitives, one control plane. Compose them from the dashboard, the CLI, or four lines of SDK — then stop thinking about them.

01 / SERVERLESS

Serverless inference

Per-token pricing, cold starts under 200 ms, and autoscaling you never have to think about. Zero to 40K requests per second, back to zero.

02 / DEDICATED

Dedicated capacity

Reserved GPU pools with hard latency floors for the traffic you can predict. Burst into serverless for the traffic you can't.

03 / TUNING

Fine-tuning

LoRA and full-parameter jobs from a single CLI command. Your data never trains anyone else's model, and the weights stay yours.

04 / EVALS

Evals & guardrails

Regression-test model behaviour in CI and block bad outputs before a user ever sees them. Break the build, not the brand.

05 / OBSERVE

Observability

Trace every request from token one — latency, cost, and answer quality in a single pane. Export everywhere your team already looks.

06 / BATCH

Batch & prefix cache

Queue millions of async jobs at half price, and cache repeated prompt prefixes at the edge so you stop paying for the same tokens twice.

02 DEVELOPERS +

Your first token in under a minute.

No console safari, no sales call. Export a key, pick a model, stream. The demo on the right is the entire integration.

  • A
    Drop-in compatible

    Point your existing OpenAI-style client at Heliox. Change one line, keep your whole codebase.

  • B
    Streaming-native

    Server-sent tokens, tool calls, and structured JSON output — first-class on day one, every model.

  • C
    SDKs that respect you

    Python, TypeScript, and Go. Typed, retried, instrumented. Under 300 KB each, zero transitive drama.

03 MODELS +

A lineup for every latency budget.

Four first-party models tuned on our own silicon, plus 40+ open-weights models served at cost. Swap between them with one string.

FIRST-PARTY LINEUP — PRICES PER 1M TOKENS
Model Built for Context p50 latency Input / Output
heliox-xenon-2 Deep reasoning, agents, long-horizon plans 256K 310 ms $4.50 / $13.50
heliox-argon-3 The production workhorse — chat, RAG, extraction 128K 120 ms $0.80 / $2.40
heliox-neon-mini Real-time UX, routing, classification 64K 45 ms $0.10 / $0.30
heliox-embed-1 Search, retrieval, clustering at scale 8K 12 ms $0.02 /

▸ EVERY MODEL: streaming · tool calls · JSON mode · batch at 50% off · zero-retention option

04 NETWORK +

Built like infrastructure, because it is.

Our own scheduler on our own clusters across 26 regions, routed to whichever accelerator answers your request fastest. You see one endpoint; we sweat the rest.

99.99%measured uptime, trailing 12 mo
38 msmedian time-to-first-token
4.2Ttokens served weekly
0prompts retained in zero-retention mode
SOC 2 TYPE IIISO 27001ZERO-RETENTION MODESSO / SCIMPRIVATE NETWORKING
“We moved 100% of production traffic to Heliox over one weekend. Latency halved, the bill dropped 38%, and nobody on my team has thought about GPUs since.”
PRIYA RAMAN — VP ENGINEERING, ARRAKIA

05 PRICING +

Pay for tokens, not promises.

No platform fee until you're earning one back. Usage is metered per token, invoiced monthly, and visible in real time down to the request.

Build

$0 / month

Everything you need to find out if it works.

  • 5M free tokens every month
  • All serverless models
  • 2 fine-tune jobs included
  • Community support
Start free

Most deployed

Scale

$249 / month + usage

For traffic that pages someone when it fails.

  • Priority routing & 99.99% SLA
  • Batch API at 50% off
  • Unlimited fine-tunes
  • Evals & guardrails suite
  • Named support engineer
Get API key

Enterprise

Custom

For workloads with their own compliance team.

  • Dedicated GPU capacity
  • Private networking & VPC peering
  • Zero-retention, custom DPAs
  • Custom SLAs, 24/7 response
Talk to us

06 IGNITION

Ship your first request tonight.

A key lands in your inbox in seconds. Five million free tokens a month, forever, on every serverless model.

Request received.

Your API key is being minted. Check your inbox — then run the demo above against production.

NO CREDIT CARD · 5M FREE TOKENS / MO · CANCEL ANYTIME

Built by ECTD · see more concepts