SERVERLESS ENDPOINTS · v2.4 · Generally Available

Deploy models instantly. Pay only for Infrastructure

Serverless endpoints spin up on first request, auto-scale to any traffic spike, and vanish when idle — so you never manage infrastructure or overpay for idle GPUs.
cold start
~ 0 ms
uptime SLA
0 %
auto scale
0 → ∞
global availability
0 regions
Client Request INFERENCE GATEWAY Load Balancer · Auth · Rate Limit ~2ms routing WORKER 01 A100 · 40GB 75% vRAM WORKER 02 H100 · 80GB 85% vRAM WORKER 03 A100 · 40GB 40% vRAM RESPONSE STREAM JSON · SSE · WebSocket

CONFIGURATION

Choose your Endpoint Type

Each type is optimized for different latency, throughput,
and cost profiles.

Serverless

Zero infrastructure. Scales from 0 to thousands of concurrent requests automatically. Billed per token — no idle cost.
cold start
~ 300 ms
idle cost
$ 0 .00
idle cost
$ 0 .00

Dedicated

Reserved GPU capacity for predictable, low-latency workloads. Guaranteed availability with no cold starts.
cold start
~ 300 ms
idle cost
$ 0 .00
idle cost
$ 0 .00

Spot

Up to 80% cheaper than dedicated. Best for batch inference, fine-tuning jobs, and fault-tolerant pipelines.
cold start
~ 300 ms
idle cost
$ 0 .00
idle cost
$ 0 .00

DASHBOARD

Your Endpoints

Live status, metrics, and controls for all deployed serverless endpoints.
NAME - MODEL
STATUS
HARDWARE
REQUESTS / MIN
AVG LATENCY
llama-3-70b-instruct
prod-inference-v3
Active
H100 SXM5
1,240
184ms
mixtral-8x22b
staging-moe-test
Cold
A100 80G
0
whisper-large-v3
transcription-api
Active
A10G
327
420ms
sdxl-turbo
image--gen-exp
Paused
A100 40G
0
Model Name
sub-id
Active
H100
0
-

HOW IT WORKS

Serverless Lifecycle

From first request to streamed response — every stage is
optimized for the lowest possible latency.

Request arrives

Client sends inference request to the global endpoint URL with auth token.

Gateway routes

Global load balancer authenticates, rate-limits, and selects the nearest warm worker.

GPU warms up

If cold, container spins up in <340ms. Model weights stream from cache to GPU VRAM.

Inference runs

Model executes on GPU. Output streams back token-by-token via SSE or full JSON response.

Idle -> sleep

After your configured TTL, the container hibernates. You stop paying immediately.

PERFORMANCE

Cold Start Breakdown

We’ve obsessively optimized every millisecond of container initialization.

Total cold start time • p50

340ms
H100 • Llama 3 70B • 8-bit quantized
Oms 60ms 180ms 280ms 340ms
Container start
Model cache fetch
vRAM loading
First token
~340ms
p50 cold start
~680ms
p99 cold start
~18ms
p50 warm latency
10GB/s
model cache speed

INTEGRATION

Call your endpoints in 5 lines

OpenAI-compatible API — drop in your endpoint URL and start inferencing.

SECURITY

Enterprise-Grade Security

Every endpoint is isolated, encrypted, and auditable by default.

API Key Scoping

Fine-grained key permissions: read-only, endpoint-specific, IP-restricted, or rate-limited keys. Rotate without downtime.

End-to-End Encryption

All requests encrypted via TLS 1.3. Inference data never persisted. Container memory wiped after each request batch.

Full Audit Log

Every request logged with timestamp, token usage, latency, and originating IP. Export to S3 or stream to your SIEM.

VPC Isolation

Deploy endpoints inside your private VPC. Zero public internet exposure. Private link support for AWS, GCP, and Azure.

PRICING

Pay for what
you actually use

Billed per millisecond of actual GPU execution time. Zero charges
when your functions are idle.

Launch your first

Deployment Today

GPU-powered workloads, run serverless AI functions, and access production-
ready model endpoints — all on one unified platform.

PRODUCT

Cloud GPUs

Serverless

Public Endpoints

Hub

RESOURCES

Blogs

Case Studies

Referral Program

Articles

Pricing

COMPANY

About

Contact

Careers

Privacy Policy

Terms & Conditions