Home
Serverless
About
Endpoints
Video Generation
Job Board
Blogs
Blog
Contact
Home
Serverless
About
Endpoints
Video Generation
Job Board
Blogs
Blog
Contact
Sign In
Get Started
SERVERLESS ENDPOINTS · v2.4 · Generally Available
Deploy models instantly. Pay only for
Infrastructure
Serverless endpoints spin up on first request, auto-scale to any traffic spike, and vanish when idle — so you never manage infrastructure or overpay for idle GPUs.
Start Deploy
How it Works?
cold start
~
0
ms
uptime SLA
0
%
auto scale
0
→ ∞
global availability
0
regions
Client Request
INFERENCE GATEWAY
Load Balancer · Auth · Rate Limit
~2ms
routing
WORKER 01
A100 · 40GB
75% vRAM
WORKER 02
H100 · 80GB
85% vRAM
WORKER 03
A100 · 40GB
40% vRAM
RESPONSE STREAM
JSON · SSE · WebSocket
CONFIGURATION
Choose your
Endpoint Type
Each type is optimized for different latency, throughput,
and cost profiles.
Recommended
Serverless
Zero infrastructure. Scales from 0 to thousands of concurrent requests automatically. Billed per token — no idle cost.
cold start
~
300
ms
idle cost
$
0
.00
idle cost
$
0
.00
Dedicated
Reserved GPU capacity for predictable, low-latency workloads. Guaranteed availability with no cold starts.
cold start
~
300
ms
idle cost
$
0
.00
idle cost
$
0
.00
Spot
Up to 80% cheaper than dedicated. Best for batch inference, fine-tuning jobs, and fault-tolerant pipelines.
cold start
~
300
ms
idle cost
$
0
.00
idle cost
$
0
.00
DASHBOARD
Your
Endpoints
Live status, metrics, and controls for all deployed serverless endpoints.
NAME - MODEL
STATUS
HARDWARE
REQUESTS / MIN
AVG LATENCY
llama-3-70b-instruct
prod-inference-v3
Active
H100 SXM5
1,240
184ms
Manage
mixtral-8x22b
staging-moe-test
Cold
A100 80G
0
—
Manage
whisper-large-v3
transcription-api
Active
A10G
327
420ms
Manage
sdxl-turbo
image--gen-exp
Paused
A100 40G
0
—
Manage
Model Name
sub-id
Active
H100
0
-
Manage
HOW IT WORKS
Serverless
Lifecycle
From first request to streamed response — every stage is
optimized for the lowest possible latency.
Request arrives
Client sends inference request to the global endpoint URL with auth token.
Gateway routes
Global load balancer authenticates, rate-limits, and selects the nearest warm worker.
GPU warms up
If cold, container spins up in <340ms. Model weights stream from cache to GPU VRAM.
Inference runs
Model executes on GPU. Output streams back token-by-token via SSE or full JSON response.
Idle -> sleep
After your configured TTL, the container hibernates. You stop paying immediately.
PERFORMANCE
Cold Start
Breakdown
We’ve obsessively optimized every millisecond of container initialization.
Total cold start time • p50
340ms
H100 • Llama 3 70B • 8-bit quantized
Oms
60ms
180ms
280ms
340ms
Container start
Model cache fetch
vRAM loading
First token
~340ms
p50 cold start
~680ms
p99 cold start
~18ms
p50 warm latency
10GB/s
model cache speed
INTEGRATION
Call your endpoints in 5 lines
OpenAI-compatible API — drop in your endpoint URL and start inferencing.
SECURITY
Enterprise-Grade
Security
Every endpoint is isolated, encrypted, and auditable by default.
API Key Scoping
Fine-grained key permissions: read-only, endpoint-specific, IP-restricted, or rate-limited keys. Rotate without downtime.
End-to-End Encryption
All requests encrypted via TLS 1.3. Inference data never persisted. Container memory wiped after each request batch.
Full Audit Log
Every request logged with timestamp, token usage, latency, and originating IP. Export to S3 or stream to your SIEM.
VPC Isolation
Deploy endpoints inside your private VPC. Zero public internet exposure. Private link support for AWS, GCP, and Azure.
PRICING
Pay for what
you actually use
Billed per millisecond of actual GPU execution time. Zero charges
when your functions are idle.
Launch your first
Deployment Today
GPU-powered workloads, run serverless AI functions, and access production-
ready model endpoints — all on one unified platform.
Start Deploy
Schedule a Demo
PRODUCT
Cloud GPUs
Serverless
Public Endpoints
Hub
RESOURCES
Blogs
Case Studies
Referral Program
Articles
Pricing
COMPANY
About
Contact
Careers
Privacy Policy
Terms & Conditions
© 2025 Created with
Royal Elementor Addons
Facebook-f
X-twitter
Tumblr
Instagram