ABOUT US

We make GPU
compute accessible to every builder on Earth.

Push your AI code to a repo. We handle containers, scaling, cold
starts, and infrastructure — you ship features.

founded
2000
team members
0
global regions
0

OUR MISSION

Democratizing GPU Infrastructure

Empowering developers, startups, and enterprises with scalable
GPU compute built for modern AI workloads.

"The next breakthrough in AI shouldn't depend on who can afford a $100,000/month cloud bill."

We built NovAI after watching brilliant researchers and indie developers get priced out of the compute needed to turn ideas into reality. The hyperscalers made GPU compute complex, expensive, and opaque.
Our answer: a serverless-first platform where you call an API, your model runs on H100 silicon, and you pay only for the milliseconds of compute you actually used. No clusters to manage. No reserved capacity to overpay for.

Radical transparency

Our pricing, uptime, and incident history are always public—because transparency should be the standard.

Developer-first always

Every decision runs through one filter: does this make a developer's life meaningfully easier? If not, it doesn't ship.

Speed as a feature

Cold starts under 340ms, deploys under 90 seconds. Latency isn't a tradeoff we accept — it's an engineering problem we obsess over.

Access, not gatekeeping

A researcher at a university and an engineer at a trillion-dollar company should have the same access to GPU infrastructure. Full stop.

Observable by default

Every request is logged, every GPU is metered, and every latency spike is visible in your dashboard — no black boxes.

Sustainable compute

Our serverless model eliminates idle GPU waste by design. Every watt of GPU power we serve traces to renewable energy.

BY THE NUMBERS

Impact at GPU Scale

40K+
ACTIVE DEVELOPERS
Running production workloads across 120 countries
2.4B
INFERENCE REQUESTS / MONTH
Growing at 18% month-over-month
340ms
COLD START P50
Down from 2,400ms at launch in 2021
40K+
PLATFORM UPTIME (12 MO)
Exceeding our 99.95% SLA commitment

THE TEAM

Built by people who
felt the pain.

Created by engineers who understand the realities of building, deploying,
and scaling AI products.
Jago Davenport

Jago Davenport

Finance
Former ML infra lead at Google Brain. Built distributed training pipelines for PaLM. Obsessed with cold start latency and the Rust programming language.
Patricia Wilson

Patricia Wilson

Finance
Former ML infra lead at Google Brain. Built distributed training pipelines for PaLM. Obsessed with cold start latency and the Rust programming language.
Garfield Potter

Garfield Potter

Finance
Former ML infra lead at Google Brain. Built distributed training pipelines for PaLM. Obsessed with cold start latency and the Rust programming language.

CULTURE & WORK

A place where
infrastructure nerds thrive.

We’re remote-first, async by default, and deeply skeptical of meetings that could’ve been a Notion doc.

Remote-first, truly

Our team spans 14 timezones. No office. No headquarters. Your best work environment is wherever you do your best thinking.

Async by default

We document decisions, write before meetings, and trust each other to read. Your focus time is protected by design, not by accident.

Remote-first, truly

Our team spans 14 timezones. No office. No headquarters. Your best work environment is wherever you do your best thinking.

Async by default

We document decisions, write before meetings, and trust each other to read. Your focus time is protected by design, not by accident.

Start building on
faster infrastructure today.

Deploy your first serverless endpoint in minutes. Start building instantly with a streamlined, developer-first experience.

PRODUCT

Cloud GPUs

Serverless

Public Endpoints

Hub

RESOURCES

Blogs

Case Studies

Referral Program

Articles

Pricing

COMPANY

About

Contact

Careers

Privacy Policy

Terms & Conditions