InferX built a single-tenant architecture with VM-level isolation instead of shared clusters from day one.
Public source
Publisher name
Public post
This tracks with what we’ve seen across the industry. Most neoclouds run shared multi-tenant GPU infrastructure — which is exactly where these container escape and isola…
Company
InferX
True Serverless GPU Inference Platform
- Industry
- Technology, Information and Internet
- Location
- Seattle, US
- Company size
- 2–10 employees
About InferX
Most inference platforms sell you compute. InferX gives you control of it. We built InferX because most GPU spend is wasted — models sitting loaded in VRAM waiting for traffic that isn’t there, and painfully slow cold starts when it finally shows up. InferX is an inference runtime that virtualizes GPU execution — CUDA and NCCL — at the boundary, not just the scheduler. That’s what makes sub-second cold starts possible: instead of rebuilding CUDA state, loading weights, and re-establishing multi-GPU topology from scratch, InferX snapshots an initialized inference system and restores it in milliseconds. In production, we’ve measured 494ms cold-start-to-first-token on a 27B model at full precision — a metric we’ve submitted to NeurIPS as a standardized benchmark. One snapshot becomes four production primitives: cold start, scale-out, GPU migration, and failure recovery — all from the same mechanism. Because the runtime controls the boundary all the way through GPU communication, isolation is built into the architecture, not bolted on top. InferX deploys inside your existing Kubernetes clusters — you keep your infrastructure, we handle inference execution underneath it. Inference is a systems problem, not a chip problem. Learn more: inferx.net
See moreLatest activity
Latest activity from InferX
5 signals
Products & Services
InferX intercepts NCCL communication at the runtime boundary to validate it before reaching the underlying GPU communication stack.
Products & Services
InferX launched a new free lineup including Gemma 4 31B, GPT-OSS 20B, Qwen3.6 35B A3B, Agents A1, Devstral 2 123B, and Qwen3 Coder Next.
Products & Services
InferX launched DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal model matching DeepSeek V4 Flash on text capabilities while adding vision, bringing performance close to Opus-4.8.
Discover more
Similar signals
Similar public activity from other companies.
Products & Services
ScitiX
ScitiX features session-aware orchestration, 3-tier failover, no cold starts, and no bill surprises.
Products & Services
io.net
io.net was built to orchestrate for distributed training, fine-tuning, and serious AI workloads.
Products & Services
FLEXNODE
FLEXNODE offers a repeatable, scalable inference compute architecture built for structural adaptability, with deployments configured for site, power, and chip architecture.
Products & Services
Gimlet Labs
Gimlet Labs built an inference cloud from the ground up to tackle problems across the entire stack, including datacenter infrastructure, networking, accelerators, compilers, scheduling, and serving.
Products & Services
FlexAI