Company intelligence
Yantrion Inc
More AI. More tokens. Same GPUs.
About Yantrion Inc
Yantrion helps AI teams run more work on the GPUs they already have. Our software reduces the GPU memory needed to retain a model’s working context and optimizes how that context is used during inference. This creates room for more concurrent requests, longer context and more agents ready to resume work. We’re turning those memory savings into measured throughput gains, helping teams deliver more tokens per second from the same hardware. Built for inference providers and enterprise AI teams, Yantrion integrates with existing serving stacks, including SGLang and vLLM, on supported AMD and NVIDIA GPU configurations. More AI. More tokens. Same GPUs.
Verified activity
Signals from Yantrion Inc
3 published signals
Products & Services
Yantrion Inc started by helping AI models use GPU memory more efficiently, resulting in a promising first result with Kimi-K3.
Reported by Yantrion Inc
Research & Knowledge
Yantrion Inc measured 975.5 aggregate decode tokens per second across 56 concurrent requests on one 8× AMD Instinct MI350X node using Kimi-K3.
Reported by Yantrion Inc
Products & Services
Yantrion Inc is starting to turn GPU memory savings into more tokens per second on the same hardware.
Reported by Bhagawan Gnanapa