Mirai Labs outperforms MTPLX and llama.cpp by almost 2x and over 3x at comparable quantization levels on Apple M5 chips.
Public source
Publisher name
Public post
Today we are releasing our speculative decoding implementation in 𝘂𝘇𝘂 (https://lnkd.in/eSMheMkd). Initially, for 𝗤𝘄𝗲𝗻𝟯.𝟲 𝟮𝟳𝗕, with support for 𝗤𝘄𝗲𝗻𝟯.𝟴…
Company
Mirai Labs
- Industry
- Technology, Information and Internet
- Location
- San Francisco, US
- Company size
- 11–50 employees
About Mirai Labs
We believe in the trinity: model, inference stack, hardware. Companies that focus on a single component of this trinity lack sovereignty and are constrained by the architectural choices made by others. Most labs treat on-device models as scaled-down versions of their cloud-focused cousins. But LLM architectures that evolved for the cloud are not well-suited to on-device setups. Cloud LLMs operate in the arithmetic-bound regime. Mainstream architectures aim to maximise total token throughput by reducing the amount of computation performed per request, and they treat device memory as an unlimited resource. But for on-device deployment, memory is the main bottleneck, both in terms of throughput and the size of the resident set. When designing our on-device architecture, we focus on three core objectives: increasing the arithmetic intensity of the decoding stage, reducing the size of the resident set, and maximally utilising the GPU neural accelerators. This leads us to models that differ from traditional autoregressive transformers in a number of meaningful ways.
See moreLatest activity
Latest activity from Mirai Labs
11 signals
Products & Services
Mirai Labs achieved a 27B model run at over 100t/s on a MacBook using optimizations including activations quantization, native A4W4 and A8W8 MXU execution paths, and per-chip tuned matrix-multiplication kernels.
Products & Services
Mirai Labs released a speculative decoding implementation in its inference engine Uzu for Qwen3.6 27B, Qwen3.8 27B, and Muse Glimmer.
Products & Services
Mirai Labs outperformed MTPLX and llama.cpp by almost 2x and over 3x at comparable quantization levels on Apple M5-series chips.
Discover more
Similar signals
Similar public activity from other companies.
Products & Services
Sevren
Sevren ported FlashPrefill V2 to MLX, bringing block-sparse attention to Apple silicon, achieving 1.78× faster end-to-end prefill at 32K tokens, running Qwen3-4B on an M4 Pro from 105 seconds to under a minute.
Products & Services
ScitiX
ScitiX reports 1s average time to first token, 93.9% KV-cache hit rate, 99.9% uptime, and 72% average cost savings.
Products & Services
FriendliAI
FriendliAI delivers 301 tokens per second throughput for the Z.ai GLM-5.3 model, nearly 45% higher than the runner-up.
Products & Services
Unsloth AI
Unsloth AI made GLM-5.3-Flash run 3.3x faster locally with optimized decoding and bonus multi-token prediction support.
Products & Services
Apodex