Company intelligence
Mirai Labs
About Mirai Labs
We believe in the trinity: model, inference stack, hardware. Companies that focus on a single component of this trinity lack sovereignty and are constrained by the architectural choices made by others. Most labs treat on-device models as scaled-down versions of their cloud-focused cousins. But LLM architectures that evolved for the cloud are not well-suited to on-device setups. Cloud LLMs operate in the arithmetic-bound regime. Mainstream architectures aim to maximise total token throughput by reducing the amount of computation performed per request, and they treat device memory as an unlimited resource. But for on-device deployment, memory is the main bottleneck, both in terms of throughput and the size of the resident set. When designing our on-device architecture, we focus on three core objectives: increasing the arithmetic intensity of the decoding stage, reducing the size of the resident set, and maximally utilising the GPU neural accelerators. This leads us to models that differ from traditional autoregressive transformers in a number of meaningful ways.
Verified activity
Signals from Mirai Labs
9 published signals
Products & Services
Mirai Labs achieved a 27B model run at over 100t/s on a MacBook using optimizations including activations quantization, native A4W4 and A8W8 MXU execution paths, and per-chip tuned matrix-multiplication kernels.
Reported by Eugene Bokhan
Products & Services
Mirai Labs released a speculative decoding implementation in ๐๐๐ for ๐ฉ๐ฒ๐ป๐ป.๐ฏ ๐ฎ๐ณ๐ฉ, with support for ๐ฉ๐ฒ๐ป๐ป.๐ด ๐ฎ๐ณ๐ฉ and ๐ ๐๐๐ฒ ๐ ๐ต๐ผ๐น๐ถ๐บ๐บ๐ฒ๐ฟ coming soon.
Reported by Aleksei Savin
Products & Services
Mirai Labs provides the model ๐ ๐ถ๐ฟ๐ฎ๐ถ-๐ and ๐ ๐ถ๐ฟ๐ฎ๐ถ-๐ for running the model.
Reported by Aleksei Savin
Products & Services
Mirai Labs provides information about the speculative decoding implementation.
Reported by Aleksei Savin
Products & Services
Mirai Labs provides benchmarks for exploring the speculative decoding implementation.
Reported by Aleksei Savin
Products & Services
Mirai Labs outperforms MTPLX and llama.cpp by almost 2x and over 3x at comparable quantization levels on Apple M5 chips.
Reported by Aleksei Savin
Products & Services
Mirai Labs co-designed its speculative decoding implementation with the latest Apple M5 chips to take maximum advantage of GPU Neural Accelerators.
Reported by Aleksei Savin
Products & Services
Mirai Labs released a speculative decoding implementation in Uzu for Qwen3.6 27B, Qwen3.8 27B, and Muse Glimmer.
Reported by Alexey Moiseenkov
Products & Services
Mirai Labs outperformed MTPLX and llama.cpp by almost 2x and over 3x on Apple M5-series chips with their speculative decoding implementation.
Reported by Alexey Moiseenkov