Mirai Labs released a speculative decoding implementation in 𝘂𝘇𝘂 for 𝘩𝗲𝗻𝗻.𝟯 𝟮𝟳𝑩, with support for 𝘩𝗲𝗻𝗻.𝟴 𝟮𝟳𝑩 and 𝘠𝘂𝘀𝗲 𝘠𝗵𝗼𝗹𝗶𝗺𝗺𝗲𝗿 coming soon.
Published
Signal category
Products & Services
Quote
“Today we are releasing our speculative decoding implementation in 𝘂𝘇𝘂 (https://lnkd.in/eSMheMkd).”
— Aleksei Savin|Mirai Labs team
Company
- Industry
- Technology, Information and Internet
- Location
- San Francisco, US
- Company size
- 15 employees
We believe in the trinity: model, inference stack, hardware. Companies that focus on a single component of this trinity lack sovereignty and are constrained by the architectural choices made by others. Most labs treat on-device models as scaled-down versions of their cloud-focused cousins. But LLM architectures that evolved for the cloud are not well-suited to on-device setups. Cloud LLMs operate in the arithmetic-bound regime. Mainstream architectures aim to maximise total token throughput by reducing the amount of computation performed per request, and they treat device memory as an unlimited resource. But for on-device deployment, memory is the main bottleneck, both in terms of throughput and the size of the resident set. When designing our on-device architecture, we focus on three core objectives: increasing the arithmetic intensity of the decoding stage, reducing the size of the resident set, and maximally utilising the GPU neural accelerators. This leads us to models that differ from traditional autoregressive transformers in a number of meaningful ways.