Deel reported that DeepSeek's Multi-head Latent Attention (MLA) technique reduced the KV cache by more than 90% compared to standard attention.
Public source
Publisher name
Public post
Continuing on the thread of LLM efficiency - a clever trick called Multi-head Latent Attention (MLA) makes DeepSeek's models cheaper to run. We've covered the KV cache a…
Company
Deel
The Global-First HR Platform
- Industry
- Human Resources Services
- Location
- San Francisco, US
- Company size
- 5,001–10,000 employees
About Deel
Deel is one platform for payroll, HR, benefits, mobility, performance, and device management across 150+ countries. Built on owned infrastructure, powered by AI, and supported by thousands of local experts, Deel helps businesses scale smarter, faster, and more compliantly. Trusted by 40,000+ customers, and created to become a global brand people love. Learn more at deel.com.
See moreLatest activity
Latest activity from Deel
137 signals
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Magic
Magic developed a new pretraining recipe for Frontier that matches DeepSeek V4 Pro's pretrain using 50x less compute, roughly half the FLOPs used for GPT3, or ~$0.5M on GB200.
Research & Knowledge
Oxmiq Labs
Oxmiq Labs noted that DeepSeek's R1 model, released in January 2025, was trained on NVIDIA H800 chips and demonstrated competitive performance.
Research & Knowledge
Together AI
Together AI published research on QLoRA to compress the base model to 4 bits and reduce memory usage to 1/4, allowing the base model plus 16-bit adapters to fit on a smaller GPU.
Research & Knowledge
Weave
Weave reports that about a third of the requests running through its router went to lower-cost models like Deepseek Flash.
Research & Knowledge
Cint