Google llm-d delivered a 44% increase in total throughput and a 39.8-second P90 Time-to-First-Token (TTFT) down to 0.399 seconds on multimodal workloads compared to standard Kubernetes services.
Public source
Publisher name
Public post
🚀 Multimodal support in llm-d. Multimodal requests break traditional text inference infrastructure. While text strings are discrete and cheap to tokenize on a CPU, medi…
Company
- Industry
- Software Development
- Location
- Mountain View, US
- Company size
- 10,001+ employees
About Google
A problem isn't truly solved until it's solved for all. Googlers build products that help create opportunities for everyone, whether down the street or across the globe. Bring your insight, imagination and a healthy disregard for the impossible. Bring everything that makes you unique. Together, we can build for everyone. Check out our career opportunities at goo.gle/3DLEokh
See moreLatest activity
Latest activity from Google
1,041 signals
Presence & Recognition
Google hosted the Make AI Work for You tour on the road traveling Route 66 to bring hands-on AI training to small business owners across five states.
Products & Services
Google DeepMind introduced AlphaGenome Atlas, an AI-powered searchable database mapping predicted genetic impact changes across the human genome.
Presence & Recognition
Google employee Yujun Liang attended PyTorch Conference Community Night in Shanghai with the Bund as the backdrop.
Discover more
Similar signals
Similar public activity from other companies.
Research & Knowledge
Databricks
Databricks found that Llama 3.1 8B Instruct was the fastest model in the synthetic workload, but only 34 out of 100 passed.
Research & Knowledge
DeepLearning.AI
DeepLearning.AI published an analysis in The Batch evaluating throughput and latency for large language models.
Research & Knowledge
Cloudian Inc
Cloudian Inc published a new benchmark run with NVIDIA Dynamo that shows KV Cache offloading to Cloudian HyperStore reduces time to first token by up to 20x compared to full recompute at 120K tokens.
Research & Knowledge
Alkami Technology
Alkami Technology performed performance testing Qwen, GLM, and Llama on vLLM under production-like concurrency, resulting in 1.92x more requests meeting the same latency SLO on the same GPU budget or roughly 48% lower effective GPU cost per successful request.
Research & Knowledge
Red Hat