Modulate is hiring a Senior Machine Learning Engineer to work on voice AI problems across research and engineering to design, train, evaluate, and deploy ML models powering the conversational voice intelligence platform.
Published
Signal category
People
Quote
“We’re hiring a Senior Machine Learning Engineer at Modulate”
— Modulate team
Company
Modulate
Frontier AI lab building Velma — audio-native voice intelligence for developers and platforms
- Industry
- Software Development
- Location
- Somerville, US
- Company size
- 50 employees
Modulate is a frontier AI lab building the voice intelligence layer for the internet. We develop foundation models for audio: models that don't just transcribe speech, but understand it: who's speaking, what's being said, whether it's real, and whether it's safe. Velma is Modulate's audio-native voice intelligence model, purpose-built for audio from the ground up — not a text model retrofitted for speech. Velma powers a suite of APIs for developers building the next generation of voice-first products: → Modulate Transcribe: state-of-the-art speech-to-text. #1 on the Hugging Face Open ASR Leaderboard (out of 88 models), with top results on Sierra's μ-Bench, including #1 for Mandarin (zh-CN) → Deepfake Detect: real-time detection of AI-generated and cloned voices → AI Music Detection: identifying AI-generated audio content at scale → PII/PHI Redaction: automated sensitive-data protection for regulated industries Developers use Velma to build voice products that are faster, more accurate, and safer by default — without stitching together brittle, single-purpose audio tools. ToxMod, our proactive voice moderation system, applies this same model foundation to real-time safety — trusted by studios including Ubisoft, Activision, and Schell Games to detect toxicity and harassment without relying on player reports. Modulate was founded on a simple premise: audio is its own modality, and it deserves models built natively for it - not adapted from something else.
Founded 2019