Modulate raises $25M on a transcription lead it lost in August
Boston voice AI company Modulate announced $25 million in new funding on 28 September, and its press release says the company “recently earned the number-one position on Hugging Face’s Open ASR Leaderboard for transcription.” The version of that leaderboard published three days before the announcement put Modulate’s entry third. Zoom and Microsoft held the two places above it.
The round was led by Future Ventures with participation from Hyperplane and Lakestar, and Modulate says it brings total funding to $60 million. SiliconANGLE reported the same $60 million figure, and added that Hyperplane backed a $2 million seed and Lakestar led a $30 million Series A in 2022.
That total doesn’t reconcile with the other number in circulation. TechCrunch cited PitchBook data putting Modulate’s prior funding at $41 million, at a $170 million valuation. Add this round and you get $66 million, which is $6 million more than the company’s own figure, and neither source explains the gap. Prior-funding totals are one of the softest numbers in any announcement, as we set out in how to read a round announcement.
Where the transcription model actually sits
The Open ASR Leaderboard ranks models by average word error rate, lowest first, and it keeps every dated version of its English short-form table. We pulled the published results files for each version and sorted them on that metric ourselves. Modulate entered at the top on 10 July and hasn’t led since the version dated 5 August.
| Leaderboard version | Modulate entry | Avg WER | Rank | First place |
|---|---|---|---|---|
| 10 Jul 2026 | modulate/vfast | 4.43 | 1 | modulate/vfast |
| 5 Aug 2026 | modulate/vfast | 4.43 | 1 | modulate/vfast |
| 25 Aug 2026 | modulate/vfast | 4.13 | 2 | elevenlabs/scribe_v2 |
| 28 Aug 2026 | modulate/vfast | 4.05 | 4 | elevenlabs/scribe_v2 |
| 25 Sep 2026 | modulate/multilingual | 3.84 | 3 | zoom/scribe_v2_pro |
| 2 Oct 2026 | modulate/multilingual | 3.84 | 3 | zoom/scribe_v2_pro |
On the current version, dated 2 October, Modulate’s model scores 3.84 against 3.59 for Zoom’s scribe_v2_pro and 3.81 for Microsoft’s azure-speech-07-2026. The gap is narrow, but the model is still third. The entry name also changed along the way, from modulate/vfast to modulate/multilingual, which is the sort of detail that makes a headline score hard to trace back to a specific product. We hit the same problem reading Apple’s bilingual speech results.
Read closely, the press release is more careful than the coverage of it. It says Modulate “recently earned” first place for transcription, then says the company “currently ranks first” on Hugging Face’s deepfake speech benchmark. That tense change is doing real work, because only one of those two claims is written in the present tense.
Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript.
Carter Huffman, CEO and co-founder, Modulate, via the company’s press release
The prices nobody quoted
SiliconANGLE mentioned batch transcription at three cents an hour, which the press release confirms as $0.03. It didn’t carry the rest of the price list on Modulate’s own benchmarks page, which puts streaming transcription at $0.06 per hour and deepfake detection at $0.25 per hour. The same page prices Resemble AI’s enterprise deepfake tier at $29 per hour and other providers at $30 to $120. Those are Modulate’s figures for its competitors, not ours.
The architecture claim underneath all of this is that Modulate’s Ensemble Listening Model orchestrates more than 100 small audio models rather than one large one, which the release says delivers up to 1,000 times greater efficiency. It also claims the Velma platform gets twice the accuracy of traditional large language models on true positives, with seven times fewer false positives. No rival model is named in either comparison, so there’s nothing in the release to check the ratios against.
What the money buys is spelled out: new SDKs and APIs, industry-specific models, partner integrations, and more deployment options. TechCrunch reported 40 to 45 staff today and about 10 more hires planned. The next leaderboard version is the thing to watch, because Modulate has now tied its own pitch to a table that updates roughly every week. Funding rounds in voice AI keep landing on benchmark claims like this, as Wispr’s $280 million round showed.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
