Models & Research

Modulate raises $25M on a transcription lead it lost in August

Boston voice AI company Modulate announced $25 million in new funding on 28 September, and its press release says the company “recently earned the number-one position on Hugging Face’s Open ASR Leaderboard for transcription.” The version of that leaderboard published three days before the announcement put Modulate’s entry third. Zoom and Microsoft held the two places above it.

The round was led by Future Ventures with participation from Hyperplane and Lakestar, and Modulate says it brings total funding to $60 million. SiliconANGLE reported the same $60 million figure, and added that Hyperplane backed a $2 million seed and Lakestar led a $30 million Series A in 2022.

That total doesn’t reconcile with the other number in circulation. TechCrunch cited PitchBook data putting Modulate’s prior funding at $41 million, at a $170 million valuation. Add this round and you get $66 million, which is $6 million more than the company’s own figure, and neither source explains the gap. Prior-funding totals are one of the softest numbers in any announcement, as we set out in how to read a round announcement.

Where the transcription model actually sits

The Open ASR Leaderboard ranks models by average word error rate, lowest first, and it keeps every dated version of its English short-form table. We pulled the published results files for each version and sorted them on that metric ourselves. Modulate entered at the top on 10 July and hasn’t led since the version dated 5 August.

Leaderboard versionModulate entryAvg WERRankFirst place
10 Jul 2026modulate/vfast4.431modulate/vfast
5 Aug 2026modulate/vfast4.431modulate/vfast
25 Aug 2026modulate/vfast4.132elevenlabs/scribe_v2
28 Aug 2026modulate/vfast4.054elevenlabs/scribe_v2
25 Sep 2026modulate/multilingual3.843zoom/scribe_v2_pro
2 Oct 2026modulate/multilingual3.843zoom/scribe_v2_pro
Ranks calculated by Rundowns AI from the leaderboard’s published English short-form results files. The dataset mix changed from seven test sets in July to ten in October, so the ranks compare within a version, not across them.

On the current version, dated 2 October, Modulate’s model scores 3.84 against 3.59 for Zoom’s scribe_v2_pro and 3.81 for Microsoft’s azure-speech-07-2026. The gap is narrow, but the model is still third. The entry name also changed along the way, from modulate/vfast to modulate/multilingual, which is the sort of detail that makes a headline score hard to trace back to a specific product. We hit the same problem reading Apple’s bilingual speech results.

Read closely, the press release is more careful than the coverage of it. It says Modulate “recently earned” first place for transcription, then says the company “currently ranks first” on Hugging Face’s deepfake speech benchmark. That tense change is doing real work, because only one of those two claims is written in the present tense.

Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can’t be solved from a transcript.

Carter Huffman, CEO and co-founder, Modulate, via the company’s press release

The prices nobody quoted

SiliconANGLE mentioned batch transcription at three cents an hour, which the press release confirms as $0.03. It didn’t carry the rest of the price list on Modulate’s own benchmarks page, which puts streaming transcription at $0.06 per hour and deepfake detection at $0.25 per hour. The same page prices Resemble AI’s enterprise deepfake tier at $29 per hour and other providers at $30 to $120. Those are Modulate’s figures for its competitors, not ours.

The architecture claim underneath all of this is that Modulate’s Ensemble Listening Model orchestrates more than 100 small audio models rather than one large one, which the release says delivers up to 1,000 times greater efficiency. It also claims the Velma platform gets twice the accuracy of traditional large language models on true positives, with seven times fewer false positives. No rival model is named in either comparison, so there’s nothing in the release to check the ratios against.

What the money buys is spelled out: new SDKs and APIs, industry-specific models, partner integrations, and more deployment options. TechCrunch reported 40 to 45 staff today and about 10 more hires planned. The next leaderboard version is the thing to watch, because Modulate has now tied its own pitch to a table that updates roughly every week. Funding rounds in voice AI keep landing on benchmark claims like this, as Wispr’s $280 million round showed.

Get the daily rundown

One email each weekday with the AI news that matters, every claim linked to its primary source.

Free, one email each weekday, unsubscribe in one click. We never sell or share your address.

Rundowns AI Desk

Rundowns AI Desk covers artificial intelligence: model releases, research, funding and policy. Every story is written from primary sources, with each claim linked to the announcement, filing or paper it came from, and checked against those sources before publication.

Leave a Reply

Your email address will not be published. Required fields are marked *