Funding
Modulate Raises $25M to Build the Intelligence Layer for Voice AI

Modulate has raised $25 million in new funding as the Boston-based company looks to capitalize on a growing challenge in artificial intelligence: teaching machines to understand not only what people say, but how they say it.
Future Ventures led the round, with participation from Hyperplane and Lakestar. The financing brings Modulate’s total funding to $60 million and will be used to expand its AI research, engineering teams, developer tools, partnerships, and deployment options.
The investment comes as voice increasingly becomes an interface for AI agents, customer service platforms, gaming environments, and security systems. While speech-to-text technology has improved dramatically, Modulate is betting that transcripts alone leave out much of the information contained in human speech.
Its technology instead analyzes the underlying audio for signals such as emotion, tone, intent, pauses, synthetic speech, and conversational behavior.
Moving Beyond Speech-to-Text
Much of today’s voice AI stack follows a relatively straightforward architecture: convert speech into text and then send the resulting transcript to a large language model (LLM).
That works well when the words themselves contain most of the relevant information. It becomes more limiting when meaning depends on sarcasm, frustration, hesitation, stress, interruptions, or changes in tone.
Modulate has built its flagship Velma platform around the idea that these signals should be analyzed directly from the audio rather than reconstructed afterward from a transcript.
The company describes Velma as an audio-native conversation intelligence system capable of producing transcripts while simultaneously identifying speakers, emotions, conversational topics, sentiment, and more than 150 behaviors. Developers can also define custom behaviors using natural-language descriptions.
That opens the technology to applications ranging from identifying a frustrated customer before a call deteriorates to detecting suspicious behavior during a financial transaction or identifying when an AI voice agent is behaving unexpectedly.
In June, Modulate opened Velma directly to developers through an API, widening access beyond its earlier enterprise deployments.
Modulate Is Taking an Ensemble Approach to Audio AI
The architecture underneath Velma is what Modulate calls an Ensemble Listening Model, or ELM.
Rather than using one enormous model to perform every task, ELM coordinates more than 100 specialized models, with individual components analyzing different characteristics of an audio stream.
Those models can examine basic properties such as speaker changes and pauses alongside higher-level signals such as emotion, perceived intent, synthetic voice markers, confusion, or behavioral patterns. An orchestration layer then combines those observations into a broader interpretation of the conversation.
It is a notably different philosophy from the increasingly common strategy of applying ever-larger multimodal foundation models to audio.
Modulate argues that specialization allows its architecture to deliver more transparent outputs while reducing the amount of computation required. For enterprise applications such as fraud detection and compliance, the ability to understand why a system raised an alert can be nearly as important as the alert itself.
According to the company, the approach can be up to 1,000 times more computationally efficient than using a single large model for some audio-analysis workloads.
Voice AI Creates a New Monitoring Problem
The timing of the funding also reflects a broader shift taking place across enterprise AI.
AI systems are increasingly capable of conducting complete voice conversations with customers. That means companies now have to monitor interactions involving not only human employees, but autonomous agents operating across thousands or potentially millions of conversations.
A voice agent might technically follow a script while still repeatedly interrupting customers, failing to recognize frustration, misunderstanding intent, or responding poorly to emotional situations.
Modulate sees an opportunity to become an independent listening layer sitting alongside those agents.
Its voice-agent technology is designed to identify signals such as interruptions, pauses, emotions, accents, and other characteristics that are difficult to capture from text alone, allowing developers to adjust an agent’s behavior while conversations are still happening.
That same infrastructure can be applied to human conversations where organizations need to identify fraud, compliance risks, harassment, or other potentially significant events in real time.
From Gaming Moderation to Broader Voice Intelligence
Modulate’s current ambitions represent a significant expansion from its earlier focus on online gaming.
One of its best-known products, ToxMod, was developed to analyze live voice communications on gaming and social platforms and identify harassment, threats, grooming, and other harmful behavior.
That remains an important part of the business. Modulate says ToxMod can integrate directly with existing voice infrastructure while routing potentially problematic interactions to moderation systems or human reviewers.
The company has worked with major gaming companies including Activision, Riot Games, Rockstar Games, and Rec Room.
But the underlying technology required to understand toxicity in a multiplayer game can also be adapted to other environments where context matters.
Modulate now positions its technology across fraud prevention, contact centers, AI-agent supervision, deepfake detection, customer experience, and trust and safety.
Its developer platform currently offers multiple model families spanning transcription, deepfake detection, audio redaction, conversation intelligence, and other audio-analysis capabilities.
Deepfakes Add Another Reason to Understand Audio Directly
Synthetic speech is becoming another important part of that opportunity.
As increasingly convincing voice cloning becomes available, organizations handling financial transactions or sensitive information need to determine not only what a caller is saying but whether the voice itself is authentic.
Modulate has consequently expanded into deepfake detection, combining synthetic-voice analysis with the wider contextual information available through its audio models.
The company says its technology currently analyzes more than 10 million hours of audio per month and has processed more than 600 million hours overall.
Its expansion into areas such as AI-generated music detection also illustrates how broadly the underlying ensemble architecture could potentially be applied. Modulate’s music detection system, for example, separately evaluates the presence of music, AI-generated vocals, and AI-generated instrumentation before combining the signals into an interpretable result.
The Next Challenge for Voice AI Is Understanding, Not Speaking
The $25 million investment gives Modulate additional resources to expand its research, engineering, developer tools, and partnerships. More broadly, it reflects a growing challenge for voice AI: speaking naturally is only half of a conversation.
As voice agents move into customer service, healthcare, financial services, and other fields, systems will need to understand signals that transcripts often miss, including frustration, hesitation, emphasis, tone, and possible synthetic speech.
That could create a new intelligence layer around voice interactions, helping systems detect fraud, recognize when customers are dissatisfied, or identify when an AI agent is misunderstanding or mishandling a conversation.
But the technology also raises questions around privacy, accuracy, and bias. Inferring emotion or intent from speech is more subjective than transcribing words, particularly across different accents, languages, and cultures.
As AI becomes increasingly capable of speaking like a person, the next frontier may be determining how well machines can actually listen—and how responsibly those capabilities are used.












