Interviews
Phil Marshall, Founder and CEO of Spoken – Interview Series

Phil Marshall, Founder and CEO of Spoken, is a physician, inventor, technology entrepreneur, and science fiction writer whose career has spanned healthcare, digital media, artificial intelligence, and product innovation. Before launching Spoken in 2024, Marshall co-founded Conversa Health, a personalized patient engagement and conversational AI company that was acquired by Amwell in 2021 and became part of its automated virtual care business. Earlier in his career, he spent more than a decade at WebMD as Vice President of Product Strategy, later served as Senior Vice President of Product Management at Press Ganey Associates, and founded collaborative video technology startup JumperCut. His work increasingly focuses on the intersection of AI, storytelling, brain-computer interfaces, and knowledge technologies, combining his background as a technologist and entrepreneur with his interests in science fiction and emerging forms of digital interaction.
Spoken is an AI-powered audiobook platform designed to help authors, publishers, and rights holders transform manuscripts into professionally produced audio without relying on a traditional recording studio. Its technology analyzes a manuscript’s language, style, narrative structure, and characters before assigning distinct voices for narration and dialogue, supporting single-narrator, dual-narrator, and multi-cast productions. Authors can use custom-generated voices, select from more than 100 paid professional voice actors, or clone their own voice, while retaining ownership of their work, masters, and distribution rights. Spoken also offers an application programming interface (API) aimed at publishers seeking to produce audiobooks across larger catalogs, while emphasizing an ethical AI model under which it says written works are not used to train its systems and participating human voice actors are compensated.
Your career has taken you from medicine and product strategy at WebMD to founding Conversa Health and now Spoken. What problem in audiobook creation convinced you to found Spoken, and how did your previous work in conversational AI influence the platform you are building today?
With Conversa, our mission was to extend the care team’s voice by automating personalized outreach, follow-up questions, and guidance that a clinical department at a large health system would provide directly to patients by phone, if they had the time. Initially, we anthropomorphized the bot with a persona, Cate, who referred to itself in the first person. Ultimately, however, I decided that feigning humanity was disingenuous. Cate wasn’t a person, and I didn’t want patients to think it was. So, we removed the persona name and shifted to a plural “We” for it to be seen as an extension of the care team, which it was. The experience taught me that automated solutions can still feel like an authentic extension of human care without pretending to be human.
Fast forward to Spoken, where we focus on authentic, human-centered expression. The problem we help solve is obvious – most stories can’t reach the fastest growing audience (listeners) with their full audio potential due to time and cost constraints – but my longstanding philosophy to be a genuine extension of humans played a big role. Unlike AI podcasts that try to manufacture banter or agents that simulate compassion, Spoken’s Multi-Cast narration brings an author’s true words to life. Because we are telling stories, not trying to replicate human interactions, there’s no pretense that our AI is a person. These are characters in a story, so we are fortunate to avoid any confusion and conflict.
As both a hard science fiction writer and an AI entrepreneur, how did you once expect artificial intelligence to evolve, and where has the reality of AI’s development diverged most sharply from the future you imagined?
This question is near and dear to my heart. I love living at the intersection of art, science, and tech, and depicting automated systems in my near-future sci-fi is a great example of that. In fact, every future company I write about – and there are several – has a domain that I’ve owned, a product that I’ve designed, and a concept that I believe could someday be a real company. Even the sculpture in front of my house is modeled after the modernist shape of The Kite Factory logo, a company I imagine will one day create the anti-gravity technology that will transform the planet!
Focusing on AI, I have a problem when future sci-fi doesn’t address how people derive, use and share information. In the year 2100, are they really still using a phone to call people to answer questions of fact? My philosophy in writing, and in real life, is simple: What can be automated will be automated, and we will always trend toward more real-time, more collaborative, and more physically integrated information. I should add, however, that I have a personal aversion to implants. This is why, in the backstory of my book, when Elon Musk hooked up his Starlink satellites to his Neuralink implants and began broadcasting messages directly into people’s brains, we outlawed brain implants!
To the question of divergence: because I believe what can be automated will be automated, automating questions of fact is happening as a simple inevitability. However, when it comes to matters of the heart – art, music and stories – there has been a bit of a divergence. Because AI is trained on pre-existing patterns, by definition, it will not ever be able to create something truly new. Yet people are writing stories, creating art, and composing music using AI. How can that be? Because art does not have to be truly new or unique. That, I believe, will mean that in the future, those creations that do break the mold will be even more precious. I hope we’ll still be able to recognize when that happens.
A recent Edison Research study compared Spoken’s multi-cast production with a professionally produced, single-narrator human version. How much of Spoken’s advantage do you attribute to the underlying AI technology, and how much comes from giving each character a distinct voice? How might the results differ against a full human cast?
Analyzing the story and every character to derive their perfect voice, is in fact, a big part of our AI. We go through an intensive agentic process to derive the underlying layers of the story, from rudiments like genre and language to the cadence dictated by accents and style, as well as the scene cohesion of multiple characters interacting. All of these AI-driven layers inform the actual character voice narration. It’s what we call “Magic Mode”, and you can think of it like blending paint colors.
When it comes to how the results might differ from a full human cast, it really comes down to cost and workflow. Single-click and hundreds of dollars, compared to a full cast, is not accessible to indie authors and most publishers, which is why it wasn’t our point of comparison. As for quality, however, I believe we’re already approaching parity with full cast, so quality will be indistinguishable.
Spoken Multi-Cast received higher ratings for character-driven scenes, while human narration performed better during exposition. What makes non-dialogue passages particularly difficult for AI narration, and how are you working to close that gap?
Actually, it’s not that non-dialogue passages are difficult for AI narration; it’s just that the number of dynamics the AI has to work with is still less than what a human voice actor has to work with. Think of every nuance of inflection and delivery that a single narrator can bring to a passage. With multi-cast, however, the number of dynamics is far greater because of the varied, personalized timbres of character voices and that process of mixing them, again, like paint colors. That means multi-cast scenes can bring forward the realism of interactions between multiple people in a way a single narrator would be hard pressed to match.
As for closing the gap, the number of dynamics used in single-narrator AI narration will continue to increase, and the gap will continue to shrink.
Before hearing the samples, only 31% of participants expressed interest in AI-narrated audiobooks, but that figure rose to 65% after they experienced Spoken Multi-Cast. What elements of the performance do you believe were most responsible for changing their perceptions?
Actually, it was the other way around, and that design element of the survey is super important. 65% said they’d listen to a whole audiobook with this narration after hearing the excerpts, before knowing it was AI. We then asked, more generally, whether they’d be willing to listen to audiobooks narrated using AI, and only 31% said yes. Next, we asked whether they thought what they had heard was AI. Only 39% of those who heard Spoken Multi-Cast suspected it might be AI, compared to 35% of those who heard the human versions and thought theirs might be AI. This made our next question pop: Now that you know this, would you be willing to listen to an audiobook narrated with AI? That willingness jumped to 43% post-exposure, and we are actively optimizing to close the gap toward our 65% benchmark.
Can you walk us through the agentic AI workflow that transforms an uploaded manuscript into a multi-voice audiobook, including how the system identifies speakers, interprets characters, assigns voices, and determines emotion, timing, and delivery?
Given our pending patents on the process, I’ll simply summarize it at a high level by saying it is an intensely agentic process that delivers multiple layers. It’s a process we’ve honed over more than two years. The rudiments of the story, the cadence of the story, the scene cohesion, and finally the precise character voices, or as I like to call them, “the coat of paint,” are all part of it. Getting the arc of emotional delivery and timing is critical, and we work on that more than anything. It often surprises people that the AI used to prepare for narration is about 5x the amount of AI used in the actual narration.
How much creative control does an author retain over the final production, and what tools are available when the AI misinterprets a character, pronunciation, emotional beat, or narrative intention?
The author has absolute control over the final production. Whether it’s the emotional inflections, accents, voice timbre, or timing, they control it all. Our goal is to get them as close to exquisite in a single click as we can. If the author, publisher or producer needs a very specific delivery of a passage, they can use “Speak It” to speak into their microphone and demonstrate how they want it delivered, and the character voice will obey beautifully.
We believe authors should feel full ownership of their work, down to the voice used in their opening and closing credits. Even “Made with Spoken” will use the voice of their choice, including their own personal voice if they choose so.
Spoken allows authors to use custom-generated voices as well as approved voice clones from professional narrators who are compensated when their voices are used. How are consent, ownership, compensation, revocation rights, and protection against unauthorized cloning built into that model?
We only use voice partners who adhere to a strict code of ethics, which includes detection and prohibition of unauthorized use of a voice (especially important to protect notable voices of celebrities). For example, we have many of the top voices from ElevenLabs in our library. Those voice actors get paid for every use by our authors.
Do you see AI narration primarily expanding the audiobook market by making previously uneconomical titles viable, or will it also replace portions of traditional production? Which roles will remain distinctly human as the technology matures?
We envision a vastly expanded audiobook market that attracts new readers, especially given our new, vivid multi-cast capabilities. In the end, that expanded market will be driven by a hybrid of technology and humans. Time will tell, but this transformation may be less about whether it is human voice talent or AI, and more about the workflow of author/producer control and the ever-changing listening habits of the audience.
Spoken combines audiobook production with publishing, discovery, streaming, community, and monetization. How could AI ultimately change storytelling itself, rather than simply making the existing audiobook production process faster and less expensive?
What does it mean for a reader or listener to “get lost in story?” That’s what this is really all about: giving audiences an immersive audio experience at a time when demand for character-driven fiction is growing and audio is becoming the dominant modality of consuming stories. In the end, whether it’s a short story, personal memoir, fanfic or epic fantasy, the speed and low cost of creating vivid, natural delivery will have profound long-term implications. The current audiobook market is the near-term opportunity, but short-form, vertical shorts, serials, fan fic, and any number of offshoots will thrive as a result. And when they do, we want Spoken Multi-CastTM to be at their heart.
Thank you for the great interview, readers who wish to learn more should viist Spoken.












