Audibly Silenced

Audibly Silenced

AI Narration vs. the Human Voice

 

words by Rydwan Anwar for

Literature

Audiobooks have been having a ‘moment’ for a while now. Their rise in recent decades has shaken the literary world and the wider cultural spheres of the performing arts and broadcast. It has changed the ways writers, publishers and booksellers work and opened up a range of possibilities in literature and production. As with any trending phenomenon, commercial interest comes hot on its heels with new entrants seeking to cash in. Enter Artificial Intelligence. While the debate on AI-written novels rages on, the less prominent but equally divisive conversation about AI narration of audiobooks is taking place. As AI narration becomes cheaper, more realistic and commonplace, what happens to the relationship between the human voice and literature itself? Besides the valid ethical questions and concerns surrounding generative AI, we have to ask ourselves; why does human narration matter?   

One of my earliest memories is of my mother reading me bedtime stories. We had a pile of hand-me-down English storybooks, from Ladybird fairy tales to Southeast Asian legends to Aesop’s Fables. My mother began her education in English towards the end of British colonial rule in Malaya and Singapore. English was not her native tongue and did not come naturally to her, but she tried her best to read to me in English, in addition to Malay. It wasn’t perfect, but she made those stories come alive, and instilled in me a love for language and literature. Unlike a novel, a children’s storybook is intended to be read aloud. The adult reader taps into their own experience — how would an elderly grandmother speak? Or into their imagination — what does a mermaid or the Gruffalo sound like? There is history and lineage in the way an adult reader narrates to a child, adding their own spin and personal touch to bring the story to life and keep their audience enraptured.

The audiobook is another form of oral storytelling. Although the modern novelist seldom hears their words read aloud, traditional storytelling has existed since ancient times — oral history and epic tales were passed down through generations long before the advent of the written form and the printing press. As late as the early 20th century, Singapore had street storytellers who would charge a few cents to tell stories to their mostly illiterate audience. As technology caught up, some of those storytellers later moved into radio broadcasting. The audiobook is a dramatized reading of a novel and, unlike solitary reading, is experienced through the prism of another human being. In recent years, the boom in audiobooks has been bolstered by the casting of celebrities as narrators. Elisabeth Moss narrated The Handmaid’s Tale, Benedict Cumberbatch has an entire library shelf to his name, and there was even a cast of 166 famous names for Lincoln in the Bardo. Audiobooks are aural theatre, with a similar creation process. At the heart of it is the human experience. So, when the AI voice replaces the human voice, what do we lose?  

In 2026, a thirteen-hour audiobook of Homer’s The Odyssey was released. It had a cast of synthetic voices, AI-generated music and sound effects. Its selling point was the officially licenced AI voice replica of the actor Michael Caine. Caine consented to the use of his vocal likeness, but did not perform the text in a recording studio. Caine’s voice tells the tale, but it is not his personal interpretation of the lines. As a former theatre producer, I had many questions upon hearing of this. My cynical first thought was that it was a commercial gimmick, banking on a celebrity name for virality. But I mulled over it further. What is the intention and motivation of using a particular actor’s voice? What does Michael Caine’s distinctive cockney accent add to a reading of The Odyssey? Was it a statement to break down class stratifications and pre-conceived notions of how classical texts should be delivered? What is voice without context?  I don’t think these were considerations for the producers. If a listener does not know Michael Caine or his work, would his voice matter? Why ask Michael Caine, besides the fact that he is famed for his oratorical trait? So, I listened to an excerpt. Sir Michael would not lose sleep over my opinion, and I’m certain he has his reasons for allowing his voice to be cloned. But I will state adamantly that the AI voice did him a disservice, and he would have done a much better job of it himself. 

As a theatre producer, I sat through hundreds of auditions. It is one of the most interesting parts of the casting process; a first glimpse of the creativity with which an actor will bring a text to life. It is not merely about how they sound, but their ability to understand the text, its subtext and nuance. Casting the right actors means the difference between an acceptable production and a great one. Another exciting milestone in theatre-making is the table read, when a script is read by the cast together for the first time. Prior to this moment, having read the script alone, the story and its characters dwell only in their heads. They know the plot and its twists and turns, the punchlines, and where the emotions will hit, but the table read will bring surprises. The actors discover new ways of expressing their lines as they hear others’ differing interpretations. They might react differently and develop a richer interpretation of their characters. Each actor brings their own interpretation, drawn from their training and their own life experiences. At later rehearsals, they explore different ways of delivering the text — a longer breath here, a shorter pause there, a gasp, a crack in the voice. It is not enough for a narrator to sound like a famous actor. Human talent makes the creative process dynamic and alive, whereas AI is trained on formulas. It regurgitates hollow mimicry perfectly, but without depth and understanding. 

When the AI voice replaces the human voice, what do we lose? AI users don’t need to find the right narrator, just the ‘right’ voice. But how well is the AI voice able to move with tonal shifts and emphasis? Already AI is able to write presentation scripts with notes as to where to pause or take a break. Any narrator can follow those notes, and the result will be technical narration. AI will be a competent enough ‘actor’ who can deliver the lines in the way stipulated by prompts, but it will not be able to bring fresh interpretations and originality, nor will it allow listeners to discover new things about the text the way a human narrator could. A poor performance by an actor may be described as wooden, flat, robotic. It is tempting to add ‘AI-like’ to this list of adjectives. Human emotion and artistry is what sets aside a good actor from a bad one. A competent actor can recite their lines and mimic emotions externally, but may not fully embody them. A good actor makes their lines come alive. The words become their own, because they externalise heartfelt emotions through their faces and their bodies, not just their voices. A proficient musician or dancer can be technically competent, but a great artist brings depth and layers which no machine can replicate. The discerning audience will be able to tell. 

Back to that hand-me-down pile of books. Although those were the days before audiobooks became widely available, I have a vague recollection of seeing audiobook cassette tapes in there. My mother could have outsourced her storytelling to an audiobook. If I had been born a few decades later, she could perhaps have lulled me to sleep with Goldilocks and the Three Bears narrated by a vocal replication of Morgan Freeman. But I would have missed my mother’s heritage and life experience in her accent, her imperfect delivery and every mispronounced word. And I would have been poorer for it. 

Rydwan Anwar is a Malay-Singaporean writer based in Newcastle Upon Tyne. After 20 years leading arts programmes at major cultural institutions as a theatre producer in Singapore, he now writes about theatre, culture, and creative practice. He holds an MA in Cultural & Creative Industries from King's College London.