Making an audio story from a text you wrote
The first decision shapes everything else: a human reading is better once, a generated voice is better every time you change something afterwards.

Short answer
Making an audio story starts with one decision: whether a person reads it or a synthesised voice generates it. A human reading has more warmth but changing one sentence later means matching the microphone, room and tone; a generated voice regenerates from the text, so audio and script never drift apart. Then edit the text for the ear — short sentences, spoken attributions, nothing visual.
On this page
You wrote something — a bedtime story, a company history, a set of instructions — and you want it as audio. Not a podcast with music and editing, just the words, well read, in a file people can play.
Making an audio story is a smaller job than it appears, and the decision that shapes everything is the first one: whether a person reads it or a synthesised voice does. That choice determines the time, the cost, and what you can change afterwards, and it is worth making deliberately rather than by default.
Should you read it yourself or generate it?
For an audio story the trade is not simply quality against convenience.
| You read it | Generated voice | |
|---|---|---|
| Warmth and character | Better | Flatter |
| Time for 2,000 words | An hour with retakes | Minutes |
| Fixing one sentence | Re-record, match tone | Change the text |
| Consistency across chapters | Hard | Perfect |
| Cost | Free | Free to modest |
The row that decides it for most people is the third. A recorded narration is a performance: change one sentence six months later and you need the same microphone, the same room and roughly the same voice, or the edit is audible. With a generated voice, changing a sentence means changing a sentence — a tool like SpeakFile regenerates from the text, so the audio and the script cannot drift apart.
A human reading is better the first time. A generated voice is better every time after that.
For an audio story you will publish once and never touch, read it yourself. For anything that will be revised, translated, or produced as a series, generate it.
How do you write for the ear?
Text written to be read silently does not work aloud, and four changes fix most of it.
- Shorten the sentences. A sentence with three subordinate clauses is fine on a page and unfollowable in audio, because the listener cannot go back.
- Say who is speaking. "He said" and "she asked" are redundant on a page where quotation marks do the work, and necessary in audio where they do not.
- Replace anything visual. "As shown above", "the following list", "see the diagram" — none of these exist for a listener.
- Read it aloud before recording it. Every awkward phrase, every place you run out of breath, every accidental rhyme reveals itself immediately. This is the single most useful step and it costs one pass.
The fourth point applies equally to generated audio: if you stumble reading it, the synthesiser will produce something equally awkward, because the problem is in the sentence.
Recording, if you are reading it
Recording an audio story well comes down to five things, in order of how much difference each makes.
- Get close to the microphone. Twenty centimetres, slightly off-axis so plosives do not hit it directly. This matters more than the microphone itself.
- Choose a soft room. Curtains, carpet, a wardrobe of clothes. Hard rooms produce reverb that cannot be removed afterwards.
- Turn off everything that hums. Air conditioning, fans, the fridge. Continuous noise is what listeners notice.
- Record in sections, one scene or one chapter at a time. A mistake then costs a section rather than the whole take.
- Leave a few seconds of silence at the start. It gives you a sample of the room's noise, which is what noise reduction needs to work from.
And one habit that saves more time than any of them: when you fumble a line, pause, then say it again from the start of the sentence. Do not stop and restart the recording. The edit is trivial and the flow is preserved.
Assembling and exporting
The technical half of an audio story is smaller than people expect.
Format. MP3 for compatibility, or AAC. Both are fine for speech, and speech does not need a high bitrate — a spoken-word file at a modest setting is indistinguishable from a large one.
Levels. Consistent volume across chapters matters more than absolute loudness. A listener adjusting the volume between chapters is the most common complaint about home-made audio.
Silence. A short gap at the start and end of each file, and a beat between sections. Audio that begins mid-word feels broken.
Chapters. For anything long, separate files with clear names beat one long file, because a listener who loses their place in a fifty-minute file has lost it entirely.
Where the story is for children specifically, shorter files matter more still — an app like Tales Home, which pairs each short story with a complete narration, reflects how the format is actually used: one story, one sitting, played again tomorrow.
Adding sound, or not
The instinct is to add music. Usually it is worth resisting.
Music under speech competes with it, and at a level low enough not to compete it is barely audible — which raises the question of why it is there.
Rights matter. A commercial track in something you publish invites a takedown, and platforms detect it automatically. Use genuinely licence-free material or nothing.
Effects work in specific places. A door, a storm, a bell at a scene change. Sparingly, and always at a level well below the voice.
Silence is a tool. A pause before a revelation does more than any effect, and it costs nothing.
If you do use music, put it at the beginning and end rather than underneath. A short opening and closing gives the piece shape without fighting the words.
Where do people actually listen?
Worth knowing before you finish an audio story, because it changes decisions you have already made.
On a phone speaker, in a room with other noise. Not headphones in silence. This is why consistent levels matter more than subtlety, and why a quiet passage recorded beautifully may simply vanish.
In the car. Long stretches, no ability to look at anything, and road noise underneath. Anything requiring visual reference is lost, and so is any dynamic range.
At bedtime, quietly. For children's stories specifically this is the dominant case, and it argues for even levels and a calm ending rather than a dramatic one.
In fragments. Someone listens for six minutes, stops, and comes back tomorrow. Clear chapter divisions and file names are what make that possible, and a single long file makes it impossible.
The practical consequence of all four is the same: compress your dynamic range more than feels right. A recording with beautiful quiet moments and loud peaks is a recording where half the words are inaudible on a phone in a kitchen. Even levels sound flat in headphones and correct everywhere people actually listen.
One more: check the very first sentence carefully. It is the part most often played while someone is still adjusting the volume, so it needs to be clear rather than clever — and if the file begins with a title, keep it short enough that nobody skips past the opening line to get to the story.
A workable sequence
For a two-thousand-word audio story, about ninety minutes end to end.
- Edit the text for the ear, reading it aloud once as you go.
- Decide voice or narration, using the revision question above.
- Record or generate in sections.
- Assemble with consistent levels and a beat between sections.
- Listen to the whole thing once, on the device people will use — phone speaker, not headphones.
- Export, name it properly, and keep the source text beside the audio.
Step five catches almost everything: a phone speaker in a normal room is an unforgiving test, and it is where most people will actually hear it.
More creator tools in creator tools, the text side in text tools, and general picks in online tools. The WebVTT specification covers captions if the audio is going alongside video.
The short version
Decide first whether a person reads it or a voice generates it, and decide on revisions: a human reading is better once, a generated voice is better every time you change something afterwards.
Then edit the text for the ear — short sentences, spoken attributions, nothing visual — read it aloud before recording, get close to the microphone in a soft room, and keep the levels consistent across sections. Resist music under the words, and listen to the finished file on a phone speaker, because that is where it will actually be heard.
Frequently asked questions
- Should I read it myself or use a generated voice?
- Read it yourself for something you will publish once and never touch. Use a generated voice for anything you will revise, translate or produce as a series, because changing a sentence then means changing a sentence rather than matching a recording.
- What is the single best thing I can do before recording?
- Read the text aloud once. Every awkward phrase, every place you run out of breath and every accidental rhyme reveals itself immediately — and it applies to generated audio too, since the synthesiser inherits the same problem.
- What should I do when I fumble a line?
- Pause, then say it again from the start of the sentence without stopping the recording. The edit is trivial and the flow is preserved, whereas restarting costs the take.
- Should I add background music?
- Usually not. Music low enough not to compete with speech is barely audible, and commercial tracks invite automatic takedowns. Put a short piece at the start and end rather than underneath the words.
Sources
- SpeakFile: Text to Voice — Tecno Blocks
- Tales Home: Audio Stories — Tecno Blocks
- WebVTT: The Web Video Text Tracks Format — W3C
Loopara
Practical guides, free tools, workflows, and resources for productivity, files, images, video, text, creators, and everyday digital tasks.
About the publication