← Files AI Film Pipeline MasterARCHIVED FILE

skills/ai-film-pipeline-master/references/phase-04-audio-narration/formats-audio.md

36.9 KB · Oct 7, 2026 · 00:35 UTC

↓ Download file

# Formats — audio

Entry point into the database. Each kit is a ranked short list; open the `cat-*.md` named in the row for the full card.

A format kit answers one question: **given this format's constraints, which techniques earn their place, and why this one rather than the obvious one.** The ranking is the argument.

**Shelf life.** The craft in these kits is durable. The platform numbers — median episode lengths, completion rates, listening contexts — are not. Review dated 2027-09.

---

<!-- GENERATED:toc — do not edit by hand. Generated in the source package -->

## Contents

- [Narrative Podcast](#narrative-podcast)
- [Interview Podcast](#interview-podcast)
- [Radio and Audio Drama](#radio-and-audio-drama)
- [Audiobook Narration](#audiobook-narration)
- [Voiceover Script](#voiceover-script)
- [Coverage](#coverage)
- [Sources](#sources)
- [On the visual kit](#on-the-visual-kit)

<!-- /GENERATED:toc -->

## Narrative Podcast

**Constraints:** median episode length 41 minutes, down from 48 in 2022; roughly a third of shows sit between 20 and 40 minutes, 16% run over an hour and 16% under ten. Narrative genres complete best — true crime and audio fiction run above 85% episode completion, against a long tail of chat shows that do not. Speech runs at 140 to 160 words per minute in conversational and hosted podcast delivery, so a 35-minute episode is about 5,000 words of script and tape combined, and one page of script is roughly ninety seconds of air. **That figure is for the talk, not for a documentary read** — this form's examples below include the audio documentary, and a scripted documentary or heritage narration is slower: take it from `pacing.wpm_documentary`, which owns that rate for the whole package, rather than from this line. Getting that wrong is not a small error at this length; over a 22-minute episode the two figures are about 880 words apart, which is three or four minutes of programme. 86% of listening is on a phone, 49% happens in a car, and 59% of listeners are doing something else while it plays. There are no faces: every speaker must be identifiable by voice, name or introduction, and a listener who loses track cannot glance back. They can only rewind, and they will not.

*(Speaking rates are owned by [`pacing-cadence-wpm.md`](pacing-cadence-wpm.md); this file cites them and does not restate them.)*

**Grammar:** A narrative podcast is built from two materials — tape (people speaking as themselves) and track (the host's written narration) — and the writing is mostly the decision about which one carries each beat. Tape is evidence, track is argument; a story told entirely in track is a lecture and a story told entirely in tape is a mess. The narration contract has to be stated early and honoured: whether the host is a reporter, a participant, a guide or an unreliable party decides what they are allowed to withhold later. Because a distracted listener drops the thread every few minutes, information is delivered on a need-to-know ladder and repeated in different words rather than assumed, which reads as clumsy on the page and correct in the ear. Every act break is a question, not a pause — a listener in a car decides to stay at every natural stopping point. Silence and ambient sound are punctuation, not atmosphere; a two-second gap does what a paragraph break does on a page.

**Palette:** 35 minutes · ~5,000 words · 140–160 wpm · 3–5 acts · 1 question per act · 4–6 recurring voices · a cold open under 90 seconds · scenes of 2–4 minutes

**Rules:** State the narration contract in the first two minutes — who is talking and why they know | Decide tape or track for every beat before you write the line | Introduce a voice the first three times it appears, then stop | Put a question at every act break, never a summary | Repeat the key fact in different words twice per episode; a driver missed it the first time | Write track to be spoken, not read — one idea per sentence, subject early | Use a two-second silence where a page would use a section break | Never introduce more than one new name per minute | End on the next question, and make it answerable next week

**Examples:** investigative limited series, the true-crime season, the historical narrative series, audio documentary, the reported-feature show

**Avoid:** The over-scripted host who narrates what the tape just said, the fastest way to teach a listener that track can be skipped. Four voices introduced in ninety seconds, after which nobody knows who is speaking. Building a season around a mystery you cannot resolve — audio audiences punish this more than readers do, because they gave you six hours. Music used to manufacture feeling the reporting did not earn. Ending an episode on a conclusion, which is the same as telling the listener they are free to go.

**The kit (ranked):**

1. **The Narration Contract** (Exposition) `expo.doc_narration_contract`
   Who the narrator is and what they are allowed to know. In audio this must be audible, not implied.
2. **Writing for the Ear** (Voice) `voice.write_for_the_ear`
   A sentence that reads well can be unsayable. Everything in this format is composed to be spoken once.
3. **Need-to-Know Ladder** (Exposition) `expo.need_to_know_ladder`
   A driver cannot re-read. Information arrives only at the moment it becomes necessary, never before.
4. **The Cold Open** (Openings) `open.cold_open`
   Ninety seconds of the most compelling tape you have, before any host speaks. It is the whole audition.
5. **The Micro-Reveal Ladder** (Information) `info.micro_reveal_ladder`
   Weekly release means every episode must pay something, or the audience learns that waiting is free.
6. **The Sequence Question** (Sequence) `seq.sequence_question`
   Each act is a sequence with its own question. Named out loud, it is what a distracted listener holds on to.
7. **Rhythm in Voice-Over and Audio** (Rhythm) `rhythm.vo_breath_line`
   Line length is breath length. Track written without breath points makes a good host sound bad.
8. **The Suspended Ending** (Tension) `tens.suspended_ending`
   The act break and the episode close. Pressure, not resolution, is what converts a listener into a subscriber.
9. **The Knowledge Ledger** (Information) `info.knowledge_ledger`
   Across eight hours you will lose track of who knows what and when the listener learned it. Keep the ledger.
10. **The Confiding Voice** (Voice) `voice.confiding_voice`
    The register the format runs on — one person, in your ear, telling you something. It fails the moment it is performed.
11. **The Specific Detail** (Emotion) `emo.specific_detail`
    With no image, the one exact detail is the only picture the listener gets. Spend words there and nowhere else.
12. **The Reporter's Stance** (POV) `pov.reportage_stance`
    The default stance, and the one that decides whether the host may speculate, feel, or take a side.
13. **The Dangling Cause** (Sequence) `seq.dangling_cause`
    A cause planted in one episode and paid two later is what makes a season feel designed rather than serialised.
14. **Dramatic Irony** (Information) `info.dramatic_irony`
    Letting the listener know before the people in the tape did is the narrative podcast's most reliable tension.
15. **Passage Marker** (Time) `time.passage_marker`
    Audio has no dateline. Time moves only when a voice says so, and it must say so every time.
16. **Start From a Question** (Theme) `theme.start_from_a_question`
    A season is an investigation. If the question was answerable in one episode, it was not a season.
17. **The Understating Voice** (Voice) `voice.understating_voice`
    In true crime especially, the restrained host is trusted and the dramatic one is not.
18. **Recurring Fragment** (Time) `time.recurring_fragment`
    The same piece of tape returned to across episodes, meaning more each time. Only audio can repeat it verbatim.
19. **The Fair-Play Contract** (Information) `info.fair_play_contract`
    If the season is a mystery, the listener must have had the pieces. Six hours buys them that right.
20. **The Callback** (Endings) `end.callback`
    The line from the cold open of episode one, returned in the finale. The season's cleanest available close.

---

## Interview Podcast

**Constraints:** 30 to 60 minutes is standard, with the 20-to-40 band the most listened; median across all podcasting is 41 minutes. The content is produced live by someone who is not you, so the only things you write in advance are the structure, the questions and the framing. There are no faces, so two similar voices are two problems — a guest and a host of the same register, age and accent will blur inside ten minutes for a listener who is driving. 59% of listeners are multitasking. Completion rates are materially lower than for narrative shows, so the middle of the episode is where the audience is lost. Editing is cheap, so the recorded conversation is raw material, not the product.

**Grammar:** An interview is a scene, and it obeys scene rules: someone wants something, someone resists, and a value turns. The host's want is a specific admission; the guest's want is usually to deliver their prepared version, and that conflict is the whole show. The real content lives in follow-ups, not in the prepared list — a prepared question produces a prepared answer, and the fourth question on one subject produces the episode. Questions are clustered by theme rather than run in a flat list, because a listener follows topics and not agendas. Silence is the host's strongest instrument and the hardest to use: a three-second pause after an answer gets more than any follow-up. The host's job on the page is to design an arc — a warm entry, a middle with real friction, a landing — and then abandon the page whenever the guest says something unplanned.

**Palette:** 45 minutes · 3–5 themed clusters · 5–8 prepared questions · 1 admission you want · 1 planned disagreement · 2 open loops promised early and closed late · a cold open pulled from minute 30

**Rules:** Name the one admission you want before recording; it decides everything else | Cluster questions by theme, never run a flat list | Ask the fourth question on a subject — the first three are the press release | Open the episode on the best 40 seconds of the conversation, wherever it occurred | Promise something early and pay it after the midpoint | Leave the pause; count to three before rescuing an answer | Never ask a question the guest has answered in another interview unless you are asking why they always answer it that way | Introduce the guest's relevance, not their résumé | Cut the whole first eight minutes if the conversation did not start until then

**Examples:** the long-form conversation show, the industry or craft interview, the celebrity promo interview, the two-host show with rotating guests

**Avoid:** The question list read in order, audible to the listener as a form being filled. Questions that are actually the host's essay with a question mark added. The unclosed loop: "we'll come back to that" and then not. Letting a guest deliver a fifteen-minute prepared story without a single interruption, which is not respect, it is abdication. Publishing in recording order — the conversation's best moment is almost never where it happened.

**The kit (ranked):**

1. **Who Asks the Questions** (Dialogue) `dial.question_control`
   The whole format in one card. Control of the question is control of the episode, and guests take it back constantly.
2. **The Scene Question** (Scene) `scene.scene_question`
   Every cluster is a scene with a question under it. A cluster with no question is filler and sounds like it.
3. **Give the Listener a Reason to Resist** (Exposition) `expo.resistant_listener`
   A host who accepts every answer produces a press release. Friction is what makes the information land.
4. **Silence as Reply** (Dialogue) `dial.silence_as_reply`
   Three seconds of nothing extracts more than any prepared follow-up. The hardest technique to actually execute.
5. **The Redirected Beat** (Beat) `beat.redirect`
   The tactical unit of interviewing: the guest goes somewhere, you move them, and the listener hears the move.
6. **Deflection** (Dialogue) `dial.deflection`
   Learn the shape of a deflection so you can hear it live. Every rehearsed guest has two or three.
7. **The Beat as Tactic** (Beat) `beat.tactic_beat`
   Each question is a tactic, not a topic. Naming the tactic is how you prepare without scripting.
8. **The Topic Change** (Dialogue) `dial.topic_change`
   Where a guest changes subject is more informative than what they say. Mark it and come back later.
9. **The Stated Promise** (Openings) `open.stated_promise`
   Say in the first minute what this episode will deliver. In a feed of forty options that is the entire pitch.
10. **Enter on the Answer** (Openings) `open.enter_on_the_answer`
    Cold-open the episode on a reply from minute thirty. The listener will stay for the question.
11. **The Micro-Reveal Ladder** (Information) `info.micro_reveal_ladder`
    The middle of an interview is where audiences leave. Something has to be paid every few minutes.
12. **The Change of Register** (Relationship) `rel.register_change`
    The moment a guest drops their public register is the moment of the episode. Recognise it and slow down.
13. **The Status Transaction** (Character) `char.status_transaction`
    Interviews are status negotiations end to end, and the listener hears status even when they miss the content.
14. **Information the Listener Already Has** (Dialogue) `dial.known_information`
    The trap of the well-researched host: asking things you know, in a way that shows you know them.
15. **The Half-Truth** (Dialogue) `dial.half_truth`
    What a rehearsed answer actually is. Naming the technique is how you find the unrehearsed one underneath.
16. **Refusing the Frame** (Dialogue) `dial.refusing_the_frame`
    Both directions: the guest rejecting your premise, and you declining theirs. The best exchanges are here.
17. **Talking to Avoid Talking** (Subtext) `sub.talk_to_avoid`
    A long answer is often an avoidance. Length is a signal, not a contribution.
18. **Writing for the Ear** (Voice) `voice.write_for_the_ear`
    Applies to your written questions: a question that needs a comma to parse will not survive being spoken.
19. **The Unanswered Line** (Dialogue) `dial.unanswered_line`
    Leaving a question hanging rather than filling the gap — audible, uncomfortable, and productive.
20. **The Callback** (Endings) `end.callback`
    Returning at minute forty to something said at minute four. It is also the cheapest proof that you listened.

---

## Radio and Audio Drama

**Constraints:** BBC slots run to 15, 30, 45, 60 and 90 minutes, and a page of dialogue script is roughly one minute of air. Four to five speaking characters per scene is the working maximum — beyond that, listeners who cannot see faces stop tracking who is talking. Everything exists through four materials: voice, sound effect, music, silence. There is no narrator unless you build one, and no way to show a room except by making someone move through it. Listeners are frequently doing something else, and often listening in a car, so a scene that requires close attention to a whispered line will be lost. Scripts are produced by a cast and an engineer, so every sound cue is a line item someone has to make.

**Grammar:** In audio, a character who is silent does not exist — presence is created by speech, breath or movement, and a listener will forget anyone who has not made a sound for two minutes. Names therefore do the work faces do, which means characters address each other by name far more often than in life, and the trick is making that sound like need rather than exposition. Acoustics are scene description: the reverb of a stairwell, the deadness of a car, the fact that rain means outside. A scene change is a sound change, not a slug line. Distinct voices must be designed at the script stage — different sentence lengths, different vocabulary, different rhythms — because casting cannot rescue two characters who write identically. Silence is the format's most powerful instrument and the one most likely to be cut by a nervous producer: a held two seconds of nothing is a close-up.

**Palette:** 45 minutes · ~45 pages · 4–5 voices per scene · 6–10 scenes · 1 acoustic per location · a silence of 2–3 seconds as a close-up · 1 recurring sound as a motif

**Rules:** Give every character a different sentence length and vocabulary before you cast | Use names early in a scene, in motivated lines, until the listener has the room | Change acoustics to change scenes; never rely on a music sting alone | Keep four voices to a scene and five as the ceiling | Write silence as a cue with a duration, or it will not survive the edit | Make every sound effect a story event, not an atmosphere | Let a character narrate their own action only when the action would be silent | Establish a recurring sound early so it can carry meaning late | Never write a reaction that is only a facial expression

**Examples:** BBC Radio 4 afternoon drama, the audio fiction podcast, the full-cast serialised audio drama, the adapted stage play for radio

**Avoid:** Characters describing what they are doing, the format's signature failure — "I'm opening the door now" is a sound cue that lost an argument with a writer. Crowd scenes, which turn to noise. Two similar voices in one scene, which is a casting problem created at the keyboard. Silence written as "(beat)" and therefore ignored. Sound effects used as wallpaper: if the rain is always there, the rain stops meaning anything, and the one moment it should mean something has been spent.

**The kit (ranked):**

1. **Writing for the Ear** (Voice) `voice.write_for_the_ear`
   The whole format's first principle. A line is not written until it has been said aloud by someone else.
2. **The Name Weapon** (Dialogue) `dial.name_weapon`
   In audio, names carry identification as well as aggression. This card is how to make necessary naming sound motivated.
3. **Silence as Content** (Subtext) `sub.silence_as_content`
   Audio's close-up. Two seconds of nothing does what a held shot does, and it is the first thing an editor cuts.
4. **The Beat of Silence** (Beat) `beat.silent_beat`
   The silence written as a beat with a job, so it survives the studio and the edit.
5. **The Voice of a Place or an Object** (Voice) `voice.voice_of_a_place`
   Audio can give a room, a city or a machine a voice with no justification. Almost no other format can.
6. **The Sound That Repeats** (Tension) `tens.repeating_sound`
   A recurring sound is the format's motif system, and its most efficient dread generator.
7. **Keeping the Group Scene Legible** (Relationship) `rel.ensemble_scene_legibility`
   Written for the page, essential here: four voices is the ceiling and this card is how to stay under it.
8. **Exposition Under Pressure** (Exposition) `expo.exposition_under_pressure`
   Radio must explain more than screen and has less patience for it. Pressure is the only cover available.
9. **The Interrupted Beat** (Beat) `beat.interrupted_beat`
   Interruption is how audio shows two people in the same space. Written carelessly it is just noise.
10. **Interruption** (Dialogue) `dial.interruption`
    Status readable in one exchange, with no visual cue needed — which is the only kind of cue this format has.
11. **The Threshold Scene** (Scene) `scene.threshold_scene`
    Doors, stairs, cars. A crossed threshold is the cleanest audible scene change in the medium.
12. **Restricted Narration** (Information) `info.restricted_narration`
    Audio drama is natively restricted: the listener knows only what is audible. Build on that instead of fighting it.
13. **The Bomb Under the Table** (Tension) `tens.bomb_under_table`
    Tell the listener what the characters cannot hear, and ordinary dialogue becomes unbearable. Cheap in audio.
14. **Waiting as a Weapon** (Tension) `tens.waiting_as_weapon`
    Time in audio is real time. A wait the listener endures is a wait they feel, unlike a cut to later.
15. **Smell Before Sight** (Tension) `tens.smell_before_sight`
    Non-visual sense detail is the format's native register, and it survives a listener who is only half-attending.
16. **The Clipped Line** (Dialogue) `dial.clipped_line`
    Short lines keep a scene legible when there are no faces to hold attention between long speeches.
17. **Talking to the Room** (Dialogue) `dial.talking_to_the_room`
    The legitimate version of talking to oneself, which audio needs more often than any other dramatic form.
18. **Passage Marker** (Time) `time.passage_marker`
    There is no caption reading THREE WEEKS LATER. Someone says it, or it did not happen.
19. **The Physical Signature** (Character) `char.physical_signature`
    Translated to sound: a cough, a limp, a habit of clearing the throat. Identification without dialogue.
20. **The Last Line** (Endings) `end.last_line`
    Audio drama ends on a voice, then nothing. The last line is the only image the listener leaves with.

---

## Audiobook Narration

**Constraints:** the planning rate is ACX's published average, which `pacing.wpm_narrated_fiction` in [`pacing-cadence-wpm.md`](pacing-cadence-wpm.md) owns and sources; a reflective non-fiction read runs slower, and the render, measured, is the length. An 80,000-word novel is therefore about 8.6 finished hours, and one hour of finished audio takes several hours to record and produce. The listener cannot skim, cannot glance back at a name, and is very often walking, driving or working. Many listen at 1.25x to 2x speed. One narrator usually performs every character, so a twelve-person dinner scene is a twelve-voice problem for one throat. Chapter breaks are the only navigation the listener has.

**Grammar:** Text written for the eye fails in the ear in specific, predictable ways, and writing an audio-aware book means anticipating them rather than fixing them in the studio. Dialogue attribution becomes essential where a page could drop it: a listener tracking four voices needs "he said" more often than a reader does, and needs it early in the line rather than at the end. Names have to be distinct in sound, not on the page — Marion and Miriam are the same character in audio. Anything visual in the typography is gone: italics for emphasis survive only if the sentence's stress already carries it, footnotes are an interruption, and a section break is silence of an unspecified length unless the writer gave it a marker. Long sentences are fine if they are breathable; long sentences with nested clauses are where narrators run out of air and listeners lose the subject. Repetition that looks lazy on a page reads as helpful in audio, because the listener has no way to check.

**Palette:** 155 wpm · 9,300 words per finished hour · 8–10 hours for a novel · chapters of 15–25 minutes · 2–4 voices per scene · attribution early in the line · names distinct at the first syllable

**Rules:** Give every character a name distinct in sound, not just in spelling | Attribute dialogue early in the line, and more often than the page needs | Keep sentences breathable — one nested clause, not three | Never rely on italics, footnotes or typography to carry meaning | Mark section breaks in words, because silence has no length in a listener's ear | Keep scenes to two to four speaking voices; one narrator cannot do six | Budget proper nouns hard: an unfamiliar name heard once is gone | Re-state who is speaking after any narrative interruption longer than a paragraph | Read your dialogue aloud in one voice; if you lose track, so will the listener

**Examples:** single-narrator commercial fiction, full-cast literary audio, author-read memoir, non-fiction and business audiobooks, audio-first originals

**Avoid:** Eye dialect, which is unreadable aloud and turns a narrator into a phonetics problem. Six named characters in one scene. Long unbroken interiority with no attribution, after which the listener does not know whose head they are in. Similar names, the most common note narrators give writers. Assuming the listener heard the first chapter at full attention — a large share of them were doing something else, and they will not go back.

**The kit (ranked):**

1. **Writing for the Ear** (Voice) `voice.write_for_the_ear`
   The governing card. Everything below is a specific consequence of it.
2. **Reading Aloud** (Rhythm) `rhythm.read_aloud_test`
   The manuscript's only valid audio test, and it must be done by someone who is not the author.
3. **Narrator Voice Against Character Voice** (Voice) `voice.narrator_vs_character`
   One performer holds both. If they are not distinct on the page, they cannot be distinct in the booth.
4. **The Action Beat** (Dialogue) `dial.action_beat`
   The attribution that does not sound like attribution — how audio keeps four voices straight without "he said" fatigue.
5. **The Proper Noun Budget** (Exposition) `expo.proper_noun_budget`
   A name heard once is lost. This is the card that says how many a listener can actually carry.
6. **Eye Dialect** (Voice) `voice.eye_dialect`
   Read as a prohibition. Spelled-out accent is a page technique that actively damages a recording.
7. **Dialect by Syntax** (Voice) `voice.dialect_by_syntax`
   The replacement: accent carried by word order and idiom, which a narrator can perform and a reader can follow.
8. **Rhythm in Voice-Over and Audio** (Rhythm) `rhythm.vo_breath_line`
   Breath is a hard physical limit. A sentence with no place to breathe is a sentence that will be re-recorded.
9. **Beat Density in Prose Dialogue** (Beat) `beat.prose_beat_density`
   How often to break dialogue with action. The audio answer is more often than the page answer.
10. **The POV Contract** (POV) `pov.contract`
    A listener cannot flip back to check whose chapter this is. The contract has to be audible at every handoff.
11. **The POV Handoff** (POV) `pov.handoff`
    Multi-POV books live or die here in audio. Name the new viewpoint in the first sentence, every time.
12. **The Chapter Opening** (Openings) `open.chapter_opening`
    Chapters are the only navigation. Each one re-establishes place, time and viewpoint inside two sentences.
13. **Sentence Length as Tempo** (Rhythm) `rhythm.length_as_tempo`
    At 155 words a minute for nine hours, uniform sentence length becomes physically soporific.
14. **Monotone Sentence Length** (Rhythm) `rhythm.monotone_length`
    The failure card for exactly that, and far more audible than it is visible.
15. **Passage Marker** (Time) `time.passage_marker`
    A white-space time skip does not exist in audio. Someone must say how much time passed.
16. **The Psychic Distance Ladder** (Voice) `voice.psychic_distance_ladder`
    A narrator performs distance with pace and warmth. Write the rung changes or they will be invented for you.
17. **The Full Stop as Tempo Control** (Rhythm) `rhythm.full_stop_tempo`
    Punctuation is the only direction the narrator receives. A full stop is an instruction, not a convention.
18. **Parallel Construction** (Rhythm) `rhythm.parallel_construction`
    Parallel structure is a memory aid for a listener who cannot look back at the list.
19. **The Triad** (Rhythm) `rhythm.triad`
    Three items is the most a listener reliably holds in a spoken sentence. Four is a fog.
20. **The Knowledge Ledger** (Information) `info.knowledge_ledger`
    Nine hours, often across two weeks. Track what the listener knows, not what the reader would have re-read.

---

## Voiceover Script

**Constraints:** 150 words per minute is the planning baseline, so a :60 is about 150 words, a :30 about 75 and a :15 about 37 — and those counts include the pauses, so the written script must come in under them. Registers move the number: a documentary or reassurance read is slower, and its rate is not repeated here — it is on `pacing.wpm_documentary`, which owns that figure for the whole package; corporate and e-learning narration sits around 140, and the commercial rate is not repeated here either — it is on `pacing.wpm_commercial_promo`, which owns that figure for the whole package. Whichever register you are in, the word count is a planning estimate and the real duration comes from the rendered audio file, measured. The script is read cold by a stranger who has never seen the product and will make the wrong choice on any line that can be read two ways. Where there is picture, the words carry only what the picture cannot; where there is not, the words carry everything. Timings are contractual — a :30 that runs :32 does not air.

**Grammar:** A voiceover script is a set of instructions to a mouth, not a piece of prose, and the difference shows in three places: the sentence is short enough to be said in one breath, the emphasis is built into the word order rather than left to italics, and the ambiguity is engineered out because the reader gets one pass. The single idea is the format's whole discipline — at 150 words there is exactly one thing you can say, and a script saying two says neither. Front-load: the promise of the piece belongs in the first sentence, because the listener's decision to keep listening is made before the second one. Every abstraction has to be converted into something a mouth can hit — a number, a name, a concrete noun — and the strongest word goes at the end of the line, where the voice naturally drops into emphasis. Read your own script against a stopwatch before you send it; the gap between a written 150 words and a spoken minute is where most scripts fail.

**Palette:** 150 wpm · 150 words per :60 · 75 per :30 · 37 per :15 · 1 idea · 1 promise · 1 call to action · sentences of 8–14 words · the strongest word last

**Rules:** Count the words against the clock before anything else | One idea per script; a second idea is a second script | Put the promise in the first sentence | Write sentences that can be said in one breath | Build the emphasis into word order, not into italics or bold | Strip every sentence that can be read two ways | Convert abstractions to numbers, names and objects | Put the strongest word at the end of the line | Read it aloud against a stopwatch and cut 10% — every script runs long | Write the register at the top of the page so the voice knows what it is doing

**Examples:** the :30 and :60 radio and TV commercial, the explainer and product video, corporate and e-learning narration, documentary and museum narration, trailer and promo copy

**Avoid:** Writing to the word count instead of to the clock, which is how a :30 becomes a :34. Lists of features where a listener needs one benefit. Sentences with subordinate clauses at the front, which force the reader to hold a subject they do not have yet. Trusting bold and italic to carry emphasis — the voice reads the words, not the formatting. Copy that describes the picture: if the visual shows the product, the words are wasted repeating it. A tonal drift halfway through, which happens whenever a writer stops hearing the read and starts writing sentences.

**The kit (ranked):**

1. **Rhythm in Voice-Over and Audio** (Rhythm) `rhythm.vo_breath_line`
   The format's core card. Line length is breath length and breath length is the timing.
2. **Writing for the Ear** (Voice) `voice.write_for_the_ear`
   One pass, no re-reading. Every ambiguity is a take that has to be done again.
3. **Reading Aloud** (Rhythm) `rhythm.read_aloud_test`
   With a stopwatch. This is the only way to know whether a 150-word script is a :60.
4. **Short-Form Front-Loading** (Exposition) `expo.short_form_frontload`
   The promise belongs in sentence one. The listener decides before sentence two arrives.
5. **Setting the Register** (Voice) `voice.register_setting`
   A documentary read and a 180 wpm retail read are different languages. Declare which one at the top of the page, and take the documentary rate from `pacing.wpm_documentary` rather than from memory.
6. **Concrete Nouns Over Abstract** (Rhythm) `rhythm.concrete_nouns`
   A mouth cannot hit an abstraction and an ear cannot hold one. Numbers, names and objects only.
7. **Verb Strength and the Passive** (Rhythm) `rhythm.verb_load`
   The passive costs words you do not have and removes the energy a read depends on.
8. **Word Order and Where the Emphasis Lands** (Rhythm) `rhythm.end_weight`
   The voice drops into emphasis at the end of a line. Put the word you are paying for there.
9. **Short, Short, Long** (Rhythm) `rhythm.short_short_long`
   The most reliable shape for a :30: two hits, then the line that carries the turn.
10. **The Stated Promise** (Openings) `open.stated_promise`
    Say what this is for, immediately. Every second spent establishing is a second of a thirty-second budget.
11. **The Proper Noun Budget** (Exposition) `expo.proper_noun_budget`
    A brand name, a product name and a URL is already three. A fourth name is not heard.
12. **The First-Second Hook** (Openings) `open.first_second_hook`
    Broadcast and pre-roll both give you one second before the listener disengages or skips.
13. **Anglo-Saxon Against Latinate** (Rhythm) `rhythm.saxon_vs_latinate`
    Short Germanic words are faster to say and faster to understand. At 150 words that is the whole budget.
14. **The Triad** (Rhythm) `rhythm.triad`
    Three beats is the most an ear holds in one sentence, and the most a voice can build to in a :30.
15. **The Full Stop as Tempo Control** (Rhythm) `rhythm.full_stop_tempo`
    Punctuation is direction. A full stop is a hard beat; a comma is not, whatever the writer intended.
16. **The Explainer Scaffold** (Exposition) `expo.explainer_scaffold`
    Problem, mechanism, proof, action. The shape almost every product and e-learning script is secretly using.
17. **The Sentence That Turns on Its Last Word** (Rhythm) `rhythm.last_word_turn`
    The written equivalent of a button. It gives the voice somewhere to land.
18. **Tonal Drift** (Voice) `voice.tonal_drift`
    The failure card for a script that starts warm and ends corporate — always at about the 40-second mark.
19. **The Funny Word Goes Last** (Comedy) `comedy.funny_word_last`
    For comic reads, the whole technique. Word order is the only timing control a writer has over a stranger's mouth.
20. **The Last Line** (Endings) `end.last_line`
    The call to action, the tag, the brand. One line, and it is the only one anyone is asked to remember.

---

## Coverage

Five kits: Narrative Podcast · Interview Podcast · Radio and Audio Drama · Audiobook Narration · Voiceover Script. Each ranks its techniques in order. Every id is live in `LOOKUP.md`; open the `cat-*.md` named there for the full card.

## Sources

Researched 2026-09-17.

- [Podcast Statistics and Trends for 2026 — Riverside](https://riverside.com/blog/podcast-statistics)
- [17 Essential Podcast Statistics You Need to Know in 2026 — The Social Shepherd](https://thesocialshepherd.com/blog/podcast-statistics)
- [Podcast Statistics 2026: Growth, Revenue and Trends — AffMaven](https://affmaven.com/podcast-statistics/)
- [The impact of driving versus undistracted listening on podcast comprehension — PLOS ONE](https://journals.plos.org/plosone/article/file?id=10.1371/journal.pone.0331299&type=printable)
- [Interview Podcast Format: Structure That Keeps Listeners Engaged — PodRewind](https://podrewind.com/blog/interview-podcast-format-structure)
- [How to Interview Someone for a Podcast — Lower Street](https://lowerstreet.co/how-to/interview-someone-for-podcast)
- [Podcast Structure: How to Create One — Riverside](https://riverside.com/blog/podcast-structure)
- [Producing and Recording Your Audiobook — ACX Help](https://help.acx.com/s/article/producing-and-recording-your-audiobook)
- [Audiobook Length Calculator: Word Count to Listening Hours — ReadMinutes](https://readminutes.com/audiobook-length-calculator)
- [Simple Math About Audiobook Rates — Karen Commins](https://karencommins.com/2011/06/some_simple_math_about_audiobo.html)
- [How to Write a Killer Audio Drama Script — Tony Sarrecchia](https://www.tonysarrecchia.com/p/how-to-write-a-killer-audio-drama)
- [BBC Radio Drama Format Guide: Professional Script Writing 2026 — EpicScribe](https://epicscribe.io/blog/bbc-radio-drama-format-guide.html)
- [How to Write a Radio Play — EpicScribe](https://epicscribe.io/how-to/write-radio-play.html)
- [Words to Time Conversion Calculator — Voices.com](https://www.voices.com/tools/words_to_time_conversion)
- [Voiceover Script Length: Word Count to Duration — GoTeleprompter](https://goteleprompter.com/blog/voiceover-script-word-count-calculator/)

---

## On the visual kit

**These five formats have no visual kit, and that is the finished answer rather than a gap.**
Every other kit in the package carries one; an audio piece has no frame to compose, no light to
motivate and no cut to place, so a ranked list of visual cards here would be a list nobody could
use.

**Two exceptions, and they are real.** A piece with no picture still usually ships **one still** -
cover art, an episode card, a clip thumbnail - and that still is Phase 7 work governed by the
poster and thumbnail craft, opened for that one image only. And a podcast recorded on camera for
clips is not an audio format at this point: it is an interview format with an audio deliverable,
and it takes the vertical and long-form kits.

**What replaces the visual kit here is the sound design**, which these kits already rank, and
the Phase 2 Pipeline C rule that every sound is written down including the deliberate silences.

SHA-256: 564aa67acaed86e2f3f9c2b875d386bcafb15e65a788756fab6cfb740efb6598