← MoknahCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Moknah
Snapshot Sep 30, 2026 · 22:52 UTC · version 1.0.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Produce natural dialogue and pacing in Moknah - detect speakers, cast a voice per character, merge consecutive same-speaker lines, and place [break_X] pauses at speaker changes, sentence ends and scene beats. Use when the text contains dialogue, or when rendered audio sounds choppy, rushed, or abruptly cut between voices.",
"included_files": [],
"name": "moknah-dialogue-and-pacing",
"skill_md_contents": "---\nname: moknah-dialogue-and-pacing\ndescription: Produce natural dialogue and pacing in Moknah - detect speakers, cast a voice per character, merge consecutive same-speaker lines, and place [break_X] pauses at speaker changes, sentence ends and scene beats. Use when the text contains dialogue, or when rendered audio sounds choppy, rushed, or abruptly cut between voices.\n---\n\n# Dialogue and pacing in Moknah\n\nThe single biggest quality difference between amateur and professional Moknah\noutput is **line structure and pause placement**. Voice choice matters less than\nthis.\n\n## Why line structure decides quality\n\nEach line is rendered as its own TTS request. Every render starts prosody from\nscratch: fresh intonation contour, fresh energy, fresh breath. So:\n\n- Splitting one speaker's continuous speech across several lines produces audible\n restarts - the pitch resets mid-thought and the delivery sounds stitched.\n- Two adjacent lines with *different* voices butt straight up against each other\n with no gap, so the switch sounds like an edit error rather than a reply.\n\nBoth are fixed at the text/segmentation layer, before you spend a single credit.\n\n## Rule 1 - Merge consecutive lines from the same speaker\n\nIf two or more consecutive lines belong to the same speaker, **combine them into\none line.** Let the voice run continuously through the whole utterance.\n\nDo this at cleaning time, before `create_project` / `add_line`, or fix it after\nwith `update_lines`.\n\n```\nBAD (3 lines, same speaker - 3 prosody resets)\n line 1: \"ما زلت أذكر ذلك اليوم.\"\n line 2: \"كانت السماء صافية.\"\n line 3: \"ولم يكن أحد يتوقع ما سيحدث.\"\n\nGOOD (1 line, one continuous delivery)\n line 1: \"ما زلت أذكر ذلك اليوم. كانت السماء صافية. ولم يكن أحد يتوقع ما سيحدث.\"\n```\n\nStop merging when you hit any of these ceilings:\n\n| Limit | Value |\n|---|---|\n| Characters per request, emotions ON | 3,000 |\n| Characters per request, emotions OFF | 10,000 |\n| Lines per chapter | 200 |\n| Breaks per line | 20 |\n| Total pause per line | 30 s |\n\nNever merge across a speaker change. Never merge across a scene break.\n\n## Rule 2 - Always break on a speaker change\n\nWhen the voice changes, insert a pause at the **end of the outgoing speaker's\nline**. Without it the two voices collide and the listener misses the handoff.\n\n**Apply this by default, without being asked.** It is standard house behaviour,\nnot an option to raise with the user. Every speaker change gets a pause unless\nthe user has asked otherwise.\n\n```\nline 1 (Voice A): \"هل انتهيت من العمل؟ [break_0.5]\"\nline 2 (Voice B): \"لم أبدأ بعد.\"\n```\n\nUse **0.4-0.7 s**. Fast exchanges sit at the low end; a considered reply or a\nchange of tone sits at the high end.\n\n### The one exception - deliberate interruption\n\nDrop the pause only when the user wants the next voice to **cut the current\nspeaker off**: an argument, an interjection, someone talking over someone else,\noverlapping urgency. Then the two lines butt straight together with no gap and\nthe abruptness is the point.\n\nNever remove the pause on your own judgement - only when the user asks for that\neffect.\n\n## Rule 3 - Pause at sentence ends and beats\n\nPunctuation gives a small natural pause on its own. Add an explicit `[break_X]`\nonly where you want the listener to *feel* the gap - the end of a thought, a\nparagraph, a scene, a chapter opening.\n\nRecommended starting values, then tune by ear:\n\n| Position | Duration |\n|---|---|\n| Between sentences, same speaker (only if the text feels rushed) | 0.3 - 0.5 s |\n| End of paragraph / thought | 0.6 - 0.8 s |\n| Speaker change in dialogue | 0.4 - 0.7 s |\n| Scene or section break | 1.0 - 1.5 s |\n| After a chapter title, before the body | 1.0 - 1.5 s |\n| Dramatic beat (suspense, revelation) | 1.5 - 2.5 s |\n\nDo not put a break after every sentence. Over-paced narration sounds like a\ndictation exercise. Breaks are punctuation for the ear - use them where meaning\nturns.\n\nFor a standalone pause between blocks, put `[break_X]` on its own line.\n\n## Detecting speakers\n\nScan for these patterns before segmenting:\n\n- Arabic attribution verbs: `قال فلان:` / `قالت` / `أجاب` / `همس` / `صاح`\n- Quotation marks: `\"...\"` `«...»` `„...\"`\n- The dialogue dash at line start: `- ...`\n- Script format: `NAME: line`\n\nBuild a **speakers map** (character -> voice) and keep it consistent for the whole\nbook. Persist it so later chapters cast identically.\n\n## Casting\n\n- One dedicated voice per recurring character, plus a distinct **narrator** voice.\n- Adjacent speakers must be clearly distinguishable - do not cast two similar\n voices for two characters who talk to each other.\n- Attribution tags (`قال أحمد`) belong to the **narrator**, not the character,\n unless the user asks otherwise.\n- Get casting approved before rendering (this is a gate in every production mode).\n- Preview with `play_voice_sample` before the user commits.\n\n### Emotions in dialogue\n\nTurn `emotions_mode` **on** for multi-voice dialogue - utterances are short and\nthe voice keeps changing, so the variation between generations reads as\nperformance rather than inconsistency. It is a clear quality gain for dialects in\nparticular.\n\nTurn it **off** for a single-narrator production, even a dialect one. Over\nhundreds of consecutive lines in the same voice that variation becomes an audible\nstability problem. See `moknah-voice-settings`.\n\n## Markup syntax - get this right or tags get read aloud\n\nThe tag syntax depends on `emotions_mode`:\n\n- `emotions_mode: true` -> square brackets: `[happy]`, `[whispering]`, `[sighs]`\n- `emotions_mode: false` -> SSML-style angle brackets `<...>`, used sparingly\n\n**Never mix the two.** With emotions off, `[...]` tags are spoken aloud as text.\n\n`[break_X]` works in both modes. X is seconds, 0.1-10.\n\nAvailable emotion tags (emotions_mode true): `[laughing] [happy] [excited]\n[angry] [shouting] [sad] [crying] [surprised] [afraid] [calm] [serious]\n[sarcastic] [curious] [whispering] [sighs] [exhales] [wheezing] [snorts]\n[mischievously] [questioning] [confident] [swallows] [gulps]`\n\nEffects: `[applause] [clapping] [gunshot] [explosion]`\n\n**Emotion and effect tags are experimental and not fully reliable.** Tell the user\nthat up front, and validate on one sample line before applying them book-wide.\n\nA line containing only tags and no real words is skipped. Tags and breaks are not\nbilled beyond their character count.\n\n## Checklist before rendering dialogue\n\n1. Consecutive same-speaker lines merged.\n2. A `[break_]` at every speaker change.\n3. Beat pauses at paragraph, scene and chapter boundaries only.\n4. Speakers map complete; every speaking character cast to a distinct voice.\n5. Tag syntax matches `emotions_mode`.\n6. One sample line rendered per voice and approved before the full run.\n"
}SHA-256 of public snapshot: 0f74dfdb8cf905133f2c1cdf0c62618430879f3588e0936c46817c7189462071