← MoknahCONTENT HISTORY

Update to Moknah

Snapshot Sep 30, 2026 · 22:52 UTC · version 1.0.1

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "moknah-voice-settings",
  "description": "What every Moknah voice parameter does - voice_id, temperature, similarity, speed, expressiveness, emotions_mode, prerecording - how the server/user/project/line cascade resolves, and which settings to use for Standard Arabic, dialects, non-Arabic and poetry. Use when choosing or tuning voice settings, or when rendered audio sounds wrong.",
  "included_files": [],
  "skill_md_contents": "---\nname: moknah-voice-settings\ndescription: What every Moknah voice parameter does - voice_id, temperature, similarity, speed, expressiveness, emotions_mode, prerecording - how the server/user/project/line cascade resolves, and which settings to use for Standard Arabic, dialects, non-Arabic and poetry. Use when choosing or tuning voice settings, or when rendered audio sounds wrong.\n---\n\n# Moknah voice settings\n\n## The cascade\n\nSettings resolve from the most general to the most specific:\n\n```\nserver default  <  user default  <  project  <  line\n```\n\nA `null` at any level means **inherit from the level above**. Set the project\nvoice once; override individual lines only by exception. Do not stamp the same\nvalue onto 200 lines - that is what the project level is for.\n\n- `set_project_voice` - the book's default voice and settings\n- `set_line_voice_settings` - one line's override\n- `update_lines` - batch overrides, up to 200 lines in a single call\n\n## The parameters\n\n### `voice_id`\nWhich voice speaks. Must come from `list_voices` - never invent an id, and only\nvoices the user actually has access to will work. Premium voices add a\npercentage fee on top of the per-character rate, so the same text costs more.\n\n### `prerecording` - the text pre-processing pipeline\nThis is the highest-impact setting for Arabic. It selects how the text is\nprepared *before* synthesis.\n\n| Value | Mode | Use for |\n|---|---|---|\n| `1` | Normal | Dialects, all non-Arabic languages |\n| `2` | Arabic Tashkeel | Standard Arabic (fusha) only - applies contextual diacritics |\n| `3` | AI-poet | Poetry and metered verse |\n\n`prerecording: 2` markedly improves fusha pronunciation because Arabic script is\nambiguous without diacritics (`كُتب` vs `كَتب`). **Never use it on dialects or\nnon-Arabic text** - it will impose fusha diacritics on words that do not take\nthem and the output gets worse, not better. It is billed at the advanced rate.\n\n### `emotions_mode` - the expressive engine\n`true` enables the expressive engine and square-bracket emotion tags. `false`\nkeeps the standard engine.\n\nIt sounds **markedly better for dialects and for dialogue**. But every generation\ncomes out slightly different, so it carries a real stability cost.\n\n**The deciding question is how many voices the production uses:**\n\n- **Dialogue / multi-voice -> emotions ON.** Utterances are short and the voice\n  keeps changing, so variation between lines reads as natural performance.\n- **Single narrator (\"one man show\") -> emotions OFF**, even for dialect. Across\n  hundreds of consecutive lines in one voice, that same variation turns into an\n  audible inconsistency the listener notices immediately. Stability wins.\n\nThis supersedes any blanket \"dialects always use emotions\" rule: for dialect\ncontent the mode still helps, but only when more than one voice is speaking.\n\nOther consequences:\n- Tag syntax flips: `true` -> `[happy]`; `false` -> SSML `<...>`. Mixing them\n  makes the tags get read aloud as literal text.\n- The per-request character limit drops: **3,000 with emotions on, 10,000 off.**\n\n### `speed`\nDelivery rate. Range `0.7` - `1.2`; `1.0` is natural. Leave at 1.0 unless the\nuser asks. Small moves (0.95, 1.05) are usually enough - anything past ~1.15\nstarts to sound clipped.\n\n### `similarity`\nHow closely the output tracks the reference voice's timbre and identity.\nKeep it **high** as the default - that is what makes a voice sound like itself\nacross a whole book. Lower it only when the voice is a clone made from noisy\nsource audio, where a high value faithfully reproduces the noise and artefacts\nalong with the voice.\n\n### `temperature`\nVariability between renders. Lower = more consistent, steadier, more predictable\ndelivery. Higher = more variation and life, at the cost of consistency.\n\n**Hard ceiling: never go above 0.75.** Past that the model begins to\n**hallucinate** - inventing words, dropping text, mangling phrases. This is not a\n\"slightly less polished\" threshold; it is the point where the output stops being\ntrustworthy and has to be checked word by word.\n\n- Fusha narration: leave at default. Consistency across hundreds of lines matters\n  more than per-line colour.\n- Dialect: nudge slightly up - dialect delivery sounds flat when too tightly\n  constrained - but stay well below 0.75.\n\n### `expressiveness`\nHow much prosodic style is applied - emphasis, dynamic range, dramatic contour.\n\n**Keep it at 0. Never raise it on your own initiative.** This parameter is\nextremely sensitive - small increases cause hallucination, not merely a livelier\nreading.\n\nRaise it only when the user explicitly asks for it, and when they do:\n\n1. **Warn them it risks hallucinated audio.**\n2. Render a single sample line and have them check it before applying it anywhere\n   else.\n\nIf output starts inventing or dropping words, this is the first setting to put\nback to 0.\n\n## Settings by scenario\n\n`expressiveness` is `0` in every row - it is never raised except on explicit\nuser request. `temperature` never exceeds `0.75` in any row.\n\n| Scenario | prerecording | emotions_mode | temperature | expressiveness | speed |\n|---|---|---|---|---|---|\n| Fusha, single narrator | 2 (Tashkeel) | false | default | 0 | 1.0 |\n| Fusha, single-paragraph input | 2 | true | default | 0 | 1.0 |\n| Dialect, single narrator | 1 | **false** (stability) | slightly up, < 0.75 | 0 | 1.0 |\n| Dialect, dialogue / multi-voice | 1 | **true** | slightly up, < 0.75 | 0 | 1.0 |\n| Non-Arabic, single narrator | 1 | false | default | 0 | 1.0 |\n| Non-Arabic, dialogue / multi-voice | 1 | **true** | default | 0 | 1.0 |\n| Poetry / verse | 3 (AI-poet) | false | default | 0 | 1.0 |\n\nDetect the language from the text itself and apply the matching row - do not ask\nthe user to specify it if the text makes it obvious.\n\n## Voice choice by content\n\n- **Fusha book** -> a fusha voice. A dialect voice reading fusha sounds livelier\n  and is acceptable if the user prefers it.\n- **Dialect book** -> the voice *must* match the dialect. Prime it by opening\n  with a strongly dialectal word so the voice settles into the register.\n- **Dialogue** -> a distinct voice per character plus a separate narrator voice.\n  See the `moknah-dialogue-and-pacing` skill.\n\n## Diagnosing bad output\n\n| Symptom | First thing to change |\n|---|---|\n| **Invented words, dropped or mangled text (hallucination)** | `temperature` is above 0.75, or `expressiveness` is above 0. Put both back down - this is the cause almost every time |\n| Arabic words mispronounced / wrong vowels | fusha? set `prerecording: 2`. Dialect? make sure it is `1` |\n| Tags read aloud as words | `emotions_mode` is false but `[...]` tags are in the text |\n| Delivery inconsistent between lines in one voice | `emotions_mode` is on for a single-narrator book - turn it off; then lower `temperature` |\n| Flat, lifeless dialect narration | raise `temperature` slightly (stay < 0.75); turn on `emotions_mode` **only** if multi-voice |\n| Voice does not sound like the reference | raise `similarity` |\n| Clone reproduces hiss/artefacts | lower `similarity` |\n| Choppy, restarting delivery mid-thought | not a settings problem - merge the lines (dialogue skill) |\n\nChange **one** setting at a time and re-render a single line with\n`generate_lines`. Do not re-render a chapter to test a hypothesis.\n\n## Cost note\n\nVoice settings are free to change - only rendering costs credits. Iterate on\nsettings against one short sample line until the user approves, then run the\nbook. Fixing a setting after a full render costs a full re-render.\n"
}

SHA-256: efffb5588de4e67a954f4cc070ff92dead603dd6ae9e5511abaaabd7e7590401