Moknah
Moknah Group v1.0.1
Moknah turns text and documents into natural-sounding speech and complete audiobooks, with especially strong support for Arabic alongside many other languages. Paste text or upload a PDF, Word, or EPUB, choose a voice, and generate audio you can preview and download. For Arabic, it automatically adds diacritics (tashkeel) and normalizes the text so words are pronounced correctly, with optional emotion and poetry modes. Break longer works into chapters, set a different voice per section, and review quality before you export. A Moknah account is required to generate audio.
Language: English · Automatically detected from descriptions.
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package author
- Moknah Group
Package observed Sep 30, 2026.
Files & skills
File archives
Skill instructions
moknah-audiobook-production6.47 KB
---
name: moknah-audiobook-production
description: End-to-end workflow for turning a document (PDF, EPUB, Word, TXT) into a finished audiobook with Moknah - production modes, intake, cost gates, sampling, rendering and delivery. Use when the user wants a book or long document narrated. Not needed for one-off text-to-speech of a short passage.
---
# Moknah audiobook production
Take a document and deliver a finished audiobook. The governing rule: **no credit
is ever spent without an estimate shown and an explicit approval given.**
## Mental model
```
Project -> Chapters -> Lines
```
One line = one spoken unit = one TTS request. A line is "converted" once it has
rendered audio. `total_chars` on a line means **billed** characters (0 until it
renders) - it is not the length of the text.
Projects persist. If a session is interrupted, call `get_project` and continue -
never restart a book from scratch.
## Step 1 - Always ask for a production mode first
- **Full Control** - confirm every decision and every stage.
- **Guided** *(recommended default)* - ask only the 7 key decisions below, decide
everything else from these skills.
- **Express** - decide everything from the book itself. Only two touchpoints
remain: spend confirmation and sample approval.
The 7 Guided decisions:
1. Cleaning - you do it, the user does it, or skip (warn that skipping hurts quality)
2. AI-Enhanced normalization (tashkeel, costs 2x) or Basic
3. Translate to another language?
4. Segmentation style (sentences / newline / custom)
5. Read chapter titles aloud? Include the title in the chapter text?
6. One voice for the whole book, or a voice per character?
7. Emotions on or off (warn that tone may vary between lines)
## Stage 0 - Intake
Accepts pdf, epub, doc, docx, txt, xlsx. PDFs must be under 50 MB - split or
supply a lighter copy if larger.
Determine:
- **PDF type** - has a text layer (use directly) or scanned (OCR; set
`ocr_source = true` and review far more strictly)
- **Language** - fusha / dialect / non-Arabic / mixed
- **Structure** - TOC, front matter, appendices
**Gate 0:** page count, the proposed body range (skip cover, TOC and appendices
with `start_page` / `end_page`), any flags, and `estimate_book_cost`.
## Stage 1 - Extraction
Prefer extract -> clean -> `create_project(text=...)` over handing the raw file
over, because it lets you clean before anything is billed. For PDFs use
`create_project` with `start_page` / `end_page`.
## Stage 2 - Cleaning (free)
Follow the **`moknah-text-preparation`** skill. Produce a cleaning report.
**Gate 2:** the report - change counts, before/after samples, proper-names
dictionary, chapter list.
## Stage 3 - Language review
Review before any spend. Post-generation fixes cost double, because you pay to
render the mistake and again to render the correction. Be stricter when
`ocr_source` is set.
**Gate 3:** itemized in Full Control, a summary in Guided, silent in Express.
## Stage 4 - Create the project
- Name = the exact book title
- `chapter_style = Heading 1`, `include_chapter_title` per gate
- `line_split_mode` per gate
- **Normalization by language:**
- fusha -> AI-Enhanced (2x credits, best result). If the user refuses the cost,
use Basic and manually add tashkeel only to genuinely ambiguous words
- dialect -> Basic, always
- non-Arabic -> Basic
- mixed -> split by language, or follow the dominant one
- Translation (100+ languages) - the user reviews the translation *before* any
audio is generated
After creation, verify the chapters match your map. If headings were missed, fix
the source document and re-import rather than patching chapter by chapter.
## Stage 5 - Voices and settings
Follow **`moknah-voice-settings`** for parameters, and
**`moknah-dialogue-and-pacing`** for casting, line merging and pause placement.
**Gate 5a - casting approval, in every mode.** Express proposes and confirms once.
Use `play_voice_sample` so the user *hears* a voice before choosing it.
## Stage 6 - Sample, then generate
Render **one sample line per content type** - narration, dialogue, verse or
poetry, a line with converted numbers, a line with a foreign name. Pick
non-adjacent lines. With AI-Enhanced normalization, sample from the start, middle
and end to check tashkeel consistency.
**Gate 6 - the big spend gate, never skipped in any mode.** The user listens and
approves. On rejection, iterate in this order: settings -> voice -> text. Cap it
at 3 loops, then recommend human help rather than burning credits.
Then generate, chapter by chapter (better fault isolation) or whole book. If one
line comes out wrong, re-render only that line with `generate_lines` - never the
whole chapter. Log every re-render.
## Stage 7 - Review and delivery
- Spot-check the start, middle and end of each chapter (`get_chapter_audio`)
- Deliver per-chapter files (default) or one merged file via
`merge_chapters_audio`
- Matching subtitles via `merge_chapters_subtitle` (`srt` or `vtt`)
- Both merges are **free**; omit `chapter_ids` to cover the whole book in order
- Naming: `{NN} - {chapter}.mp3`, zero-padded
- Final report: durations, credits spent, lines re-rendered
## Gate matrix
| Gate | Full Control | Guided | Express |
|---|---|---|---|
| 0 intake / page range | ask | ask | auto |
| 2 cleaning report | ask | ask if AI cleaned | on demand |
| 3 corrections | itemized | summary | auto |
| 4 project options | ask all | 7 questions | auto |
| 5a voice casting | ask | ask | propose + confirm once |
| 6 sample approval | ask | ask | **ask - never skip** |
| 7 merge / format | ask | ask | auto |
| **any credit spend** | ask | ask | **ask - never skip** |
## Money
**Free:** creating, editing text, reordering, voice settings, QA status, merging
audio, merging subtitles, and every `estimate_*` call.
**Billed:** generation (TTS), PDF OCR, AI-normalization, translation,
transcription, AI-QA.
Always call the matching `estimate_*` tool and show the number before a billable
action. If the balance is short, say so plainly, and offer to reduce scope -
rendering a single chapter now is a legitimate third option.
## Jobs
Long work returns a `job_ref` shaped `<kind>:<id>` (`project:1234`,
`chapter:58210`, `task:...`). Poll `get_job` roughly every 5 seconds until
`is_terminal`, then call `get_job_result` for the artifact URL.
## Hard limits
- 200 lines per chapter
- 3,000 characters per request with emotions on; 10,000 with emotions off
- 20 breaks and 30 s of total pause per line
- Destructive tools require `confirm=true`
- You can only ever touch the signed-in user's own content
moknah-dialogue-and-pacing6.75 KB
--- name: moknah-dialogue-and-pacing description: Produce natural dialogue and pacing in Moknah - detect speakers, cast a voice per character, merge consecutive same-speaker lines, and place [break_X] pauses at speaker changes, sentence ends and scene beats. Use when the text contains dialogue, or when rendered audio sounds choppy, rushed, or abruptly cut between voices. --- # Dialogue and pacing in Moknah The single biggest quality difference between amateur and professional Moknah output is **line structure and pause placement**. Voice choice matters less than this. ## Why line structure decides quality Each line is rendered as its own TTS request. Every render starts prosody from scratch: fresh intonation contour, fresh energy, fresh breath. So: - Splitting one speaker's continuous speech across several lines produces audible restarts - the pitch resets mid-thought and the delivery sounds stitched. - Two adjacent lines with *different* voices butt straight up against each other with no gap, so the switch sounds like an edit error rather than a reply. Both are fixed at the text/segmentation layer, before you spend a single credit. ## Rule 1 - Merge consecutive lines from the same speaker If two or more consecutive lines belong to the same speaker, **combine them into one line.** Let the voice run continuously through the whole utterance. Do this at cleaning time, before `create_project` / `add_line`, or fix it after with `update_lines`. ``` BAD (3 lines, same speaker - 3 prosody resets) line 1: "ما زلت أذكر ذلك اليوم." line 2: "كانت السماء صافية." line 3: "ولم يكن أحد يتوقع ما سيحدث." GOOD (1 line, one continuous delivery) line 1: "ما زلت أذكر ذلك اليوم. كانت السماء صافية. ولم يكن أحد يتوقع ما سيحدث." ``` Stop merging when you hit any of these ceilings: | Limit | Value | |---|---| | Characters per request, emotions ON | 3,000 | | Characters per request, emotions OFF | 10,000 | | Lines per chapter | 200 | | Breaks per line | 20 | | Total pause per line | 30 s | Never merge across a speaker change. Never merge across a scene break. ## Rule 2 - Always break on a speaker change When the voice changes, insert a pause at the **end of the outgoing speaker's line**. Without it the two voices collide and the listener misses the handoff. **Apply this by default, without being asked.** It is standard house behaviour, not an option to raise with the user. Every speaker change gets a pause unless the user has asked otherwise. ``` line 1 (Voice A): "هل انتهيت من العمل؟ [break_0.5]" line 2 (Voice B): "لم أبدأ بعد." ``` Use **0.4-0.7 s**. Fast exchanges sit at the low end; a considered reply or a change of tone sits at the high end. ### The one exception - deliberate interruption Drop the pause only when the user wants the next voice to **cut the current speaker off**: an argument, an interjection, someone talking over someone else, overlapping urgency. Then the two lines butt straight together with no gap and the abruptness is the point. Never remove the pause on your own judgement - only when the user asks for that effect. ## Rule 3 - Pause at sentence ends and beats Punctuation gives a small natural pause on its own. Add an explicit `[break_X]` only where you want the listener to *feel* the gap - the end of a thought, a paragraph, a scene, a chapter opening. Recommended starting values, then tune by ear: | Position | Duration | |---|---| | Between sentences, same speaker (only if the text feels rushed) | 0.3 - 0.5 s | | End of paragraph / thought | 0.6 - 0.8 s | | Speaker change in dialogue | 0.4 - 0.7 s | | Scene or section break | 1.0 - 1.5 s | | After a chapter title, before the body | 1.0 - 1.5 s | | Dramatic beat (suspense, revelation) | 1.5 - 2.5 s | Do not put a break after every sentence. Over-paced narration sounds like a dictation exercise. Breaks are punctuation for the ear - use them where meaning turns. For a standalone pause between blocks, put `[break_X]` on its own line. ## Detecting speakers Scan for these patterns before segmenting: - Arabic attribution verbs: `قال فلان:` / `قالت` / `أجاب` / `همس` / `صاح` - Quotation marks: `"..."` `«...»` `„..."` - The dialogue dash at line start: `- ...` - Script format: `NAME: line` Build a **speakers map** (character -> voice) and keep it consistent for the whole book. Persist it so later chapters cast identically. ## Casting - One dedicated voice per recurring character, plus a distinct **narrator** voice. - Adjacent speakers must be clearly distinguishable - do not cast two similar voices for two characters who talk to each other. - Attribution tags (`قال أحمد`) belong to the **narrator**, not the character, unless the user asks otherwise. - Get casting approved before rendering (this is a gate in every production mode). - Preview with `play_voice_sample` before the user commits. ### Emotions in dialogue Turn `emotions_mode` **on** for multi-voice dialogue - utterances are short and the voice keeps changing, so the variation between generations reads as performance rather than inconsistency. It is a clear quality gain for dialects in particular. Turn it **off** for a single-narrator production, even a dialect one. Over hundreds of consecutive lines in the same voice that variation becomes an audible stability problem. See `moknah-voice-settings`. ## Markup syntax - get this right or tags get read aloud The tag syntax depends on `emotions_mode`: - `emotions_mode: true` -> square brackets: `[happy]`, `[whispering]`, `[sighs]` - `emotions_mode: false` -> SSML-style angle brackets `<...>`, used sparingly **Never mix the two.** With emotions off, `[...]` tags are spoken aloud as text. `[break_X]` works in both modes. X is seconds, 0.1-10. Available emotion tags (emotions_mode true): `[laughing] [happy] [excited] [angry] [shouting] [sad] [crying] [surprised] [afraid] [calm] [serious] [sarcastic] [curious] [whispering] [sighs] [exhales] [wheezing] [snorts] [mischievously] [questioning] [confident] [swallows] [gulps]` Effects: `[applause] [clapping] [gunshot] [explosion]` **Emotion and effect tags are experimental and not fully reliable.** Tell the user that up front, and validate on one sample line before applying them book-wide. A line containing only tags and no real words is skipped. Tags and breaks are not billed beyond their character count. ## Checklist before rendering dialogue 1. Consecutive same-speaker lines merged. 2. A `[break_]` at every speaker change. 3. Beat pauses at paragraph, scene and chapter boundaries only. 4. Speakers map complete; every speaking character cast to a distinct voice. 5. Tag syntax matches `emotions_mode`. 6. One sample line rendered per voice and approved before the full run.
moknah-studio-reference6.66 KB
---
name: moknah-studio-reference
description: Reference for operating Moknah Audio Studio - the Project/Chapter/Line data model, the tool call sequence, create_project options, batch editing, QA status codes, job polling, what costs credits, plan gating and hard limits. Use when you need the mechanics of a specific tool or option rather than a production workflow.
---
# Moknah Audio Studio reference
Mechanics and lookup tables. For workflows see `moknah-audiobook-production`;
for parameters see `moknah-voice-settings`.
## Data model
```
Project -> Chapters -> Lines
```
- A **line** is one spoken unit and one TTS request.
- A line is *converted* once it has rendered audio.
- `total_chars` on a line means **billed** characters and is `0` until the line
renders. It is not the length of the text - never use it to estimate size.
Projects persist. After an interruption call `get_project` and continue; never
restart a book.
## Core call sequence
1. `create_project(filename, file_base64, options)` -> returns a `job_ref`
2. `get_job(job_ref)` every ~5 s until `is_terminal`; on failure read `error`
3. `get_project(project_id)` to inspect chapters and lines
4. `list_voices` -> `set_project_voice` (+ `set_line_voice_settings` for exceptions)
5. `estimate_project_generation` / `estimate_edits` -> show credits, get approval
6. `generate_audio(project_id, mode)` -> `job_ref` -> poll -> `get_job_result`
Use `get_project` for an overview (chapters + counts, no line text - safe on huge
books) and `get_chapter` only for chapters you actually need to read. Never
re-fetch what you already have.
## `create_project` options
| Option | Values / meaning |
|---|---|
| `chapter_style` | `None` / `Heading 1` / `Heading 2` / `Title` - Word chapter detection |
| `include_chapter_title` | keep the heading as the chapter title |
| `normalization` | `"0"` Basic (free) · `"2"` AI-Enhanced (1 credit/char, file only, **Standard Arabic only** - applies contextual tashkeel) |
| `enable_translation` | with `source_language` / `target_language` (file only; AI-Enhanced output is Arabic-only) |
| `line_split_mode` | `sentences` / `newline` / `custom` (+ `line_split_custom`) |
| `start_page` / `end_page` | PDF body range - skip cover, TOC, appendices |
Never use AI-Enhanced normalization on dialects or non-Arabic text.
## Standalone text-to-speech
`text_to_speech(text, voice_id, settings?, confirm_spend)` needs no project and
works on **every plan**, including free.
- `confirm_spend=false` (default) returns a **free estimate** and generates nothing
- `confirm_spend=true` starts a **background** render and returns a `job_ref` -
poll `get_job` until `is_terminal`, then `get_job_result` for the public audio
URL. Rendering runs on the worker, so the call never times out even for long
text or the slower tashkeel / emotions settings.
## Inline audio player
`play_audio(audio_url= | job_ref=, title?)` renders an inline play button (the
`ui://moknah/audio-player` MCP Apps widget) in hosts that support MCP Apps -
e.g. ChatGPT. Pass the public audio URL from a finished render, or its `job_ref`
and it resolves the URL for you. **Only Moknah audio is accepted** - the URL host
must be `moknah.io` or `*.moknah.io`, so the player can't be pointed at an
external source. In hosts without MCP Apps (currently Claude) it returns the URL
instead of a rendered player.
Settings overrides: `temperature`, `similarity`, `speed`, `expressiveness`,
`emotions_mode`, `prerecording`. Omitted keys inherit the user's saved settings.
## Batch editing - prefer this
`update_lines` applies up to **200 edits in one request**. Each item is
`{line_id, text?, settings?, qa_status?}` in any combination; items apply
independently and the response reports per-item ok/error.
Use it for AI-QA corrections across a chapter, bulk revoicing, and QA sign-off.
Fall back to `update_line` / `set_line_voice_settings` only for a single line.
Never loop single-line calls when one batch call would do.
**Adding lines in bulk:** `add_lines(chapter_id, lines=[{text, position?, settings?,
qa_status?}])` creates up to 200 new lines in ONE call (FREE), with per-item
success/failure and the new line ids. Prefer it over looping `add_line`. To split
a text blob into lines automatically instead, use `add_chapter(text=...)` or
`create_project(text=...)`.
**Reading in bulk:** `get_chapter(chapter_id)` returns all of a chapter's lines;
`get_project(project_id, with_tree=true)` returns every line with text;
`get_line_audio(line_ids=[...])` returns many lines' audio at once. Never fetch
lines one at a time.
## QA status codes
| Code | Meaning |
|---|---|
| 1 | NotStarted |
| 2 | InitialOutputReady |
| 3 | UnderReview |
| 4 | RevisionsRequired |
| 5 | AwaitingReview |
| 6 | Finalized |
## Jobs
`job_ref` is `<kind>:<id>` - e.g. `project:1234`, `chapter:58210`,
`transcription:<uuid>`, `task:<uuid>`.
Poll `get_job` until `is_terminal` (`completed` / `failed` / `cancelled`), then
`get_job_result` for the artifact URL. `get_job_result` returns a URL, not bytes.
## What costs credits
**Free:** creating projects, editing text, reordering, voice settings, QA status,
`merge_chapters_audio`, `merge_chapters_subtitle`, and every `estimate_*` call.
**Billed:** generation (TTS), PDF OCR, AI-normalization, translation,
transcription, AI-QA.
Always call the matching `estimate_*` tool and show the number before a billable
action. Never spend credits the user did not explicitly approve.
### When credits run out
`INSUFFICIENT_CREDITS` reports the shortfall. Tell the user plainly how much is
missing, offer to **reduce scope** as a real option (rendering one chapter now is
a legitimate answer), and point them to their Moknah account to manage their plan.
Treat it as a next step, never a dead end.
## Access rules - absolute
- **Owner or team only.** Every tool re-checks permissions on each object. Other
users' content is invisible and untouchable.
- **No admin.** Staff and superuser privileges are stripped over MCP. There is no
tool for any admin page or feature and none will succeed. Never offer one.
- **Plan gate.** Studio tools raise `UPGRADE_REQUIRED` for plans without Audio
Studio.
Tools that work on **every** plan, including free:
`text_to_speech`, `list_voices`, `play_voice_sample`, `estimate_text_to_speech`,
`check_credits`, `get_balance`, `get_job`, `get_job_result`, `download_result`.
## Hard limits
| Limit | Value |
|---|---|
| Lines per chapter | 200 |
| Characters per request, emotions ON | 3,000 |
| Characters per request, emotions OFF | 10,000 |
| Breaks per line | 20 |
| Total pause per line | 30 s |
| PDF upload | 50 MB |
| Batch edits per `update_lines` call | 200 |
Destructive tools (`delete_project`, `delete_chapter`, `delete_line`) require
`confirm=true`.
moknah-text-preparation4.66 KB
--- name: moknah-text-preparation description: Clean and prepare raw extracted text for narration in Moknah - strip page furniture, verbalize numbers, dates and abbreviations, build a proper-names dictionary, and segment into lines. Use after extracting text from a PDF/Word/EPUB and before creating a Moknah project or generating audio. Arabic-first but applies to any language. --- # Preparing text for narration Cleaning is **free**. Rendering is not. Every defect you leave in the text gets paid for twice: once to render it wrong, once to render it again. Do this work before any credit is spent. Log every change you make. Silently deleting the author's text is the worst failure mode in this whole workflow - always be able to show the user what you removed and why. ## 1. Structural cleanup Remove what belongs to the page, not to the book: - Lines that are only a page number - Running headers and footers that repeat across pages - Footnote blocks, plus their in-text markers (`(1)`, `¹`) - **only** when a matching footnote actually exists - Table-of-contents fragments that leaked into the body Then repair the extraction: - Merge lines broken mid-sentence (no terminal punctuation at the break) - Rejoin hyphen-split words across line ends - Collapse repeated spaces, blank lines and duplicated paragraphs ## 2. Symbols and punctuation - Bullets and list markers -> rewrite as flowing sentences. A narrator cannot read a bullet. - `/`, `\`, `_` -> a comma, or the word "or", whichever the sense requires - Keep a single hyphen used as an aside; drop decorative dashes and rules - `...` -> a period, unless the hesitation is deliberate - Strip tatweel (ـــ) - Keep `﴿ ﴾` around Qur'anic verses (confirm with the user - it affects delivery) - Normalize curly/fancy quotes to a consistent pair ## 3. Verbalization - always yours to do AI-Enhanced normalization applies **tashkeel only**. It does not turn symbols into words. You must: - **Numbers, dates, percentages, currency** -> written out in words, with correct i'rab for Arabic. `1995` -> `ألف وتسعمئة وخمسة وتسعين`, inflected to fit the sentence. - **Abbreviations** -> expanded: `د.` -> `الدكتور`, `ﷺ` -> `صلى الله عليه وسلم`, `(رض)` -> `رضي الله عنه` - **URLs and emails** -> drop them, or spell them out if the meaning depends on it - **Foreign names** -> transliterate with the `پ / گ / چ` convention (`Google` -> `گوگل`) and record every one in a **proper-names dictionary**. The names dictionary is the single most valuable artefact you produce. Keep it consistent for the entire book - a character whose name is pronounced two different ways in chapter 3 and chapter 9 is an obvious defect. Persist it and reuse it across books by the same author or in the same series. ## 4. Structure - Map TOC headings to chapter titles - Subheadings: merge into the body, or keep as their own chapter. Ask in Full Control and Guided; merge silently in Express. - Detect dialogue and build a speakers map - see `moknah-dialogue-and-pacing` ## 5. Segmentation into lines A line is one spoken unit and one TTS request. Segment where a human reader would breathe. Rules: - **Never split mid-sentence.** - Merge consecutive lines belonging to the same speaker into one line (see the dialogue skill) - this is the most common cause of choppy audio. - Max **200 lines per chapter**. - Respect the per-request character ceiling: **3,000 with emotions on, 10,000 off.** - Prefer shorter lines where the text is risky (OCR'd, heavy with numbers or foreign names) - a bad line then costs one cheap single-line re-render instead of a whole chapter. - Add missing punctuation, especially before conjunctions, so the engine knows where to breathe. `line_split_mode` options at project creation: `sentences`, `newline`, or `custom` (with `line_split_custom`). ## 6. OCR text needs stricter review If the source was scanned and OCR'd, flag it (`ocr_source = true`) and review harder. OCR reliably confuses: - `ب / ت / ث / ن` - the dotted forms - hamza placement (`أ / إ / ء / ئ`) - `ة / ه` at word end If more than ~5% of characters are unreadable, retry the extraction once, then show the user a sample rather than pushing bad text forward. ## 7. Report before you proceed Hand the user a cleaning report before creating the project: - Counts per operation (lines merged, footnotes removed, numbers verbalized...) - Before/after samples of the more aggressive edits - The proper-names dictionary - Verses or quotations detected - The resulting chapter list Get approval, then create the project. Corrections made now are free; the same correction after rendering costs a full re-render.
moknah-voice-settings7.36 KB
---
name: moknah-voice-settings
description: What every Moknah voice parameter does - voice_id, temperature, similarity, speed, expressiveness, emotions_mode, prerecording - how the server/user/project/line cascade resolves, and which settings to use for Standard Arabic, dialects, non-Arabic and poetry. Use when choosing or tuning voice settings, or when rendered audio sounds wrong.
---
# Moknah voice settings
## The cascade
Settings resolve from the most general to the most specific:
```
server default < user default < project < line
```
A `null` at any level means **inherit from the level above**. Set the project
voice once; override individual lines only by exception. Do not stamp the same
value onto 200 lines - that is what the project level is for.
- `set_project_voice` - the book's default voice and settings
- `set_line_voice_settings` - one line's override
- `update_lines` - batch overrides, up to 200 lines in a single call
## The parameters
### `voice_id`
Which voice speaks. Must come from `list_voices` - never invent an id, and only
voices the user actually has access to will work. Premium voices add a
percentage fee on top of the per-character rate, so the same text costs more.
### `prerecording` - the text pre-processing pipeline
This is the highest-impact setting for Arabic. It selects how the text is
prepared *before* synthesis.
| Value | Mode | Use for |
|---|---|---|
| `1` | Normal | Dialects, all non-Arabic languages |
| `2` | Arabic Tashkeel | Standard Arabic (fusha) only - applies contextual diacritics |
| `3` | AI-poet | Poetry and metered verse |
`prerecording: 2` markedly improves fusha pronunciation because Arabic script is
ambiguous without diacritics (`كُتب` vs `كَتب`). **Never use it on dialects or
non-Arabic text** - it will impose fusha diacritics on words that do not take
them and the output gets worse, not better. It is billed at the advanced rate.
### `emotions_mode` - the expressive engine
`true` enables the expressive engine and square-bracket emotion tags. `false`
keeps the standard engine.
It sounds **markedly better for dialects and for dialogue**. But every generation
comes out slightly different, so it carries a real stability cost.
**The deciding question is how many voices the production uses:**
- **Dialogue / multi-voice -> emotions ON.** Utterances are short and the voice
keeps changing, so variation between lines reads as natural performance.
- **Single narrator ("one man show") -> emotions OFF**, even for dialect. Across
hundreds of consecutive lines in one voice, that same variation turns into an
audible inconsistency the listener notices immediately. Stability wins.
This supersedes any blanket "dialects always use emotions" rule: for dialect
content the mode still helps, but only when more than one voice is speaking.
Other consequences:
- Tag syntax flips: `true` -> `[happy]`; `false` -> SSML `<...>`. Mixing them
makes the tags get read aloud as literal text.
- The per-request character limit drops: **3,000 with emotions on, 10,000 off.**
### `speed`
Delivery rate. Range `0.7` - `1.2`; `1.0` is natural. Leave at 1.0 unless the
user asks. Small moves (0.95, 1.05) are usually enough - anything past ~1.15
starts to sound clipped.
### `similarity`
How closely the output tracks the reference voice's timbre and identity.
Keep it **high** as the default - that is what makes a voice sound like itself
across a whole book. Lower it only when the voice is a clone made from noisy
source audio, where a high value faithfully reproduces the noise and artefacts
along with the voice.
### `temperature`
Variability between renders. Lower = more consistent, steadier, more predictable
delivery. Higher = more variation and life, at the cost of consistency.
**Hard ceiling: never go above 0.75.** Past that the model begins to
**hallucinate** - inventing words, dropping text, mangling phrases. This is not a
"slightly less polished" threshold; it is the point where the output stops being
trustworthy and has to be checked word by word.
- Fusha narration: leave at default. Consistency across hundreds of lines matters
more than per-line colour.
- Dialect: nudge slightly up - dialect delivery sounds flat when too tightly
constrained - but stay well below 0.75.
### `expressiveness`
How much prosodic style is applied - emphasis, dynamic range, dramatic contour.
**Keep it at 0. Never raise it on your own initiative.** This parameter is
extremely sensitive - small increases cause hallucination, not merely a livelier
reading.
Raise it only when the user explicitly asks for it, and when they do:
1. **Warn them it risks hallucinated audio.**
2. Render a single sample line and have them check it before applying it anywhere
else.
If output starts inventing or dropping words, this is the first setting to put
back to 0.
## Settings by scenario
`expressiveness` is `0` in every row - it is never raised except on explicit
user request. `temperature` never exceeds `0.75` in any row.
| Scenario | prerecording | emotions_mode | temperature | expressiveness | speed |
|---|---|---|---|---|---|
| Fusha, single narrator | 2 (Tashkeel) | false | default | 0 | 1.0 |
| Fusha, single-paragraph input | 2 | true | default | 0 | 1.0 |
| Dialect, single narrator | 1 | **false** (stability) | slightly up, < 0.75 | 0 | 1.0 |
| Dialect, dialogue / multi-voice | 1 | **true** | slightly up, < 0.75 | 0 | 1.0 |
| Non-Arabic, single narrator | 1 | false | default | 0 | 1.0 |
| Non-Arabic, dialogue / multi-voice | 1 | **true** | default | 0 | 1.0 |
| Poetry / verse | 3 (AI-poet) | false | default | 0 | 1.0 |
Detect the language from the text itself and apply the matching row - do not ask
the user to specify it if the text makes it obvious.
## Voice choice by content
- **Fusha book** -> a fusha voice. A dialect voice reading fusha sounds livelier
and is acceptable if the user prefers it.
- **Dialect book** -> the voice *must* match the dialect. Prime it by opening
with a strongly dialectal word so the voice settles into the register.
- **Dialogue** -> a distinct voice per character plus a separate narrator voice.
See the `moknah-dialogue-and-pacing` skill.
## Diagnosing bad output
| Symptom | First thing to change |
|---|---|
| **Invented words, dropped or mangled text (hallucination)** | `temperature` is above 0.75, or `expressiveness` is above 0. Put both back down - this is the cause almost every time |
| Arabic words mispronounced / wrong vowels | fusha? set `prerecording: 2`. Dialect? make sure it is `1` |
| Tags read aloud as words | `emotions_mode` is false but `[...]` tags are in the text |
| Delivery inconsistent between lines in one voice | `emotions_mode` is on for a single-narrator book - turn it off; then lower `temperature` |
| Flat, lifeless dialect narration | raise `temperature` slightly (stay < 0.75); turn on `emotions_mode` **only** if multi-voice |
| Voice does not sound like the reference | raise `similarity` |
| Clone reproduces hiss/artefacts | lower `similarity` |
| Choppy, restarting delivery mid-thought | not a settings problem - merge the lines (dialogue skill) |
Change **one** setting at a time and re-render a single line with
`generate_lines`. Do not re-render a chapter to test a hypothesis.
## Cost note
Voice settings are free to change - only rendering costs credits. Iterate on
settings against one short sample line until the user approves, then run the
book. Fixing a setting after a full render costs a full re-render.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 1, 2026 · 12:00 UTC
- Collection status
- Collected
plugin_asdk_app_6a58bb03727c8191a4607c78a18be2c6
Download listing JSON