← MoknahCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Moknah
Snapshot Sep 30, 2026 · 22:52 UTC · version 1.0.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "moknah-text-preparation",
"description": "Clean and prepare raw extracted text for narration in Moknah - strip page furniture, verbalize numbers, dates and abbreviations, build a proper-names dictionary, and segment into lines. Use after extracting text from a PDF/Word/EPUB and before creating a Moknah project or generating audio. Arabic-first but applies to any language.",
"included_files": [],
"skill_md_contents": "---\nname: moknah-text-preparation\ndescription: Clean and prepare raw extracted text for narration in Moknah - strip page furniture, verbalize numbers, dates and abbreviations, build a proper-names dictionary, and segment into lines. Use after extracting text from a PDF/Word/EPUB and before creating a Moknah project or generating audio. Arabic-first but applies to any language.\n---\n\n# Preparing text for narration\n\nCleaning is **free**. Rendering is not. Every defect you leave in the text gets\npaid for twice: once to render it wrong, once to render it again. Do this work\nbefore any credit is spent.\n\nLog every change you make. Silently deleting the author's text is the worst\nfailure mode in this whole workflow - always be able to show the user what you\nremoved and why.\n\n## 1. Structural cleanup\n\nRemove what belongs to the page, not to the book:\n\n- Lines that are only a page number\n- Running headers and footers that repeat across pages\n- Footnote blocks, plus their in-text markers (`(1)`, `¹`) - **only** when a\n matching footnote actually exists\n- Table-of-contents fragments that leaked into the body\n\nThen repair the extraction:\n\n- Merge lines broken mid-sentence (no terminal punctuation at the break)\n- Rejoin hyphen-split words across line ends\n- Collapse repeated spaces, blank lines and duplicated paragraphs\n\n## 2. Symbols and punctuation\n\n- Bullets and list markers -> rewrite as flowing sentences. A narrator cannot\n read a bullet.\n- `/`, `\\`, `_` -> a comma, or the word \"or\", whichever the sense requires\n- Keep a single hyphen used as an aside; drop decorative dashes and rules\n- `...` -> a period, unless the hesitation is deliberate\n- Strip tatweel (ـــ)\n- Keep `﴿ ﴾` around Qur'anic verses (confirm with the user - it affects delivery)\n- Normalize curly/fancy quotes to a consistent pair\n\n## 3. Verbalization - always yours to do\n\nAI-Enhanced normalization applies **tashkeel only**. It does not turn symbols into\nwords. You must:\n\n- **Numbers, dates, percentages, currency** -> written out in words, with correct\n i'rab for Arabic. `1995` -> `ألف وتسعمئة وخمسة وتسعين`, inflected to fit the\n sentence.\n- **Abbreviations** -> expanded: `د.` -> `الدكتور`, `ﷺ` -> `صلى الله عليه وسلم`,\n `(رض)` -> `رضي الله عنه`\n- **URLs and emails** -> drop them, or spell them out if the meaning depends on it\n- **Foreign names** -> transliterate with the `پ / گ / چ` convention\n (`Google` -> `گوگل`) and record every one in a **proper-names dictionary**.\n\nThe names dictionary is the single most valuable artefact you produce. Keep it\nconsistent for the entire book - a character whose name is pronounced two\ndifferent ways in chapter 3 and chapter 9 is an obvious defect. Persist it and\nreuse it across books by the same author or in the same series.\n\n## 4. Structure\n\n- Map TOC headings to chapter titles\n- Subheadings: merge into the body, or keep as their own chapter. Ask in Full\n Control and Guided; merge silently in Express.\n- Detect dialogue and build a speakers map - see `moknah-dialogue-and-pacing`\n\n## 5. Segmentation into lines\n\nA line is one spoken unit and one TTS request. Segment where a human reader would\nbreathe.\n\nRules:\n\n- **Never split mid-sentence.**\n- Merge consecutive lines belonging to the same speaker into one line (see the\n dialogue skill) - this is the most common cause of choppy audio.\n- Max **200 lines per chapter**.\n- Respect the per-request character ceiling: **3,000 with emotions on, 10,000\n off.**\n- Prefer shorter lines where the text is risky (OCR'd, heavy with numbers or\n foreign names) - a bad line then costs one cheap single-line re-render instead\n of a whole chapter.\n- Add missing punctuation, especially before conjunctions, so the engine knows\n where to breathe.\n\n`line_split_mode` options at project creation: `sentences`, `newline`, or\n`custom` (with `line_split_custom`).\n\n## 6. OCR text needs stricter review\n\nIf the source was scanned and OCR'd, flag it (`ocr_source = true`) and review\nharder. OCR reliably confuses:\n\n- `ب / ت / ث / ن` - the dotted forms\n- hamza placement (`أ / إ / ء / ئ`)\n- `ة / ه` at word end\n\nIf more than ~5% of characters are unreadable, retry the extraction once, then\nshow the user a sample rather than pushing bad text forward.\n\n## 7. Report before you proceed\n\nHand the user a cleaning report before creating the project:\n\n- Counts per operation (lines merged, footnotes removed, numbers verbalized...)\n- Before/after samples of the more aggressive edits\n- The proper-names dictionary\n- Verses or quotations detected\n- The resulting chapter list\n\nGet approval, then create the project. Corrections made now are free; the same\ncorrection after rendering costs a full re-render.\n"
}SHA-256: a49103aec05654616a2278f57e04dc90c220d20eb13a5c8c7a685f3bac2d70bf