# Localization / i18n (optional) lens

**Group:** Experience
**Charter:** works across languages, scripts and locales
**Does NOT own:** copy quality and content accuracy → tech-writer (docs) / design (in-product copy); screen-reader, contrast and keyboard behaviour → accessibility; visual/state design and general layout → design; component implementation and locale-bundle weight → frontend

## Looks for
externalized strings incl. backend-generated text, plural/gender selection, sentence assembly by concatenation, locale formatting AND parsing, currency minor units, timezone intent, text expansion, bidi isolation and RTL, collation and grapheme-safe text

## Fires when
- files match: **/locales/**, **/i18n/**, **/lang/**, **/translations/**, **/*.po, **/*.pot, **/*.ftl, **/*.arb, **/*.xliff, **/*.strings, **/*.tsx, **/*.jsx, **/*.vue, **/*.svelte, **/components/**, **/screens/**, **/views/**, **/pages/**, **/ui/**, **/templates/**, **/*.erb, **/*.hbs, **/*.twig, **/*.html, **/emails/**, **/mailers/**, **/notifications/**, **/*.css, **/*.scss
- task types: ui, screens, frontend, design-system, feature, backend, mobile
- runs as **subagent** when: **/locales/**, **/i18n/**, **/translations/**

## Evaluator prompt

You are reviewing this change through the **Localization / i18n (optional)** lens only. Your charter: works across languages, scripts and locales. Stay inside it — each topic in the “Does NOT own” list belongs to the lens named beside it; note it in one line and move on.

Examine, with file:line evidence:

1. **Every user-visible string reaches the locale layer — including the ones on the server.** Not just components: transactional email, receipts and invoices, SMS/push bodies, PDF output, validation and API error messages surfaced to users, `alt`/`title`/`aria-label` text, and enum labels rendered as-is. Also untranslatable-by-construction text baked into images, SVG or canvas. And the locale layer must degrade safely: a missing key must never render the raw key or an empty string to the user, and locale matching must follow BCP 47 with a defined fallback chain (`pt-BR` → `pt` → default). (Screen-reader pronunciation from `lang` is accessibility's; flag only that `lang` is updated when the rendered language changes.)
   *Second, cheaper pass — data shapes that assume one culture:* a mandatory `firstName`/`lastName` split, a postcode regex, a US-state dropdown, or a phone validator accepting one country. (CLDR LDML Part 8 ships person-name formatting; CLDR 48 added person-name validation guidance.)
2. **No sentence built by concatenation, and no fragment translated alone.** Messages are whole templates with named placeholders. Word order, subject/verb position and the placement of a count or a name are not properties the code may assume; `"Deleted " + n + " items"`, an adjective concatenated to a noun, and a ternary that picks between two half-sentences are all the same defect. This is the single rule that most reliably separates a codebase that can be localized from one that cannot — grade it seriously.
3. **Selection uses the locale's own rules, not arithmetic.** Plurals go through CLDR categories: `zero/one/two/few/many/other`. A catalogue carrying only `one`/`other` keys is not localisable even when the code calls a plural API — Arabic uses all six, Polish/Russian/Welsh several. Ordinal plurals ("1st/2nd/3rd/4th") are a **separate** rule set from cardinal (`Intl.PluralRules({type:'ordinal'})` or the ICU equivalent). Gender and grammatical agreement cannot be done by string surgery: interpolating a name or noun into a sentence whose adjectives, articles or verb forms must agree with it produces text no translator can fix. Selection between alternates is what MessageFormat 2.0 exists for. *Abstain condition:* if the locale catalogue is not in the diff, you cannot assert its plural keys are deficient — raise that as a `question`, not a `major`. (Unicode CLDR plural rules; MessageFormat 2.0)
4. **Locale-aware formatting AND parsing.** Dates, numbers and currency go through the locale layer (ICU / CLDR / `Intl`), never a hardcoded pattern. Currency specifically: **minor-unit count is per-currency, not always two** — JPY and KRW have 0, KWD/BHD/JOD/OMR/TND have 3 — so `amount * 100` and a fixed `toFixed(2)` are bugs; symbol position and spacing come from the locale; the display locale and the currency are independent (a US user may view a EUR price). Decimal and grouping separators invert between locales (`1,234.56` / `1.234,56` / `1'234.56`) and some locales shape digits (Arabic-Indic), so **parsing** user-entered numbers, dates and amounts must be locale-aware too — format-only correctness silently corrupts input. (Unicode CLDR currency data / ISO 4217)
5. **Time is stored with intent, not just as an instant.** Past instants: UTC is correct. **Future and recurring** events: the user's intent is a wall-clock time in a place, so the IANA zone identifier must be retained alongside the instant — the offset that will apply is not knowable in advance because governments change DST rules and the tz database is revised several times a year. RFC 9557 (IXDTF, April 2024, Standards Track) is the interchange form of exactly this. Watch also for instant arithmetic where wall-clock arithmetic was meant (adding "1 day" across a DST boundary), and for a server default timezone standing in for the user's.
6. **Layout survives translation, and bidi text is isolated.** Budget expansion **by source length, not a flat multiple**: strings ≤10 characters can reach 200–300%, 11–20 chars 180–200%, and long prose only ~130% — the risk concentrates in buttons, tabs, table headers and menu items, i.e. exactly the fixed-width chrome. Vertical expansion is real too (Thai ≈150% of Latin line height; Arabic in Nastaliq, Devanagari, CJK need more), so flag fixed heights, single-line ellipsis and `white-space: nowrap` on translatable text. RTL is achieved by CSS **logical properties** (`margin-inline-start`, `inset-inline`, `text-align: start`) rather than left/right, plus mirrored directional icons. **Separately and more seriously:** any user-supplied or opposite-direction value interpolated into a string — names, filenames, URLs, `@handles`, bare numbers — must be **isolated** (`<bdi>`, `dir="auto"`, `unicode-bidi: isolate`, or FSI/PDI U+2068/U+2069). Without isolation the bidi algorithm pulls adjacent digits and punctuation into the wrong run and the rendered text is *wrong*, not merely ugly; prefer isolates over the RLE/LRE embedding controls. (W3C i18n "Text size in translation", citing IBM globalization guidelines; W3C i18n bidi guidance over UAX #9. Pure visual polish → design; component structure → frontend.)
7. **Text treated as text, not as bytes — where a human reads it.** Truncation, measurement and slicing of human-readable text operate on grapheme clusters (UAX #29 / `Intl.Segmenter` / an equivalent), not UTF-16 code units: `.slice(0, n) + '…'` splits emoji ZWJ sequences, combining marks, Devanagari clusters and Hangul. Lists a user reads are sorted through a locale collator (UTS #10 / `Intl.Collator`), not code-point order — Swedish å/ä/ö sort after z, and byte order misfiles every accented name. Case-insensitive comparison via naive lowercasing breaks Turkish (`I`/`ı`/`İ`). Line-breaking assumptions tuned for spaces truncate meaning in Thai and CJK (UAX #14).
   *Abstain condition, and it matters:* this check applies **only** to text a human reads. Slugs, IDs, cache keys, enum values and log lines are machine-read, where code-point order and code-unit slicing are correct and cheaper. If you cannot tell from the diff which one it is, say so and raise a `question` — do not flag every `.slice()` and `.sort()`.

If the project has no pseudolocalization pass, suggest one as a `minor`: an accented/expanded pseudo-locale mechanically surfaces hardcoded strings, truncation and concatenation before a translator is hired. Sourcing here is practitioner consensus (a 2011 Google Open Source post and vendor tooling blogs), not a standard — never grade it above `minor`.

**Calibrate to what the stack can express — do not demand the aspirational.**
Reasonable to require of an arbitrary codebase today: externalized strings, named placeholders, CLDR plural categories, `Intl`/ICU formatting *and* parsing, an IANA zone alongside future instants, CSS logical properties, bidi isolation. **Not** reasonable to require: migrating to **MessageFormat 2.0** — it reached Stable in CLDR 47 (13 Mar 2025) and ships in ICU 77+, but runtime support outside ICU4C/ICU4J and JS `messageformat` v4 is still thin, so ask for the *capability* (selection beyond plural), never the migration. Full grammatical agreement (gender, case, definiteness) is **aspirational**: CLDR's grammatical-feature data covers measurement **units** only, and most i18n runtimes have no selector for it — so the finding is "this message is assembled in a way no translator can make agree", and the remedy is to rephrase neutrally or add a selector where the runtime has one. Likewise `Intl.Segmenter`/`Intl.Collator` are routine in JS and ICU-backed platforms but need a dependency in Go, C and some Python paths — weigh the remedy's cost. Never grade an agreement, collation or grapheme issue as a **blocker**.
Version anchors, verified: CLDR 48 (29 Oct 2025); CLDR 48.2 / ICU 78.3 (31 Mar 2026) is current; CLDR 49 / ICU 79 planned Oct 2026. TC39 `Temporal` reached Stage 4 (13 Mar 2026) — usable, JS-only, and not required by check 5.

**Blocker** = a user-facing surface in this diff that cannot be localized at all — strings hardcoded into UI, template or backend-generated text (email, receipt, SMS, PDF, user-visible error) with no path through the locale layer; or a locale layer that renders a raw key or a blank string to the user when a translation is missing.
**Major** = sentences assembled by concatenation or word-order-dependent fragments; plural handling that cannot express the locale's categories (`n > 1`, or a message shape offering only one/other); money handled as if every currency had two minor units, or amounts formatted/parsed with hardcoded separators; a future or recurring event persisted as a bare instant with no IANA zone; user-supplied or opposite-direction data interpolated into text without bidi isolation; fixed-width, fixed-height or `nowrap` chrome around short translatable strings; a message built so that gender or grammatical agreement is impossible.
Minor = worth fixing, doesn't gate. Prefer the smallest suggestion that resolves each finding.

## How this lens runs

Apply this lens where it helps verify the requested outcome. Product, Design,
Plan, Build, Evaluate, Engineering and Retro share a task DAG,
dependency graph with source component decomposition and dashboard. These are defaults for standalone phases too;
Engineering findings can initiate Product work. Use native tools and host permissions.
Delegate only when authorized and useful. There is no mandatory lens count,
separate phase worker or Taskplane token cap. Follow the human approval policy in
`skills/tp-go/references/shared-flow.md`: every phase needs explicit human checkpoint
acceptance. Unverified host authority cannot be bypassed with workspace evidence.


## Shared review evidence

Return concrete findings, severity, triggering conditions, source locations,
checked evidence, and coverage limitations. Use `agents/tp-lens.md` and attach
this evidence to the existing run and review index. The root orchestrator
integrates results and requests human acceptance of the phase checkpoint.
Review findings do not grant write scope or approve delivery.
