← Files VeraARCHIVED FILE
modules/journal-bank-reconciliation/references/workflow-reference.md
20.3 KB · Oct 3, 2026 · 06:30 UTC
# Journal-Bank Reconciliation Reference This reference documents the deterministic boundary for the plugin. Codex reads it only when a run needs more detail than the main skill. ## Stable Columns Normalized bank and journal outputs use these canonical columns: - `side` - `transaction_id` - `transaction_date` - `amount_signed` - `amount_abs` - `description` - `beneficiary` - `reference` - `movement_number` - `account` - `currency` - `unit` - `entity_ref` - `party_ref` - `direction` - `source_file` - `source_sheet` - `source_row` `amount_signed`, `amount_abs`, `bank_amount`, `journal_amount`, and `amount_delta` are canonical non-exponent Decimal text. Binary floats are not used for reconciliation comparisons. Match outputs use: - `status` - `stage` - `bank_transaction_id` - `journal_transaction_id` - `bank_date` - `journal_date` - `date_diff_days` - `bank_amount` - `journal_amount` - `amount_delta` - `bank_description` - `journal_description` - `shared_references` - `review_note` Relationship-residual outputs use `side`, `record_ref`, `transaction_id`, `record_amount`, `allocated_amount`, `residual`, `currency`, `unit`, `entity_ref`, and `party_ref`. These are an exact projection of the allocation ledger; they do not classify, allocate, or force residuals to zero. ## Header and Mapping Authority Automatic qualification is deliberately narrow: exactly one row in the first 30 rows must contain an unambiguous set of exact, supported headers for date and either signed amount or debit/credit. A value-profiled date, fuzzy label, or numeric column position may be proposed in `suggested_recipe.json`, but it has no qualification authority and emits zero movements. For CSVs, `csv_field_delimiter` is transport syntax and is distinct from `decimal_separator` and `thousands_separator`. The bounded profile tests only comma, semicolon, tab, and pipe over at most 128 KiB and 100 non-empty records. A candidate is unique only when exactly one strict parse produces at least two constant-width records with more than one field. Only the uniquely profiled comma default may participate in the exact automatic contract. A non-default delimiter or an explicit delimiter that does not uniquely match the profile requires a current reviewed mapping receipt. Ambiguous delimiters return `needs_review`; unsupported delimiters return `unsupported_source_layout`. Both emit zero rows. LF, CRLF, and CR record terminators are mechanically normalized to LF through a chunked private transport copy before the full parse. A record terminator is not a mapping field and is distinct from both field and numeric separators. The full CSV parse is strict: errors are not ignored and ragged rows are not truncated, including malformed rows beyond the bounded profile. When the exact contract does not apply, Codex must review the physical header row, mapping, CSV field delimiter, and numeric separator convention. It then uses `build_mapping_review_receipt` to seal that content against the current content-addressed source artifact reference and adapter version. Hand-editing the mapping without a valid receipt does not qualify it. A changed source, mapping, header row, separator convention, or adapter makes the receipt stale. The v6 mapping receipt also binds `date_convention`. Native date/datetime cells, valid compact `YYYYMMDD`, valid year-first text, and integral spreadsheet serial dates in the bounded Excel range are mechanical. Day/month text is evaluated under both supported interpretations. If both are valid, the source emits zero rows until the reviewer seals exactly `day_first` or `month_first`; parser-list order has no authority. A populated invalid date fails the complete source even if the row has a stable reference. Only a truly blank date with a stable reference can be emitted as `emitted_reference_only`. The additive v7 mapping contract supports full Italian textual-month dates only after the reviewer seals `date_locale: it` against the current source. It accepts the exact full-month vocabulary and valid Gregorian dates; unknown, abbreviated, mixed-language, embedded, or invalid forms fail closed. A v7 mapping receipt may also bind `non_movement_summary_labels`. An exact reviewed label excludes a monetary row only when its mapped date is blank and its explicit reference and movement-number fields contain no stable token. Labels never override an actual date or stable reference. Every reviewed path must declare `potential_monetary_columns` exactly as derived from the current parsed source and must provide `excluded_monetary_columns`, even when it is empty. Each potential monetary column must be mapped to `amount`, `debit`, or `credit`, or explicitly excluded. Incomplete, extra, or stale dispositions have no review authority and emit zero rows. ## Mapping Fields Use `amount` when the file has a single signed amount column. Use `debit` and `credit` when the file splits debit and credit. The script calculates signed amount as debit minus credit. Reference fields can contain document numbers, CRO/TRN, invoice references, IBAN fragments, or other stable identifiers. Only explicit `reference` and `movement_number` fields participate in the reference stage. Descriptions and beneficiary names remain review context. Generic words such as `invoice`, `payment`, `document`, or `transfer` are not stable identifiers; reference tokens must contain a non-generic identifier with digits. `currency`, `unit`, `entity_ref`, `party_ref`, and `direction` define the relationship perimeter. Missing required values may be supplied only through a reviewed policy default. An explicit direction column may contain canonical `positive`, `negative`, and `zero` values directly. Other categorical values require a reviewed, source-bound `direction_value_mapping` that exactly covers the observed non-canonical vocabulary. The mapping translates each source label to one canonical direction and is sealed in the mapping receipt. No universal debit/credit polarity is assumed. A missing label mapping, an extra unobserved label, or disagreement between the mapped direction and exact signed amount withholds the complete source. ## Reviewed Relationship Policy Every run requires a `journal_bank_relationship` reviewed decision receipt, sealed with `build_relationship_review_receipt` against the current bank and journal source artifact references. The supported relationship shapes are `one_to_one`, `one_to_many`, `many_to_one`, and `many_to_many`; evidence reuse is disabled, and currency and unit equality are mandatory. The policy also records: - whether entity and party must agree; - `absolute_amount`, `same_sign`, or `opposite_sign` direction treatment; - any currency, unit, entity, or party defaults; - exact amount tolerance; - date window in calendar days. Execution arguments must equal the reviewed tolerance and date window. Relationship tolerance accepts canonical decimal text, `Decimal`, or integer values and is persisted as canonical Decimal text. Floats, booleans, localized/noncanonical text, non-finite values, and negative values are rejected rather than guessed. The relationship adapter is `journal_bank.relationship.v3`, version `3`. Version 3 seals both batch-safe singleton allocation and conserved grouped reference allocation. Older relationship receipts are stale and must be reviewed again. ## Matching Stages Each wave is computed from a snapshot of all currently unmatched rows. Only bank rows with exactly one eligible candidate whose target journal row is not also the singleton target of another bank row enter the batch. The complete batch is accepted together; source row order never breaks target collisions. 1. `reference`: conflict-free singleton candidates with an explicit shared reference or movement token and amount inside the exact tolerance. Date evidence is optional for this explicit-identifier stage. Conflict-free reference waves repeat until no further safe reference singleton remains. 2. `reference_group`: one bank movement to many journal rows, or many bank movements to one journal row, only when a shared stable reference or a complete explicit identifier list defines the group. Lists come only from the mapped reference field; whitespace, comma, semicolon and pipe separate identifiers, while internal punctuation is preserved. Each identifier must contain letters and digits. At most 100 distinct identifiers are admitted; missing or duplicated opposite-side identifiers withhold the entire group. The reviewed shape must permit it, every row must be inside the reviewed perimeter, and exact Decimal group totals agree within tolerance. Any row participating in more than one possible group keeps all overlapping groups unmatched. 3. `amount_date_unique`: the first conflict-free singleton amount/date batch after reference matching is exhausted. Both rows require actual dates inside the configured date window. 4. `amount_date_single`: later conflict-free singleton waves containing only candidates that became singleton after an earlier amount/date batch removed other candidates. Later waves repeat until no safe singleton remains. Rows are not reused. Ambiguity stays unmatched. Multiple singleton bank rows targeting the same journal row remain ambiguous; none receives a row-order preference. Candidates outside the reviewed currency, unit, entity, party, or direction perimeter are never matched. ## Codex-Only Residual Resolution Funnel After a qualified deterministic run, `semantic_review.py prepare` projects the unmatched partitions into a bounded candidate graph and records the user's required certainty threshold. Preparation first validates the current output receipts and material-value ledger replay. An edge exists only when the core candidate predicate accepts the exact amount and reviewed currency, unit, entity, party, and direction perimeter, and the pair also has either a shared stable explicit reference or two actual dates inside the reviewed date window. Description and beneficiary similarity cannot create an edge. The graph groups eligible rows into deterministic connected components. The preparer admits the complete residual only when every component fits one packet inside the fixed bank-row, journal-row, edge, component-count, graph-byte, and prompt-byte limits. It does not select a partial packet or automatically chunk the remainder. A one-bank/one-journal component that should already have been resolved by deterministic singleton matching, any over-cap component, or an over-cap complete residual defers the entire model step. The preparer reports `worker_required: false`, records the reason, and preserves every bank movement in the human-review queue. Packet projection happens after source-column mapping. The worker receives only populated canonical fields for unresolved bank rows and their hard-compatible candidates. Unmapped raw columns, empty canonical fields, physical source locators, and the mechanically derived absolute amount remain outside model context. The opaque transaction ID preserves local linkage and the signed amount remains available as evidence. Only the main Codex chat may orchestrate the optional model pass. It keeps its existing model and invokes `semantic_review.py run-worker`, which requests one separate ephemeral `gpt-5.6-luna` process at max reasoning for the complete selected packet. Never reconstruct or run the underlying `codex exec` command directly. The launcher is fail-closed to the pinned macOS build, Codex CLI and hash, Seatbelt and canary hashes, and deny-default profile hash. It uses a mode-`0700` capsule, prompt stdin, bounded parent capture, a read-only inner sandbox, ignored project rules, absent-or-empty global Codex instructions, and enumerated feature disables. Exact-read and outside-read-denial canaries run before the worker; the qualified boundary also denied the hidden `view_image` path access to an outside nonce image. Do not launch a worker per candidate, change the model of the current chat, or use Luna to run source qualification, matching, ledger replay, assurance, or final review. The boundary permits exact Codex authentication and installation-ID access and outbound network access, so the packet is transmitted to the OpenAI Codex service. The installation-ID file has the narrow write permission required by the pinned build but must remain byte-identical. The capsule is deleted after the turn. The response, JSONL, stderr, and content-bound launch receipt are published together only after packet, source, runtime-input, executable, and boundary checks replay successfully. Worker decisions are limited to `suggest_match`, `ambiguous`, `no_match`, and `needs_evidence`, plus a strongest supported resolution level and optional classification or identified counterparty. A suggested journal row must be an existing graph neighbor; other verdicts cannot name a journal row. Validation requires the current graph digest, exactly one review of every selected component and bank row, global non-reuse of journal evidence, bounded evidence and rationale fields, a strict thread/turn/item lifecycle, no JSONL-visible forbidden item, and equality between the final event message and retained response. Any violation rejects the entire response. JSONL visibility is incomplete and therefore cannot prove that no hidden tool path ran; the pinned outer filesystem boundary is the primary containment control. JSONL also does not attest the selected model or reasoning effort, so worker metadata records both as requested rather than observed and validation requires the launch receipt. Validated results live outside the canonical reconciliation directory in its real sibling named `semantic-review`. `run-all` launches at most one worker for the complete admitted residual. Raw Luna output remains advisory until validation. The validator then applies accepted decisions to `semantic_resolution_application.json`, `resolution_funnel.json`, and `human_review_queue.json`, and writes `operational_review_payload.json` as the bank-side human workbench input. Unmatched journal rows remain strict evidence and candidate context but are not standalone work items. The funnel levels are `classified`, `candidate_match`, `beneficiary_match`, `identifier_match`, and `perfect_match`; “at least” totals are cumulative even though each movement is assigned only its highest level. Levels express sufficiency, not the presence of every weaker evidence field. Only deterministic replay may assign `perfect_match`; Luna may clear operational review at a lower user-selected threshold but cannot alter native matches, ledgers, artifact receipts, assurance gates, or report readiness. If the current preparation closure remains valid, worker unavailability or invalid output is recorded as `worker_failed`; source or packet tampering instead fails without writing a falsely bound status. A validated generation remains terminal until a new preparation archives its fixed-name artifacts under `semantic-review/history/`. In every failure mode, the main Codex chat and deterministic reconciliation remain unchanged. ## Source Qualification - Tabular files emit rows only after date and amount or debit/credit fields are mapped and every populated mapped monetary cell parses exactly. - CSV field delimiter and date authority follow the bounded v6 contract above. Non-default choices, profile mismatches, and full potential-monetary-column dispositions are sealed in `journal_bank.tabular.v6` mapping receipts; receipts from adapter v5 and earlier are stale. - Italian textual-month dates and exact reviewed blank-date summary labels use the additive `journal_bank.tabular.v7` receipt. Adapter selection is explicit; existing numeric-date v6 sources are not silently upgraded. - The strict full-file CSV parse follows bounded delimiter profiling. A malformed or ragged record anywhere in the population is a parser failure and emits zero rows; LF, CRLF, and CR differ only as transport syntax. - Dates accept the mechanical forms described above. Ambiguous day/month values require reviewed source-bound authority; invalid populated values fail the source and emit zero rows. - Non-canonical direction labels emit rows only after complete reviewed source-value mapping. - Every monetary candidate receives a row disposition. An invalid monetary value or a missing date without an explicit reference blocks the source and emits zero rows. A missing date with a stable explicit identifier is emitted as `emitted_reference_only` and can participate only in the reference stage. A blank-date/no-reference row may instead be `excluded_reviewed_summary` only through an exact receipt-bound v7 summary label. - Ambiguous separator syntax such as `1.000` is rejected unless the recipe explicitly declares the separator convention. - `journal_bank.pdf_table.v1` accepts ruled or positionally aligned text-PDF tables only when they have an explicit date header and at least one monetary header, every page repeats the same physical header, and a current source-bound mapping receipt approves amount/debit/credit roles, signs, and every excluded monetary column. Page/table/row coordinates remain lineage. - Generic PDF text, inconsistent page tables, and OCR-only input emit zero rows with `unsupported_source_layout`. A date plus numeric tokens on a free-form line is not a movement. Narrow balance, total, scalare, and conditions classifications may still be retained as non-movement review evidence. - A supplied sample must contain movement identifiers and select journal rows. Empty, invalid, or nonmatching samples block instead of falling back to the full journal. - Parser failure is reported separately from a readable but unsupported layout; available receipts, qualifications, gates, lineage, diagnostics, and audit are still written before the run returns a block. ## Assurance Artifacts Inspection and reconciliation bind source bytes to content-addressed `input_receipts.json` records and `source_qualifications.json`. Reviewed mapping and relationship decisions are collected in `reviewed_decisions.json`. `lineage.json` points every emitted transaction to the physical workbook sheet and row (or physical CSV row). `relationship_ledger.json` records exact non-reusing allocations and residuals, including permitted one-to-many and many-to-one reference groups. `relationship_residuals.csv` projects the record, allocated, and residual components for every bank and journal row. `material_value_ledger.json` freshly replays matching and the relationship ledger, then binds every declared match and residual field to the exact prepared row, CSV row/column, and XLSX cell. A blocked run still writes the complete reviewable native package, including an empty `relationship_residuals.csv` and an explicit blocked relationship ledger. Only `material_value_ledger.json` is absent when source qualification or relationship authority blocks, because material reconciliation never ran. The execution boundary closes an exact 24-file contract covering launcher configuration, UI assets, Python/Node code, and the shared assurance kernel. Supported Python entries validate the physical tree before local imports; MCP does so before manifest parsing and invokes Python with `-I -B`. Unowned bytecode caches and other physical entries therefore fail before execution. Those receipts establish consistency, not package-publisher or reviewer authentication. The assurance gates are independent: - source qualification; - preparation and sample perimeter; - exact reconciliation closure; - professional semantic review; - reporting integrity; - publication, which remains outside this component. Unmatched rows or residuals set reconciliation to `withheld`, even if a reviewer accepts every review item. Reporting becomes passed only after source, preparation, reconciliation, and semantic review are passed and artifact receipts validate. Authorized review edits are resealed; an unexpected changed output keeps its old failing receipt and blocks readiness. Blocked runs still write the available receipts, qualifications, lineage, gates, audit, and narrow PDF non-movement evidence before returning an error. ## Codex Review Boundary Codex may: - decide whether a generated mapping is credible; - ask a targeted mapping question; - inspect source rows; - explain why unmatched rows need manual review; - propose deterministic improvements. Codex must not: - alter a match solely because it "looks right"; - hide unresolved ambiguity; - make direct OpenAI API calls from helper scripts; - ask the user to operate CLI scripts directly.
SHA-256: 2b11e97bcecb31f82d2d6ba974a67f0d9a0a5f36f763a832910dfffc2db5b937