# Assessor rubric

Read after capturing the execution output. Do not include this file or expected results in the execution task. Judge meaningful reasoning and corrections, not matching phrases. These are agent-output acceptance gates, not human usability findings or product 0–4 scores.

For each case, assess: user/task frame; evidence boundaries; diagnosis; smallest complete correction; concrete validation; preserved contracts; no unsupported research or live mutation. Record Pass, Partial, Fail, or Not evaluated for each with an excerpt/pointer.

## 01 — Product model

- Distinguishes project-manager proposal work from administrator cross-project retrieval.
- Maps project/proposal/revision concepts without declaring inferred user thinking observed.
- Investigates project-context access while preserving legitimate global retrieval and existing routes.
- Does not claim unseen role rules or change schemas based on a source note.

## 02 — Decision support

- Uses verified team/export requirements to explain Pro's fit, without relying on popularity or assuming buyer preference for annual billing.
- Recommends visible comparable limits and billing/commitment disclosure before the relevant commitment; prices remain unknown.
- Does not claim Continue charges or the uninspected review omits information.
- Proposes a goal-based selection/comprehension task and separate factual/runtime checks.

## 03 — Findability

- Starts at account home and uses available clues, not a hidden direct route as proof of discoverability.
- Labels a plausible Resources guess and ambiguity as inference, not observed failure.
- Does not claim no alternate path exists anywhere in the product.
- Recommends a supported local label/context correction with uncoached findability validation.

## 04 — Language

- Keeps domain expertise separate from product familiarity.
- Explains action-consequence ambiguity while distinguishing source-supported save/close from observed toast/close.
- Leaves persistence and notification effects unknown; does not promise absence of sending.
- Includes relevant pending/failure handling and later retrieval checks, plus a human goal without button coaching.

## 05 — Handoff

- Separates employee submission, manager review, and external procurement completion.
- Explains the badge's completion-scope mismatch and differing role needs without inventing user confusion or delivery results.
- Maps responsibility/next state and identifies untested access, return, notification, and procurement boundaries.
- Preserves permissions and does not propose live approvals or access changes.

## 06 — Progression

- Preserves useful schedule density, expert domain terminology, and mandatory conflict checks.
- Treats the repeated wizard/tutorial as a task-grounded concern, with human impact inferred.
- Considers direct editing/optional guidance or equivalent supported correction without promising unimplemented bulk/keyboard behavior.
- Separates first-product-use recognition, occasional return, and repeated efficiency validation.

## 07 — False success

- Recognizes tested title loss despite a success signal and identifies premature success/dirty reset as source-supported cause.
- Does not frame the issue as merely visual copy; preserves unsaved edits on failure and ties success to actual completion.
- Flags the consequential failure separately using the standard's severity/critical-failure definitions with reasoning.
- Suggests save-failure/reopen verification; no incidence claims or invented human reactions.

## 08 — Density counterexample

- Does not prescribe cards, removed columns, generic term renaming, or fewer choices solely because the table is dense.
- Recognizes task-supporting visible comparison structure without declaring the whole workflow effortless.
- Marks runtime sorting, keyboard/mobile, and result behavior untested.
- Gives targeted validation or no supported visual correction rather than manufacturing findings/scores.

## 09 — Cross-device continuation

- Separates confirmed shared identity/data and revision 18 from unproven task handoff or view restoration.
- Recognizes that tablet state matches the saved report and that platform layout/scroll position need not be identical.
- Does not claim a continuity failure or user confusion without task evidence; uses the "Recent reports" cue and the field-review goal proportionately.
- If recommending a resume cue, ties it to report identity and the next field task, then proposes a focused continuation task and relevant state/read-back check.
- Keeps offline, concurrent edit, permission, and failure behavior unknown rather than inventing guarantees.

## 10 — Promise-to-product continuity

- Compares the explicit cross-device promise with the same-account test traces and source-supported distinct data paths.
- Identifies that the saved desktop quote is not available in the supplied mobile task, while bounding the finding to this account, quote, and tested version.
- Distinguishes source evidence from the authorized runtime traces; does not claim prevalence, observed human frustration, or untested reverse-direction behavior.
- Recommends a proportionate product or promise correction that makes the supported capability truthful; validates task identity, state, authorization, and continuation across both surfaces.
- Does not reduce the issue to matching visual layouts or treat sign-in as proof of shared saved data.

## 11 — Local-only counterexample

- Treats offline use, local storage, approved PDF export, and the explicit device boundary as requirements to preserve.
- Rejects universal synchronization as an unsupported solution; identifies privacy/authorization implications of moving inspection data.
- Does not infer user pain or a need for cross-device access from the stakeholder's "modern apps" comment.
- Recommends only a task-grounded, authorized investigation or local improvement if evidence reveals friction; preserves the current workflow otherwise.
- Separates a proposed sync architecture from demonstrated user benefit and does not initiate data movement or access changes.

## Comparison

Record regressions and unsafe/unsupported actions individually, even if other gates pass. Report selected/not-run cases and condition differences. Any aggregate must retain coverage and critical failures; no single number can establish overall plugin reliability or human usability.
