← TuistCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Tuist
Snapshot Sep 30, 2026 · 23:09 UTC · version 1.0.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "compare-test-runs",
"description": "Compares two test runs to identify new failures, newly flaky tests, fixed tests, and duration regressions. Can be invoked with test run IDs, dashboard URLs, or branch names.",
"included_files": [],
"skill_md_contents": "---\nname: compare-test-runs\ndescription: Compares two test runs to identify new failures, newly flaky tests, fixed tests, and duration regressions. Can be invoked with test run IDs, dashboard URLs, or branch names.\n---\n\n# Compare Test Runs\n\n## Quick Start\n\nYou'll typically receive two test run identifiers. Follow these steps:\n\n1. Run `tuist test show <id> --json` for both base and head test runs.\n2. Run `tuist test module list <test-run-id> --json` and `tuist test suite list <test-run-id> --json` to get module and suite breakdowns.\n3. Run `tuist test case run list <identifier> --json` to get individual test case results.\n4. Compare failures, flaky tests, durations, and overall status.\n5. Inspect failing test cases with `tuist test case run show <id> --json`.\n6. Summarize findings with actionable recommendations.\n\n## Step 1: Resolve Test Runs\n\n### If base/head are test run IDs or dashboard URLs\n\nFetch each directly:\n\n```bash\ntuist test show <base-id> --json\ntuist test show <head-id> --json\n```\n\n### If base/head are branch names\n\nList recent test runs on each branch to identify test run IDs:\n\n```bash\ntuist test list --git-branch <base-branch> --json --page-size 5\ntuist test list --git-branch <head-branch> --json --page-size 5\n```\n\nPick the latest test run ID from each branch's results.\n\n### Defaults\n\n- If no base is provided, use the project's default branch (usually `main`).\n- If no head is provided, detect the current git branch.\n\n## Step 2: Compare Top-Level Metrics\n\nAfter fetching both test runs, compare:\n\n| Metric | What to check |\n|---|---|\n| `status` | Flag if base passed but head failed |\n| `duration` | Flag if head is >10% slower |\n| `total_test_count` | Note if test count changed (new or removed tests) |\n| `failed_test_count` | Compare failure counts |\n| `flaky_test_count` | Compare flaky counts |\n| `avg_test_duration` | Flag significant changes |\n\n## Step 3: Get Module and Suite Breakdowns\n\nFetch module and suite-level results for both test runs to understand which areas regressed:\n\n```bash\ntuist test module list <base-test-run-id> --json\ntuist test module list <head-test-run-id> --json\n\ntuist test suite list <base-test-run-id> --json\ntuist test suite list <head-test-run-id> --json\n```\n\nMatch modules and suites by name across both runs to identify areas with new failures or duration regressions.\n\n## Step 4: Get Individual Test Case Results\n\nFetch test case runs for both test runs:\n\n```bash\ntuist test case run list <identifier> --json --page-size 100\n```\n\nMatch test cases by their `name` + `module_name` + `suite_name` across both runs.\n\n## Step 5: Classify Changes\n\nGroup test cases into categories:\n\n1. **New failures**: Tests that passed in base but failed in head.\n2. **Fixed tests**: Tests that failed in base but passed in head.\n3. **Newly flaky**: Tests not flaky in base but flaky in head.\n4. **No longer flaky**: Tests that were flaky in base but stable in head.\n5. **New tests**: Tests present in head but not in base.\n6. **Removed tests**: Tests present in base but not in head.\n7. **Duration regressions**: Tests with >50% duration increase.\n\n## Step 6: Inspect Failures\n\nFor each new failure, get detailed information:\n\n```bash\ntuist test case run show <test-case-run-id> --json\n```\n\nKey fields to examine:\n- `failures[].message` -- the assertion or error message\n- `failures[].path` -- source file path\n- `failures[].line_number` -- exact line of failure\n- `failures[].issue_type` -- type of issue\n- `repetitions` -- if present, shows retry behavior (flaky detection)\n- `crash_report` -- crash data if test runner crashed\n\n## Step 7: Inspect Attachments\n\nThe `tuist test case run show` output includes attachment and crash report information. Review:\n- Screenshots or UI test artifacts\n- Log files or crash reports\n- Any diagnostic data attached to failing runs\n\n## Summary Format\n\nProduce a summary with:\n\n1. **Overall verdict**: Better, worse, or neutral compared to base.\n2. **New failures**: List each with failure message, file path, and line number.\n3. **New flaky tests**: List with flakiness context.\n4. **Fixed tests**: List tests that are now passing.\n5. **Duration**: Overall and notable per-test regressions.\n6. **Recommendations**: Actionable next steps for each issue.\n\nExample:\n\n```\nTest Run Comparison: base (run-123 on main) vs head (run-456 on feature-x)\n\nStatus: success -> failure -- REGRESSION\nDuration: 120.5s -> 145.2s (+21%)\nTests: 342 -> 345 (3 new tests)\nFailures: 0 -> 2 (2 new failures)\nFlaky: 1 -> 3 (2 newly flaky)\n\nNew Failures:\n1. AuthModuleTests/LoginTests/test_login_with_expired_token\n Message: \"Expected status 401, got 500\"\n File: Tests/AuthModuleTests/LoginTests.swift:42\n Likely cause: Server error handling changed for expired tokens\n\n2. NetworkTests/RetryTests/test_retry_on_timeout\n Message: \"Timed out waiting for retry\"\n File: Tests/NetworkTests/RetryTests.swift:87\n Likely cause: Timeout threshold too low after network layer refactor\n\nNewly Flaky:\n1. CacheTests/WriteCacheTests/test_concurrent_writes (flaky in 3/5 runs)\n\nRecommendations:\n- Fix expired token handling in AuthModule\n- Increase timeout in RetryTests or mock the network layer\n- Investigate concurrent write synchronization in CacheTests\n```\n\n## Done Checklist\n\n- Resolved both base and head test runs\n- Compared top-level metrics\n- Fetched module and suite breakdowns for both runs\n- Identified new failures, fixed tests, and flaky changes\n- Inspected failure details for new failures\n- Provided actionable recommendations with file paths\n"
}SHA-256: 60c1273c75e6c541b8a895dc621defde9ea3678215c761cd8026330146d64408