← astronomer-dataCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to astronomer-data
Snapshot Sep 30, 2026 · 23:17 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "debugging-dags",
"description": "Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations. Use when deep failure investigation is needed, a DAG fails to import/parse or 'airflow dags list' errors on a file; a task or run is failing and must be diagnosed and fixed; requests like 'why did X fail', 'my dag keeps failing — find and fix it', or fixing a broken DAG so it loads cleanly. For simple 'why did it fail / show logs', the airflow skill handles it directly.",
"included_files": [],
"skill_md_contents": "---\nname: debugging-dags\ndescription: Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations. Use when deep failure investigation is needed, a DAG fails to import/parse or 'airflow dags list' errors on a file; a task or run is failing and must be diagnosed and fixed; requests like 'why did X fail', 'my dag keeps failing — find and fix it', or fixing a broken DAG so it loads cleanly. For simple 'why did it fail / show logs', the airflow skill handles it directly.\n---\n\n# DAG Diagnosis\n\nYou are a data engineer debugging a failed Airflow DAG. Follow this systematic approach to identify the root cause and provide actionable remediation.\n\n## Running the CLI\n\nThese commands assume `af` is on PATH. Run via `astro otto` to get it automatically, or install standalone with `uv tool install astro-airflow-mcp`.\n\n---\n\n## Step 1: Identify the Failure\n\nIf a specific DAG was mentioned:\n- Run `af runs diagnose <dag_id> <dag_run_id>` (if run_id is provided)\n- If no run_id specified, run `af dags stats` to find recent failures\n\nIf no DAG was specified:\n- Run `af health` to find recent failures across all DAGs\n- Check for import errors with `af dags errors`\n- Show DAGs with recent failures\n- Ask which DAG to investigate further\n\n## Step 2: Get the Error Details\n\nOnce you have identified a failed task:\n\n1. **Get task logs** using `af tasks logs <dag_id> <dag_run_id> <task_id>`\n2. **Look for the actual exception** - scroll past the Airflow boilerplate to find the real error\n3. **Categorize the failure type**:\n - **Data issue**: Missing data, schema change, null values, constraint violation\n - **Code issue**: Bug, syntax error, import failure, type error\n - **Infrastructure issue**: Connection timeout, resource exhaustion, permission denied\n - **Dependency issue**: Upstream failure, external API down, rate limiting\n\n## Step 3: Check Context\n\nGather additional context to understand WHY this happened:\n\n1. **Recent changes**: Was there a code deploy? Check git history if available\n2. **Package version changes**: Was a package upgraded — in the image, in a venv-style operator, or at the index? See [Package version changes](#package-version-changes) below.\n3. **Data volume**: Did data volume spike? Run a quick count on source tables\n4. **Upstream health**: Did upstream tasks succeed but produce unexpected data?\n5. **Historical pattern**: Is this a recurring failure? Check if same task failed before\n6. **Timing**: Did this fail at an unusual time? (resource contention, maintenance windows)\n\nUse `af runs get <dag_id> <dag_run_id>` to compare the failed run against recent successful runs.\n\n### Package version changes\n\nA common cause of failures with no git activity is dependency drift — the user's code didn't change, but a package they depend on did. Check in this order:\n\n1. **Worker image diff** (preferred when available). Every Astro deploy = new image tag, so the registry has a \"before\" and \"after\". Diff `pip freeze` between current and previous image — that's ground truth for what changed:\n ```\n docker run --rm <current_image> pip freeze > /tmp/now.txt\n docker run --rm <previous_image> pip freeze > /tmp/prev.txt\n diff /tmp/prev.txt /tmp/now.txt\n ```\n Also compare `docker run --rm <image> python --version` between the two — a Python minor-version bump (3.11 → 3.12, or even a patch) can break wheel compatibility even when `pip freeze` looks identical. `af config providers` lists currently installed provider versions, useful for cross-checking against modules named in the traceback.\n\n2. **Venv-style operators bypass the worker image.** `@task.virtualenv`, `PythonVirtualenvOperator`, `ExternalPythonOperator`, and `KubernetesPodOperator` build their environment per task run, so an image diff won't catch failures inside them. If the failed task is one of these, read its `requirements` / `image` / `python_version` / `python` args directly:\n - Unbounded specifier (e.g. `pandas>=2.0.0` with no upper bound, or no specifier at all) → a new upstream release is the prime suspect.\n - `image=\"foo:latest\"` or no tag → the image moved underneath you.\n - `python_version=\"3.11\"` (on `@task.virtualenv` / `PythonVirtualenvOperator`) or a `python` path (on `ExternalPythonOperator`) resolving to a different interpreter than it used to — a Python minor-version change can break wheel compatibility for unchanged `requirements`. Same vector applies to the worker image itself if the base Python changed there.\n\n Fix is to pin: `pandas>=2.0.0,<3.0.0`, a lockfile, a specific image SHA, or a fully-qualified Python version (`python_version=\"3.11.7\"` instead of `\"3.11\"`).\n\n3. **Index lookup** when image diff isn't conclusive (no image history, or a venv-style operator). Identify the configured index first — it may not be PyPI:\n - Env vars: `UV_INDEX_URL`, `PIP_INDEX_URL`, `PIP_EXTRA_INDEX_URL`\n - `pyproject.toml` → `[[tool.uv.index]]`\n - `~/.pip/pip.conf`, `/etc/pip.conf`\n - `Dockerfile` `--index-url` flags\n\n Then query for releases of the suspect package since the first failure started. PyPI:\n ```\n curl -s https://pypi.org/pypi/<pkg>/json | jq '.releases | to_entries | map({version: .key, uploaded: .value[0].upload_time}) | sort_by(.uploaded) | reverse | .[:5]'\n ```\n Private indexes usually expose the same `/pypi/<pkg>/json` shape; fall back to the Simple API (`/simple/<pkg>/`) or ask the user if neither works.\n\nA release timestamp landing between the last green run and the first red run, for a package named in the traceback, is the answer.\n\n### On Astro\n\nIf you're running on Astro, these additional tools can help with diagnosis:\n\n- **Deployment activity log**: Check the Astro UI for recent deploys — a failed deploy or recent code change is often the cause of sudden failures\n- **Astro alerts**: Configure alerts in the Astro UI for proactive failure monitoring (DAG failure, task duration, SLA miss)\n- **Observability**: Use the Astro [observability dashboard](https://www.astronomer.io/docs/astro/airflow-alerts) to track DAG health trends and spot recurring issues\n\n### On OSS Airflow\n\n- **Airflow UI**: Use the DAGs page, Graph view, and task logs to inspect recent runs and failures\n\n## Step 4: Provide Actionable Output\n\nStructure your diagnosis as:\n\n### Root Cause\nWhat actually broke? Be specific - not \"the task failed\" but \"the task failed because column X was null in 15% of rows when the code expected 0%\".\n\n### Impact Assessment\n- What data is affected? Which tables didn't get updated?\n- What downstream processes are blocked?\n- Is this blocking production dashboards or reports?\n\n### Immediate Fix\nSpecific steps to resolve RIGHT NOW:\n1. If it's a data issue: SQL to fix or skip bad records\n2. If it's a code issue: The exact code change needed\n3. If it's infra: Who to contact or what to restart\n\n### Prevention\nHow to prevent this from happening again:\n- Add data quality checks?\n- Add better error handling?\n- Add alerting for edge cases?\n- Update documentation?\n- Pin dependencies (constraints file, lockfile, or upper-bound specifiers on venv/external/pod operators) to avoid silent upstream drift?\n\n### Quick Commands\nProvide ready-to-use commands:\n- To clear and rerun the entire DAG run: `af runs clear <dag_id> <run_id>`\n- To clear and rerun specific failed tasks: `af tasks clear <dag_id> <run_id> <task_ids> -D`\n- To delete a stuck or unwanted run: `af runs delete <dag_id> <run_id>`\n"
}SHA-256: 475e7fb92fe6d5a61229c211c8864902c5c9c86ef6eb863434068bc115fef804