{"id":23786,"plugin_id":"plugins_6ab2f25e4928819184294ebadcbe38ab","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:17:47.525Z","digest":"2622d7a8b5e589db74d4902453b4b446ce011b3c2c38f704213b172abb9fe128","against":null,"payload":{"name":"troubleshooting-astro-deployments","description":"Troubleshoot Astronomer production deployments with Astro CLI. Use when investigating deployment issues, viewing production logs, analyzing failures, or managing deployment environment variables.","included_files":[],"skill_md_contents":"---\nname: troubleshooting-astro-deployments\ndescription: Troubleshoot Astronomer production deployments with Astro CLI. Use when investigating deployment issues, viewing production logs, analyzing failures, or managing deployment environment variables.\n---\n\n# Astro Deployment Troubleshooting\n\nThis skill helps you diagnose and troubleshoot production Astronomer deployments using the Astro CLI.\n\n> **For deployment management**, see the **managing-astro-deployments** skill.\n> **For local development**, see the **managing-astro-local-env** skill.\n\n---\n\n## Quick Health Check\n\nStart with these commands to get an overview:\n\n```bash\n# 1. List deployments to find target\nastro deployment list\n\n# 2. Get deployment overview\nastro deployment inspect <DEPLOYMENT_ID>\n\n# 3. Check for errors\nastro deployment logs <DEPLOYMENT_ID> --error -c 50\n```\n\n---\n\n## Viewing Deployment Logs\n\nUse `-c` to control log count (default: 500). Log flags cannot be combined — use one component or level flag per command.\n\n### Component-Specific Logs\n\nView logs from specific Airflow components:\n\n```bash\n# Scheduler logs (DAG processing, task scheduling)\nastro deployment logs <DEPLOYMENT_ID> --scheduler -c 50\n\n# Worker logs (task execution)\nastro deployment logs <DEPLOYMENT_ID> --workers -c 30\n\n# Webserver logs (UI access, health checks)\nastro deployment logs <DEPLOYMENT_ID> --webserver -c 30\n\n# Triggerer logs (deferrable operators)\nastro deployment logs <DEPLOYMENT_ID> --triggerer -c 30\n```\n\n### Log Level Filtering\n\nFilter by severity:\n\n```bash\n# Error logs only (most useful for troubleshooting)\nastro deployment logs <DEPLOYMENT_ID> --error -c 30\n\n# Warning logs\nastro deployment logs <DEPLOYMENT_ID> --warn -c 50\n\n# Info-level logs\nastro deployment logs <DEPLOYMENT_ID> --info -c 50\n```\n\n### Search Logs\n\nSearch for specific keywords:\n\n```bash\n# Search for specific error\nastro deployment logs <DEPLOYMENT_ID> --keyword \"ConnectionError\"\n\n# Search for specific DAG\nastro deployment logs <DEPLOYMENT_ID> --keyword \"my_dag_name\" -c 100\n\n# Find import errors\nastro deployment logs <DEPLOYMENT_ID> --error --keyword \"ImportError\"\n\n# Find task failures\nastro deployment logs <DEPLOYMENT_ID> --error --keyword \"Task failed\"\n```\n\n---\n\n## Complete Investigation Workflow\n\n### Step 1: Identify the Problem\n\n```bash\n# List deployments with status\nastro deployment list\n\n# Get deployment details\nastro deployment inspect <DEPLOYMENT_ID>\n```\n\nLook for:\n- Status: HEALTHY vs UNHEALTHY\n- Runtime version compatibility\n- Resource limits (CPU, memory)\n- Recent deployment timestamp\n\n### Step 2: Check Error Logs\n\n```bash\n# Start with errors\nastro deployment logs <DEPLOYMENT_ID> --error -c 50\n```\n\nLook for:\n- Recurring error patterns\n- Specific DAGs failing repeatedly\n- Import errors or syntax errors\n- Connection or credential errors\n\n### Step 3: Review Scheduler Logs\n\n```bash\n# Check DAG processing\nastro deployment logs <DEPLOYMENT_ID> --scheduler -c 30\n```\n\nLook for:\n- DAG parse errors\n- Scheduling delays\n- Task queueing issues\n\n### Step 4: Check Worker Logs\n\n```bash\n# Check task execution\nastro deployment logs <DEPLOYMENT_ID> --workers -c 30\n```\n\nLook for:\n- Task execution failures\n- Resource exhaustion\n- Timeout errors\n\n### Step 5: Verify Configuration\n\n```bash\n# Check environment variables\nastro deployment variable list --deployment-id <DEPLOYMENT_ID>\n\n# Verify deployment settings\nastro deployment inspect <DEPLOYMENT_ID>\n```\n\nLook for:\n- Missing or incorrect environment variables\n- Secrets configuration (AIRFLOW__SECRETS__BACKEND)\n- Connection configuration\n\n---\n\n## Common Investigation Patterns\n\n### Recurring DAG Failures\n\nFollow the complete investigation workflow above, then narrow to the specific DAG:\n\n```bash\nastro deployment logs <DEPLOYMENT_ID> --keyword \"my_dag_name\" -c 100\n```\n\n### Resource Issues\n\n```bash\n# 1. Check deployment resource allocation\nastro deployment inspect <DEPLOYMENT_ID>\n# Look for: resource_quota_cpu, resource_quota_memory\n# Worker queue: max_worker_count, worker_type\n\n# 2. Check for worker scaling issues\nastro deployment logs <DEPLOYMENT_ID> --workers -c 50\n\n# 3. Look for out-of-memory errors\nastro deployment logs <DEPLOYMENT_ID> --error --keyword \"memory\"\n```\n\n### Configuration Problems\n\n```bash\n# 1. Review environment variables\nastro deployment variable list --deployment-id <DEPLOYMENT_ID>\n\n# 2. Check for secrets backend configuration\n# Look for: AIRFLOW__SECRETS__BACKEND, AIRFLOW__SECRETS__BACKEND_KWARGS\n\n# 3. Verify deployment settings\nastro deployment inspect <DEPLOYMENT_ID>\n\n# 4. Check webserver logs for auth issues\nastro deployment logs <DEPLOYMENT_ID> --webserver -c 30\n```\n\n### Import Errors\n\n```bash\n# 1. Find import errors\nastro deployment logs <DEPLOYMENT_ID> --error --keyword \"ImportError\"\n\n# 2. Check scheduler for parse failures\nastro deployment logs <DEPLOYMENT_ID> --scheduler --keyword \"Failed to import\" -c 50\n\n# 3. Verify dependencies were deployed\nastro deployment inspect <DEPLOYMENT_ID>\n# Check: current_tag, last deployment timestamp\n```\n\n---\n\n## Environment Variables Management\n\n### List Variables\n\n```bash\n# List all variables for deployment\nastro deployment variable list --deployment-id <DEPLOYMENT_ID>\n\n# Find specific variable\nastro deployment variable list --deployment-id <DEPLOYMENT_ID> --key AWS_REGION\n\n# Export variables to file\nastro deployment variable list --deployment-id <DEPLOYMENT_ID> --save --env .env.backup\n```\n\n### Create Variables\n\n```bash\n# Create regular variable\nastro deployment variable create --deployment-id <DEPLOYMENT_ID> \\\n  --key API_ENDPOINT \\\n  --value https://api.example.com\n\n# Create secret (masked in UI and logs)\nastro deployment variable create --deployment-id <DEPLOYMENT_ID> \\\n  --key API_KEY \\\n  --value secret123 \\\n  --secret\n```\n\n### Update Variables\n\n```bash\n# Update existing variable\nastro deployment variable update --deployment-id <DEPLOYMENT_ID> \\\n  --key API_KEY \\\n  --value newsecret\n```\n\n### Delete Variables\n\n```bash\n# Delete variable\nastro deployment variable delete --deployment-id <DEPLOYMENT_ID> --key OLD_KEY\n```\n\n**Note**: Variables are available to DAGs as environment variables. Changes require no redeployment.\n\n---\n\n## Key Metrics from `deployment inspect`\n\nFocus on these fields when troubleshooting:\n\n- **status**: HEALTHY vs UNHEALTHY\n- **runtime_version**: Airflow version compatibility\n- **scheduler_size/scheduler_count**: Scheduler capacity\n- **executor**: CELERY, KUBERNETES, or LOCAL\n- **worker_queues**: Worker scaling limits and types\n  - `min_worker_count`, `max_worker_count`\n  - `worker_concurrency`\n  - `worker_type` (resource class)\n- **resource_quota_cpu/memory**: Overall resource limits\n- **dag_deploy_enabled**: Whether DAG-only deploys work\n- **current_tag**: Last deployment version\n- **is_high_availability**: Redundancy enabled\n\n---\n\n## Investigation Best Practices\n\n1. **Always start with error logs** - Most obvious failures appear here\n2. **Check error logs for patterns** - Same DAG failing repeatedly? Timing patterns?\n3. **Component-specific troubleshooting**:\n   - Worker logs → task execution details\n   - Scheduler logs → DAG processing and scheduling\n   - Webserver logs → UI issues and health checks\n   - Triggerer logs → deferrable operator issues\n4. **Use `--keyword` for targeted searches** - More efficient than reading all logs\n5. **The `inspect` command is your health dashboard** - Check it first\n6. **Environment variables in `inspect` output** - May reveal configuration issues\n7. **Log count default is 500** - Adjust with `-c` based on needs\n8. **Don't forget to check deployment time** - Recent deploy might have introduced issue\n\n---\n\n## Troubleshooting Quick Reference\n\n| Symptom | Command |\n|---------|---------|\n| Deployment shows UNHEALTHY | `astro deployment inspect <ID>` + `--error` logs |\n| DAG not appearing | `--error` logs for import errors, check `--scheduler` logs |\n| Tasks failing | `--workers` logs + search for DAG with `--keyword` |\n| Slow scheduling | `--scheduler` logs + check `inspect` for scheduler resources |\n| UI not responding | `--webserver` logs |\n| Connection issues | Check variables, search logs for connection name |\n| Import errors | `--error --keyword \"ImportError\"` + `--scheduler` logs |\n| Out of memory | `inspect` for resources + `--workers --keyword \"memory\"` |\n\n---\n\n## Related Skills\n\n- **managing-astro-deployments**: Create, update, delete deployments, deploy code\n- **managing-astro-local-env**: Manage local Airflow development environment\n- **setting-up-astro-project**: Initialize and configure Astro projects\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}