← astronomer-dataCONTENT HISTORY

Update to astronomer-data

Snapshot Sep 30, 2026 · 23:17 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "analyzing-data",
  "description": "Queries the data warehouse with SQL and answers business questions about data. Use when answering anything that needs warehouse data - counts, metrics, trends, aggregations, joins across tables, data lookups, or ad-hoc SQL analysis (for example \"who uses X\", \"how many Y\", \"show me Z\", \"find customers\", \"what is the count\").",
  "included_files": [
    {
      "relative_path": "reference/common-patterns.md",
      "size_in_bytes": 1694
    },
    {
      "relative_path": "reference/discovery-warehouse.md",
      "size_in_bytes": 4338
    },
    {
      "relative_path": "scripts/.gitignore",
      "size_in_bytes": 8
    },
    {
      "relative_path": "scripts/cache.py",
      "size_in_bytes": 9970
    },
    {
      "relative_path": "scripts/cli.py",
      "size_in_bytes": 14449
    },
    {
      "relative_path": "scripts/config.py",
      "size_in_bytes": 1961
    },
    {
      "relative_path": "scripts/connectors.py",
      "size_in_bytes": 30764
    },
    {
      "relative_path": "scripts/kernel.py",
      "size_in_bytes": 15731
    },
    {
      "relative_path": "scripts/pyproject.toml",
      "size_in_bytes": 618
    },
    {
      "relative_path": "scripts/templates.py",
      "size_in_bytes": 4269
    },
    {
      "relative_path": "scripts/tests/__init__.py",
      "size_in_bytes": 45
    },
    {
      "relative_path": "scripts/tests/conftest.py",
      "size_in_bytes": 240
    },
    {
      "relative_path": "scripts/tests/integration/__init__.py",
      "size_in_bytes": 49
    },
    {
      "relative_path": "scripts/tests/integration/conftest.py",
      "size_in_bytes": 1451
    },
    {
      "relative_path": "scripts/tests/integration/test_duckdb_e2e.py",
      "size_in_bytes": 3701
    },
    {
      "relative_path": "scripts/tests/integration/test_kernel_interrupt.py",
      "size_in_bytes": 2273
    },
    {
      "relative_path": "scripts/tests/integration/test_postgres_e2e.py",
      "size_in_bytes": 3828
    },
    {
      "relative_path": "scripts/tests/integration/test_sqlite_e2e.py",
      "size_in_bytes": 4529
    },
    {
      "relative_path": "scripts/tests/test_cache.py",
      "size_in_bytes": 7899
    },
    {
      "relative_path": "scripts/tests/test_config.py",
      "size_in_bytes": 5396
    },
    {
      "relative_path": "scripts/tests/test_connectors.py",
      "size_in_bytes": 43843
    },
    {
      "relative_path": "scripts/tests/test_kernel.py",
      "size_in_bytes": 7045
    },
    {
      "relative_path": "scripts/tests/test_utils.py",
      "size_in_bytes": 2425
    },
    {
      "relative_path": "scripts/tests/test_warehouse.py",
      "size_in_bytes": 4350
    },
    {
      "relative_path": "scripts/ty.toml",
      "size_in_bytes": 644
    },
    {
      "relative_path": "scripts/warehouse.py",
      "size_in_bytes": 1575
    }
  ],
  "skill_md_contents": "---\nname: analyzing-data\ndescription: Queries the data warehouse with SQL and answers business questions about data. Use when answering anything that needs warehouse data - counts, metrics, trends, aggregations, joins across tables, data lookups, or ad-hoc SQL analysis (for example \"who uses X\", \"how many Y\", \"show me Z\", \"find customers\", \"what is the count\").\n---\n\n# Data Analysis\n\nAnswer business questions by querying the data warehouse. The kernel auto-starts on first `exec` call.\n\n**All CLI commands below are relative to this skill's directory.** Before running any `scripts/cli.py` command, `cd` to the directory containing this file.\n\n## Workflow\n\n1. **Pattern lookup** — Check for a cached query strategy:\n   ```bash\n   uv run scripts/cli.py pattern lookup \"<user's question>\"\n   ```\n   If a pattern exists, follow its strategy. Record the outcome after executing:\n   ```bash\n   uv run scripts/cli.py pattern record <name> --success  # or --failure\n   ```\n\n2. **Concept lookup** — Find known table mappings:\n   ```bash\n   uv run scripts/cli.py concept lookup <concept>\n   ```\n\n3. **Table discovery** — If cache misses, search the codebase (`Grep pattern=\"<concept>\" glob=\"**/*.sql\"`) or query `INFORMATION_SCHEMA`. See [reference/discovery-warehouse.md](reference/discovery-warehouse.md).\n\n4. **Execute query**:\n   ```bash\n   uv run scripts/cli.py exec \"df = run_sql('SELECT ...')\"\n   uv run scripts/cli.py exec \"print(df)\"\n   ```\n\n5. **Cache learnings** — Always cache before presenting results:\n   ```bash\n   # Cache concept → table mapping\n   uv run scripts/cli.py concept learn <concept> <TABLE> -k <KEY_COL>\n   # Cache query strategy (if discovery was needed)\n   uv run scripts/cli.py pattern learn <name> -q \"question\" -s \"step\" -t \"TABLE\" -g \"gotcha\"\n   ```\n\n6. **Present findings** to user.\n\n## Kernel Functions\n\n| Function | Returns |\n|----------|---------|\n| `run_sql(query, limit=100)` | Polars DataFrame |\n| `run_sql_pandas(query, limit=100)` | Pandas DataFrame |\n| `run_sql_many(queries, limit=100)` | List of Polars DataFrames (one per query) |\n\n`pl` (Polars) and `pd` (Pandas) are pre-imported.\n\n**Run independent queries together** with `run_sql_many` — they execute concurrently (Snowflake async / connection-pool fan-out) instead of one at a time:\n\n```bash\nuv run scripts/cli.py exec \"dfs = run_sql_many(['SELECT ...', 'SELECT ...']); print(dfs[0])\"\n```\n\n`run_sql_many` is **fail-fast**: if any query errors, the call raises and the results of the queries that succeeded are discarded. Use separate `run_sql` calls if you need partial results.\n\n**Timeouts:** `exec` waits up to 120s by default, then interrupts the query and returns a \"client stopped waiting\" message (the query may still finish server-side). Raise it for known long-running queries: `uv run scripts/cli.py exec \"...\" -t 600`.\n\n**Idle kernel:** the kernel self-terminates after 2h idle (preserving state until then). Override with `ASTRO_KERNEL_IDLE_TIMEOUT` (seconds; `0` disables).\n\n## CLI Reference\n\n### Kernel\n\n```bash\nuv run scripts/cli.py warehouse list      # List warehouses\nuv run scripts/cli.py start [-w name]     # Start kernel (with optional warehouse)\nuv run scripts/cli.py exec \"...\"          # Execute Python code\nuv run scripts/cli.py status              # Kernel status\nuv run scripts/cli.py restart             # Restart kernel\nuv run scripts/cli.py stop                # Stop kernel\nuv run scripts/cli.py install <pkg>       # Install package\n```\n\n### Concept Cache\n\n```bash\nuv run scripts/cli.py concept lookup <name>                     # Look up\nuv run scripts/cli.py concept learn <name> <TABLE> -k <KEY_COL> # Learn\nuv run scripts/cli.py concept list                               # List all\nuv run scripts/cli.py concept import -p /path/to/warehouse.md   # Bulk import\n```\n\n### Pattern Cache\n\n```bash\nuv run scripts/cli.py pattern lookup \"question\"                                      # Look up\nuv run scripts/cli.py pattern learn <name> -q \"...\" -s \"...\" -t \"TABLE\" -g \"gotcha\"  # Learn\nuv run scripts/cli.py pattern record <name> --success                                # Record outcome\nuv run scripts/cli.py pattern list                                                   # List all\nuv run scripts/cli.py pattern delete <name>                                          # Delete\n```\n\n### Table Schema Cache\n\n```bash\nuv run scripts/cli.py table lookup <TABLE>            # Look up schema\nuv run scripts/cli.py table cache <TABLE> -c '[...]'  # Cache schema\nuv run scripts/cli.py table list                       # List cached\nuv run scripts/cli.py table delete <TABLE>             # Delete\n```\n\n### Cache Management\n\n```bash\nuv run scripts/cli.py cache status                # Stats\nuv run scripts/cli.py cache clear [--stale-only]  # Clear\n```\n\n## References\n\n- [reference/discovery-warehouse.md](reference/discovery-warehouse.md) — Large table handling, warehouse exploration, INFORMATION_SCHEMA queries\n- [reference/common-patterns.md](reference/common-patterns.md) — SQL templates for trends, comparisons, top-N, distributions, cohorts\n"
}

SHA-256: 2c391031cf7d91e6a2e2cfe882293c808c5ec643a408dbaf5c685ecb80d70d8f