← astronomer-dataCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to astronomer-data
Snapshot Sep 30, 2026 · 23:17 UTC · version 0.1.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "analyzing-data",
"description": "Queries the data warehouse with SQL and answers business questions about data. Use when answering anything that needs warehouse data - counts, metrics, trends, aggregations, joins across tables, data lookups, or ad-hoc SQL analysis (for example \"who uses X\", \"how many Y\", \"show me Z\", \"find customers\", \"what is the count\").",
"included_files": [
{
"relative_path": "reference/common-patterns.md",
"size_in_bytes": 1694
},
{
"relative_path": "reference/discovery-warehouse.md",
"size_in_bytes": 4338
},
{
"relative_path": "scripts/.gitignore",
"size_in_bytes": 8
},
{
"relative_path": "scripts/cache.py",
"size_in_bytes": 9970
},
{
"relative_path": "scripts/cli.py",
"size_in_bytes": 14449
},
{
"relative_path": "scripts/config.py",
"size_in_bytes": 1961
},
{
"relative_path": "scripts/connectors.py",
"size_in_bytes": 30764
},
{
"relative_path": "scripts/kernel.py",
"size_in_bytes": 15731
},
{
"relative_path": "scripts/pyproject.toml",
"size_in_bytes": 618
},
{
"relative_path": "scripts/templates.py",
"size_in_bytes": 4269
},
{
"relative_path": "scripts/tests/__init__.py",
"size_in_bytes": 45
},
{
"relative_path": "scripts/tests/conftest.py",
"size_in_bytes": 240
},
{
"relative_path": "scripts/tests/integration/__init__.py",
"size_in_bytes": 49
},
{
"relative_path": "scripts/tests/integration/conftest.py",
"size_in_bytes": 1451
},
{
"relative_path": "scripts/tests/integration/test_duckdb_e2e.py",
"size_in_bytes": 3701
},
{
"relative_path": "scripts/tests/integration/test_kernel_interrupt.py",
"size_in_bytes": 2273
},
{
"relative_path": "scripts/tests/integration/test_postgres_e2e.py",
"size_in_bytes": 3828
},
{
"relative_path": "scripts/tests/integration/test_sqlite_e2e.py",
"size_in_bytes": 4529
},
{
"relative_path": "scripts/tests/test_cache.py",
"size_in_bytes": 7899
},
{
"relative_path": "scripts/tests/test_config.py",
"size_in_bytes": 5396
},
{
"relative_path": "scripts/tests/test_connectors.py",
"size_in_bytes": 43843
},
{
"relative_path": "scripts/tests/test_kernel.py",
"size_in_bytes": 7045
},
{
"relative_path": "scripts/tests/test_utils.py",
"size_in_bytes": 2425
},
{
"relative_path": "scripts/tests/test_warehouse.py",
"size_in_bytes": 4350
},
{
"relative_path": "scripts/ty.toml",
"size_in_bytes": 644
},
{
"relative_path": "scripts/warehouse.py",
"size_in_bytes": 1575
}
],
"skill_md_contents": "---\nname: analyzing-data\ndescription: Queries the data warehouse with SQL and answers business questions about data. Use when answering anything that needs warehouse data - counts, metrics, trends, aggregations, joins across tables, data lookups, or ad-hoc SQL analysis (for example \"who uses X\", \"how many Y\", \"show me Z\", \"find customers\", \"what is the count\").\n---\n\n# Data Analysis\n\nAnswer business questions by querying the data warehouse. The kernel auto-starts on first `exec` call.\n\n**All CLI commands below are relative to this skill's directory.** Before running any `scripts/cli.py` command, `cd` to the directory containing this file.\n\n## Workflow\n\n1. **Pattern lookup** — Check for a cached query strategy:\n ```bash\n uv run scripts/cli.py pattern lookup \"<user's question>\"\n ```\n If a pattern exists, follow its strategy. Record the outcome after executing:\n ```bash\n uv run scripts/cli.py pattern record <name> --success # or --failure\n ```\n\n2. **Concept lookup** — Find known table mappings:\n ```bash\n uv run scripts/cli.py concept lookup <concept>\n ```\n\n3. **Table discovery** — If cache misses, search the codebase (`Grep pattern=\"<concept>\" glob=\"**/*.sql\"`) or query `INFORMATION_SCHEMA`. See [reference/discovery-warehouse.md](reference/discovery-warehouse.md).\n\n4. **Execute query**:\n ```bash\n uv run scripts/cli.py exec \"df = run_sql('SELECT ...')\"\n uv run scripts/cli.py exec \"print(df)\"\n ```\n\n5. **Cache learnings** — Always cache before presenting results:\n ```bash\n # Cache concept → table mapping\n uv run scripts/cli.py concept learn <concept> <TABLE> -k <KEY_COL>\n # Cache query strategy (if discovery was needed)\n uv run scripts/cli.py pattern learn <name> -q \"question\" -s \"step\" -t \"TABLE\" -g \"gotcha\"\n ```\n\n6. **Present findings** to user.\n\n## Kernel Functions\n\n| Function | Returns |\n|----------|---------|\n| `run_sql(query, limit=100)` | Polars DataFrame |\n| `run_sql_pandas(query, limit=100)` | Pandas DataFrame |\n| `run_sql_many(queries, limit=100)` | List of Polars DataFrames (one per query) |\n\n`pl` (Polars) and `pd` (Pandas) are pre-imported.\n\n**Run independent queries together** with `run_sql_many` — they execute concurrently (Snowflake async / connection-pool fan-out) instead of one at a time:\n\n```bash\nuv run scripts/cli.py exec \"dfs = run_sql_many(['SELECT ...', 'SELECT ...']); print(dfs[0])\"\n```\n\n`run_sql_many` is **fail-fast**: if any query errors, the call raises and the results of the queries that succeeded are discarded. Use separate `run_sql` calls if you need partial results.\n\n**Timeouts:** `exec` waits up to 120s by default, then interrupts the query and returns a \"client stopped waiting\" message (the query may still finish server-side). Raise it for known long-running queries: `uv run scripts/cli.py exec \"...\" -t 600`.\n\n**Idle kernel:** the kernel self-terminates after 2h idle (preserving state until then). Override with `ASTRO_KERNEL_IDLE_TIMEOUT` (seconds; `0` disables).\n\n## CLI Reference\n\n### Kernel\n\n```bash\nuv run scripts/cli.py warehouse list # List warehouses\nuv run scripts/cli.py start [-w name] # Start kernel (with optional warehouse)\nuv run scripts/cli.py exec \"...\" # Execute Python code\nuv run scripts/cli.py status # Kernel status\nuv run scripts/cli.py restart # Restart kernel\nuv run scripts/cli.py stop # Stop kernel\nuv run scripts/cli.py install <pkg> # Install package\n```\n\n### Concept Cache\n\n```bash\nuv run scripts/cli.py concept lookup <name> # Look up\nuv run scripts/cli.py concept learn <name> <TABLE> -k <KEY_COL> # Learn\nuv run scripts/cli.py concept list # List all\nuv run scripts/cli.py concept import -p /path/to/warehouse.md # Bulk import\n```\n\n### Pattern Cache\n\n```bash\nuv run scripts/cli.py pattern lookup \"question\" # Look up\nuv run scripts/cli.py pattern learn <name> -q \"...\" -s \"...\" -t \"TABLE\" -g \"gotcha\" # Learn\nuv run scripts/cli.py pattern record <name> --success # Record outcome\nuv run scripts/cli.py pattern list # List all\nuv run scripts/cli.py pattern delete <name> # Delete\n```\n\n### Table Schema Cache\n\n```bash\nuv run scripts/cli.py table lookup <TABLE> # Look up schema\nuv run scripts/cli.py table cache <TABLE> -c '[...]' # Cache schema\nuv run scripts/cli.py table list # List cached\nuv run scripts/cli.py table delete <TABLE> # Delete\n```\n\n### Cache Management\n\n```bash\nuv run scripts/cli.py cache status # Stats\nuv run scripts/cli.py cache clear [--stale-only] # Clear\n```\n\n## References\n\n- [reference/discovery-warehouse.md](reference/discovery-warehouse.md) — Large table handling, warehouse exploration, INFORMATION_SCHEMA queries\n- [reference/common-patterns.md](reference/common-patterns.md) — SQL templates for trends, comparisons, top-N, distributions, cohorts\n"
}SHA-256: 2c391031cf7d91e6a2e2cfe882293c808c5ec643a408dbaf5c685ecb80d70d8f