{"id":23734,"plugin_id":"plugins_6ab2f25e4928819184294ebadcbe38ab","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:17:46.100Z","digest":"2c391031cf7d91e6a2e2cfe882293c808c5ec643a408dbaf5c685ecb80d70d8f","against":null,"payload":{"name":"analyzing-data","description":"Queries the data warehouse with SQL and answers business questions about data. Use when answering anything that needs warehouse data - counts, metrics, trends, aggregations, joins across tables, data lookups, or ad-hoc SQL analysis (for example \"who uses X\", \"how many Y\", \"show me Z\", \"find customers\", \"what is the count\").","included_files":[{"relative_path":"reference/common-patterns.md","size_in_bytes":1694},{"relative_path":"reference/discovery-warehouse.md","size_in_bytes":4338},{"relative_path":"scripts/.gitignore","size_in_bytes":8},{"relative_path":"scripts/cache.py","size_in_bytes":9970},{"relative_path":"scripts/cli.py","size_in_bytes":14449},{"relative_path":"scripts/config.py","size_in_bytes":1961},{"relative_path":"scripts/connectors.py","size_in_bytes":30764},{"relative_path":"scripts/kernel.py","size_in_bytes":15731},{"relative_path":"scripts/pyproject.toml","size_in_bytes":618},{"relative_path":"scripts/templates.py","size_in_bytes":4269},{"relative_path":"scripts/tests/__init__.py","size_in_bytes":45},{"relative_path":"scripts/tests/conftest.py","size_in_bytes":240},{"relative_path":"scripts/tests/integration/__init__.py","size_in_bytes":49},{"relative_path":"scripts/tests/integration/conftest.py","size_in_bytes":1451},{"relative_path":"scripts/tests/integration/test_duckdb_e2e.py","size_in_bytes":3701},{"relative_path":"scripts/tests/integration/test_kernel_interrupt.py","size_in_bytes":2273},{"relative_path":"scripts/tests/integration/test_postgres_e2e.py","size_in_bytes":3828},{"relative_path":"scripts/tests/integration/test_sqlite_e2e.py","size_in_bytes":4529},{"relative_path":"scripts/tests/test_cache.py","size_in_bytes":7899},{"relative_path":"scripts/tests/test_config.py","size_in_bytes":5396},{"relative_path":"scripts/tests/test_connectors.py","size_in_bytes":43843},{"relative_path":"scripts/tests/test_kernel.py","size_in_bytes":7045},{"relative_path":"scripts/tests/test_utils.py","size_in_bytes":2425},{"relative_path":"scripts/tests/test_warehouse.py","size_in_bytes":4350},{"relative_path":"scripts/ty.toml","size_in_bytes":644},{"relative_path":"scripts/warehouse.py","size_in_bytes":1575}],"skill_md_contents":"---\nname: analyzing-data\ndescription: Queries the data warehouse with SQL and answers business questions about data. Use when answering anything that needs warehouse data - counts, metrics, trends, aggregations, joins across tables, data lookups, or ad-hoc SQL analysis (for example \"who uses X\", \"how many Y\", \"show me Z\", \"find customers\", \"what is the count\").\n---\n\n# Data Analysis\n\nAnswer business questions by querying the data warehouse. The kernel auto-starts on first `exec` call.\n\n**All CLI commands below are relative to this skill's directory.** Before running any `scripts/cli.py` command, `cd` to the directory containing this file.\n\n## Workflow\n\n1. **Pattern lookup** — Check for a cached query strategy:\n   ```bash\n   uv run scripts/cli.py pattern lookup \"<user's question>\"\n   ```\n   If a pattern exists, follow its strategy. Record the outcome after executing:\n   ```bash\n   uv run scripts/cli.py pattern record <name> --success  # or --failure\n   ```\n\n2. **Concept lookup** — Find known table mappings:\n   ```bash\n   uv run scripts/cli.py concept lookup <concept>\n   ```\n\n3. **Table discovery** — If cache misses, search the codebase (`Grep pattern=\"<concept>\" glob=\"**/*.sql\"`) or query `INFORMATION_SCHEMA`. See [reference/discovery-warehouse.md](reference/discovery-warehouse.md).\n\n4. **Execute query**:\n   ```bash\n   uv run scripts/cli.py exec \"df = run_sql('SELECT ...')\"\n   uv run scripts/cli.py exec \"print(df)\"\n   ```\n\n5. **Cache learnings** — Always cache before presenting results:\n   ```bash\n   # Cache concept → table mapping\n   uv run scripts/cli.py concept learn <concept> <TABLE> -k <KEY_COL>\n   # Cache query strategy (if discovery was needed)\n   uv run scripts/cli.py pattern learn <name> -q \"question\" -s \"step\" -t \"TABLE\" -g \"gotcha\"\n   ```\n\n6. **Present findings** to user.\n\n## Kernel Functions\n\n| Function | Returns |\n|----------|---------|\n| `run_sql(query, limit=100)` | Polars DataFrame |\n| `run_sql_pandas(query, limit=100)` | Pandas DataFrame |\n| `run_sql_many(queries, limit=100)` | List of Polars DataFrames (one per query) |\n\n`pl` (Polars) and `pd` (Pandas) are pre-imported.\n\n**Run independent queries together** with `run_sql_many` — they execute concurrently (Snowflake async / connection-pool fan-out) instead of one at a time:\n\n```bash\nuv run scripts/cli.py exec \"dfs = run_sql_many(['SELECT ...', 'SELECT ...']); print(dfs[0])\"\n```\n\n`run_sql_many` is **fail-fast**: if any query errors, the call raises and the results of the queries that succeeded are discarded. Use separate `run_sql` calls if you need partial results.\n\n**Timeouts:** `exec` waits up to 120s by default, then interrupts the query and returns a \"client stopped waiting\" message (the query may still finish server-side). Raise it for known long-running queries: `uv run scripts/cli.py exec \"...\" -t 600`.\n\n**Idle kernel:** the kernel self-terminates after 2h idle (preserving state until then). Override with `ASTRO_KERNEL_IDLE_TIMEOUT` (seconds; `0` disables).\n\n## CLI Reference\n\n### Kernel\n\n```bash\nuv run scripts/cli.py warehouse list      # List warehouses\nuv run scripts/cli.py start [-w name]     # Start kernel (with optional warehouse)\nuv run scripts/cli.py exec \"...\"          # Execute Python code\nuv run scripts/cli.py status              # Kernel status\nuv run scripts/cli.py restart             # Restart kernel\nuv run scripts/cli.py stop                # Stop kernel\nuv run scripts/cli.py install <pkg>       # Install package\n```\n\n### Concept Cache\n\n```bash\nuv run scripts/cli.py concept lookup <name>                     # Look up\nuv run scripts/cli.py concept learn <name> <TABLE> -k <KEY_COL> # Learn\nuv run scripts/cli.py concept list                               # List all\nuv run scripts/cli.py concept import -p /path/to/warehouse.md   # Bulk import\n```\n\n### Pattern Cache\n\n```bash\nuv run scripts/cli.py pattern lookup \"question\"                                      # Look up\nuv run scripts/cli.py pattern learn <name> -q \"...\" -s \"...\" -t \"TABLE\" -g \"gotcha\"  # Learn\nuv run scripts/cli.py pattern record <name> --success                                # Record outcome\nuv run scripts/cli.py pattern list                                                   # List all\nuv run scripts/cli.py pattern delete <name>                                          # Delete\n```\n\n### Table Schema Cache\n\n```bash\nuv run scripts/cli.py table lookup <TABLE>            # Look up schema\nuv run scripts/cli.py table cache <TABLE> -c '[...]'  # Cache schema\nuv run scripts/cli.py table list                       # List cached\nuv run scripts/cli.py table delete <TABLE>             # Delete\n```\n\n### Cache Management\n\n```bash\nuv run scripts/cli.py cache status                # Stats\nuv run scripts/cli.py cache clear [--stale-only]  # Clear\n```\n\n## References\n\n- [reference/discovery-warehouse.md](reference/discovery-warehouse.md) — Large table handling, warehouse exploration, INFORMATION_SCHEMA queries\n- [reference/common-patterns.md](reference/common-patterns.md) — SQL templates for trends, comparisons, top-N, distributions, cohorts\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}