← LanceDBCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to LanceDB
Snapshot Sep 30, 2026 · 23:15 UTC · version 0.1.1
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Use when writing, reviewing, debugging, or documenting LanceDB pipelines in Python or TypeScript, especially code that should work across local LanceDB OSS tables and remote LanceDB Enterprise/Cloud tables. Helps avoid non-portable full-table materialization, choose idiomatic query/search patterns, apply LanceDB performance defaults for ingestion, indexing, filtering, and diagnostics, and resolve connections to the remote server for Enterprise-only operations such as jobs.",
"included_files": [
{
"relative_path": ".DS_Store",
"size_in_bytes": 6148
},
{
"relative_path": "agents/openai.yaml",
"size_in_bytes": 273
},
{
"relative_path": "assets/icon.png",
"size_in_bytes": 23274
},
{
"relative_path": "references/branch_ops.md",
"size_in_bytes": 1863
},
{
"relative_path": "references/column_metadata.md",
"size_in_bytes": 4214
},
{
"relative_path": "references/remote_connect.md",
"size_in_bytes": 2297
},
{
"relative_path": "references/remote_jobs.md",
"size_in_bytes": 3717
},
{
"relative_path": "scripts/check_materialization.py",
"size_in_bytes": 4297
}
],
"name": "lancedb",
"skill_md_contents": "---\nname: lancedb\ndescription: Use when writing, reviewing, debugging, or documenting LanceDB pipelines in Python or TypeScript, especially code that should work across local LanceDB OSS tables and remote LanceDB Enterprise/Cloud tables. Helps avoid non-portable full-table materialization, choose idiomatic query/search patterns, apply LanceDB performance defaults for ingestion, indexing, filtering, and diagnostics, and resolve connections to the remote server for Enterprise-only operations such as jobs.\n---\n\n# Building LanceDB Pipelines\n\nUse this skill to produce LanceDB pipelines that are portable between local and remote tables (for LanceDB Enterprise/Cloud) and idiomatic for the selected SDK.\n\n## LanceDB Table Modes\n\nLanceDB has two common execution modes:\n\n- **Local table**: embedded, open source, in-process LanceDB. The client opens data from a local path or object storage URI and executes queries in the application process.\n- **Remote table**: LanceDB Enterprise/Cloud table opened through a `db://...` URI. The data may be very large, commonly backed by object storage, and queried through a remote service.\n\nDo NOT assume local-only table helpers exist on remote tables. If the user asks for LanceDB Enterprise, Cloud, `db://...`, production remote access, or a remote table, focus on the remote table path: use `search()` / `query()`, keep reads bounded with `select()` and `limit()`, and avoid table-level full materialization APIs.\n\n## Workflow\n\n1. Identify the SDK: Python, TypeScript, or both.\n2. Identify the table mode: local/embedded OSS, remote Enterprise/Cloud, or portable across both. If the user says \"LanceDB Enterprise\", choose the remote table path. If the task involves jobs in any way (listing, inspecting, creating, or canceling jobs), it is always the remote path and requires a remote server connection — see \"Connecting to the LanceDB remote server\" below before doing anything else.\n3. Read the matching topic reference before writing or changing code:\n - Column metadata authoring (both SDKs): `references/column_metadata.md`\n - Branch operations (both SDKs): `references/branch_ops.md`\n - Remote server connection resolution (jobs, raw REST): `references/remote_connect.md`\n - Job operations REST API (list/describe/cancel/query_events): `references/remote_jobs.md`\n\n There is no bundled per-language guide. For exact method names, signatures, and options, look them up in the canonical sources instead of relying on memory:\n - Python: `docs/src/python/python.md` (the hand-maintained API reference) and the source under `python/python/lancedb/` when working inside the LanceDB repo; otherwise <https://lancedb.github.io/lancedb/python/python/>.\n - TypeScript: the generated typedoc under `docs/src/js/` and the source under `nodejs/lancedb/` when working inside the LanceDB repo; otherwise <https://lancedb.github.io/lancedb/js/globals/>.\n4. Apply the SDK invariants in \"Per-SDK Invariants\" below. Read `column_metadata.md` when the task is documenting, tagging, classifying, or grouping table columns (field descriptions, `lancedb:tag:*` tags, logical column families). Read `branch_ops.md` when the task involves branch lifecycle (list/create/delete), writing to a non-main branch, or verifying a change stayed off main. Read `remote_connect.md` when the task involves jobs or direct REST access to an Enterprise deployment, and `remote_jobs.md` for the job REST methods themselves (list, describe, cancel, query_events).\n5. For Python schemas, favor Pydantic models and validate records before writing. Use PyArrow schemas when Arrow-native, streaming, or highly dynamic data makes them materially better suited.\n6. Prefer `search()` or `query()` builders with explicit `select()` and `limit()` for reads.\n7. Avoid table-level full materialization in remote or portable code. This is the main local-vs-remote read pitfall.\n8. After a successful embedded OSS ingestion, call `table.optimize()`. Do not call it for Enterprise/Cloud; remote maintenance is automatic.\n9. For remote Enterprise/Cloud writes, never drop-then-reuse or `mode=\"overwrite\"` the same table name — see \"Enterprise: never drop-then-reuse the same table name\" below. This is the main local-vs-remote write pitfall.\n10. If reviewing an existing file or repo, run `scripts/check_materialization.py` on the relevant paths and inspect each finding before editing.\n11. Cross-check unfamiliar or non-trivial API claims against the source tree instead of relying on memory.\n\n## Core Portability Rule\n\nDo not write code that assumes a local table API will exist on a remote table. Remote tables can be very large, so whole-table materialization helpers are intentionally unavailable or unsafe.\n\nThis does **not** mean result conversion is forbidden. Bounded query/search result collection is normal:\n\n- Python: `table.search(...).select([...]).limit(10).to_pandas()`\n- TypeScript: `await table.search(...).select([...]).limit(10).toArray()`\n\nThe unsafe pattern is table-level or unbounded collection, plus local-only dataset escape hatches in remote code:\n\n- Python: `table.to_pandas()`, `table.to_arrow()`, `table.to_polars()`; `table.to_lance()` is local/OSS-only dataset access, not materialization\n- TypeScript: `await table.toArrow()`, `await table.query().toArray()` without `limit()`\n\n## Per-SDK Invariants\n\nPython:\n\n- Result collectors: default to `.to_list()` (plain dicts, no extra dependency) or `.to_arrow()` (PyArrow ships with LanceDB). Use `.to_pandas()` / `.to_polars()` only when the project already declares that dependency — do not assume pandas or polars is installed.\n- Plain scans differ by client: the sync client has no `.query()` method — use `table.search()` with no argument; the async client uses `await async_table.query()`.\n\nTypeScript:\n\n- Collect bounded results with `.toArray()` (objects) or `.toArrow()` (Arrow) after `select()` and `limit()`.\n- For large reads, stream batches instead of collecting: `for await (const batch of table.query().where(...).select(...).limit(...)) { ... }`.\n\nBoth SDKs:\n\n- Ingest in bulk or in batches of thousands of rows; never write per-row in a loop — each write creates a version and fragment, slowing ingestion and later queries.\n- Build a vector index once brute-force search is too slow (rule of thumb: beyond roughly 100K vectors locally), and scalar indexes for filtered columns and merge/upsert keys. Use index defaults unless the task states recall/latency requirements.\n\n## Enterprise: never drop-then-reuse the same table name\n\nLanceDB Enterprise/Cloud splits a **control plane** (DDL: create/drop/rename) from a **data plane** (query nodes that serve reads). Query nodes cache the resolved dataset for a table name for up to `table_cache_ttl` — **default 300 seconds (5 minutes)**. After you drop or overwrite a table, the control plane updates immediately but the data plane keeps serving the *old* dataset until that cache entry expires. During the window the two planes disagree.\n\nThe failure this causes: you `drop_table(\"t\")` then immediately `create_table(\"t\", ...)` (or `create_table(\"t\", ..., mode=\"overwrite\")`). The DDL returns success, but every query against `t` returns **`500 Internal Server Error`** (the query node resolves the stale/deleted dataset), and a fresh `describe` may still show the *old* schema/version. It looks like your write silently failed; it didn't — the name is cached.\n\n**`mode=\"overwrite\"` has the same problem** — it is a drop+create of the same name under the hood.\n\nRules for portable Enterprise ingestion:\n\n1. **Never reuse a table name you just dropped/overwrote within the cache TTL.** Do not use `mode=\"overwrite\"` to replace an existing Enterprise table in place.\n2. To (re)load data, **write to a fresh table name** (e.g. `<table>_v2`, or a run-stamped suffix). A brand-new name has no cached data-plane entry, so writes and reads work immediately.\n3. Before creating, `list_tables()` and **fail loudly if the name already exists** rather than overwriting — prompt for a new name.\n4. To land on a specific final name that is currently occupied by an old table: drop the old table, **wait out the TTL (~5 min), then `rename_table(fresh_name, final_name)`**. Renaming onto a name whose old dataset is still cached hits the same race, so the wait is mandatory. `rename_table` is a supported control-plane op.\n5. When you hand a table name back to a human, tell them which step still needs the propagation wait (usually: \"the old `t` was dropped; run the rename in ~5 minutes\").\n\nThis is Enterprise/Cloud-specific. Local/OSS tables have no separate data plane, so `mode=\"overwrite\"` and immediate same-name reuse are fine there.\n\n## Connecting to the LanceDB remote server\n\nLanceDB Enterprise/Cloud deployments are served by a server implementing the lance-namespace OpenAPI spec (<https://github.com/lance-format/lance-namespace/blob/main/docs/src/spec.yaml>). Every remote (`db://...`) connection talks to such a server, and some operations exist only there. In particular, **all operations around jobs (listing, inspecting, creating, or canceling jobs) run server-side** — there is no local/OSS equivalent. Before any job work, or any direct REST call to an Enterprise deployment, read `references/remote_connect.md` to resolve the base URL, credentials, and database header and to validate the connection. Then use the four job REST methods documented in `references/remote_jobs.md` (list, describe, cancel, query_events).\n\n## Script\n\nRun the scanner when reviewing or modifying an existing codebase:\n\n```bash\npython skills/lancedb/scripts/check_materialization.py path/to/file_or_dir\n```\n\nThe script reports likely unsafe full-table materialization in Python and TypeScript. Treat results as review prompts, not automatic proof of a bug.\n"
}SHA-256 of public snapshot: 1a9fee785db248ad2816c2e821d5f8dea6ccb1590daab9804b48a933333481cc