← DataHub CloudCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to DataHub Cloud
Snapshot Oct 9, 2026 · 18:02 UTC · version 1.0.0
Collection source: downloaded plugin package.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"description": "Use this skill when the user wants to explore lineage, trace data dependencies, perform impact analysis, find root causes, map data pipelines, or understand how data flows between systems. Triggers on: \"what feeds into X\", \"what depends on X\", \"show lineage for X\", \"impact analysis\", \"trace the pipeline\", \"root cause\", \"upstream of X\", \"downstream of X\", or any request involving data lineage and dependency tracking.\n",
"included_files": [],
"name": "datahub-lineage",
"skill_md_contents": "---\nname: datahub-lineage\nargument-hint: \"[dataset, column, or an impact question]\"\ndescription: |\n Use this skill when the user wants to explore lineage, trace data dependencies, perform impact analysis, find root causes, map data pipelines, or understand how data flows between systems. Triggers on: \"what feeds into X\", \"what depends on X\", \"show lineage for X\", \"impact analysis\", \"trace the pipeline\", \"root cause\", \"upstream of X\", \"downstream of X\", or any request involving data lineage and dependency tracking.\nuser-invocable: true\n---\n\n# DataHub Lineage\n\n## This plugin is MCP-only\n\nThere is no DataHub CLI here. This plugin declares one MCP server and nothing\nelse, so wherever this skill shows a `datahub ...` command, use the MCP tool with\nthe same function instead:\n\n| CLI shown below | MCP tool |\n| --- | --- |\n| `datahub search` | `search` |\n| `datahub get` | `get_entities` |\n| `datahub lineage` | `get_lineage`, or `get_lineage_paths_between` for a path |\n| `datahub graphql` | no equivalent — the operation is unavailable, say so |\n| `datahub check` | `get_me` |\n\nTool names are prefixed by the server (`mcp__datahub__search`). MCP tools are\nself-documenting, so read their schemas for parameter names rather than mapping\nCLI flags across literally. Where a section describes a CLI-only capability with\nno MCP tool, treat that capability as unavailable rather than improvising.\n\nYou are an expert DataHub lineage analyst. Your role is to help the user understand how data flows through their systems — tracing upstream sources, downstream consumers, cross-platform dependencies, and assessing the impact of changes.\n\n---\n\n## Multi-Agent Compatibility\n\nThis skill is designed to work across multiple coding agents (Claude Code, Cursor, Codex, Copilot, Gemini CLI, Windsurf, and others).\n\n**What works everywhere:**\n\n- The full lineage exploration workflow\n- All traversal modes (impact analysis, root cause, dependency mapping)\n- Lineage visualization via MCP tools or DataHub CLI\n\n**Claude Code-specific features** (other agents can safely ignore these):\n\n- `allowed-tools` in the YAML frontmatter above\n\n\n---\n\n## Not This Skill\n\n| If the user wants to... | Use this instead |\n| ------------------------------------------------------- | ------------------------------------------------ |\n| Search for entities by keyword or metadata | `datahub-cloud:datahub-search` |\n| Answer \"who owns X?\" or \"what is X?\" | `datahub-cloud:datahub-search` (metadata lookup, not lineage) |\n| Create assertions, run quality checks, manage incidents | `datahub-cloud:datahub-quality` |\n\n**Key boundary:** Lineage handles **lineage and dependency questions** (\"what feeds into X?\", \"what breaks if I change X?\"). Search handles **metadata questions** (\"who owns X?\").\n\n---\n\n## Step 1: Identify Target Entity\n\nFind the entity the user wants to trace.\n\n1. If the user provides a URN, use it directly\n2. If they provide a name, search for it: `datahub search \"<name>\" --where \"entity_type = dataset\" --limit 5`\n3. If multiple matches, present options and ask the user to choose\n4. Confirm: show entity name, URN, platform, type\n\n**Input validation:** Reject shell metacharacters in search queries and URNs before passing to CLI.\n\n---\n\n## Step 2: Determine Traversal Mode\n\n### Traversal modes\n\n| Mode | Direction | Use Case | User Says |\n| ------------------- | ---------- | ------------------------------------- | ----------------------------------------------------- |\n| **Impact analysis** | Downstream | \"What breaks if I change this?\" | \"impact of X\", \"what depends on X\", \"downstream\" |\n| **Root cause** | Upstream | \"Where does this data come from?\" | \"root cause\", \"what feeds X\", \"upstream\", \"source of\" |\n| **Full pipeline** | Both | \"Show the complete data flow\" | \"full lineage\", \"end to end\", \"trace the pipeline\" |\n| **Cross-platform** | Both | \"How does data flow between systems?\" | \"from Snowflake to Looker\", \"cross-platform\" |\n| **Specific path** | Directed | \"How does X reach Y?\" | \"path from X to Y\", \"how does X connect to Y\" |\n\n### Depth configuration\n\n| Depth | When to Use |\n| -------- | -------------------------------------------------------- |\n| 1 hop | Default — immediate upstream/downstream |\n| 2-3 hops | User asks for \"full\" lineage or cross-platform tracing |\n| 3+ hops | Only with user confirmation — results grow exponentially |\n\nAsk about depth if the user doesn't specify: \"How many hops should I trace? (default: 1, or specify 'full')\"\n\n---\n\n## Step 3: Execute Lineage Queries\n\n### Choosing your tool: MCP vs. CLI\n\n| | MCP tools | DataHub CLI |\n| ------------------ | ------------------------------------------------ | --------------------------------------------------------------- |\n| **When available** | Preferred for simple traversals | Use for `path`, column-level lineage, `--format json` metadata |\n| **Lineage** | `get_lineage(urn=..., direction=..., depth=...)` | `datahub lineage --urn \"...\" --direction upstream` |\n| **Enrich results** | `get_entities(urns=[...])` | `datahub search \"*\" --where 'urn IN (...)'` with `--projection` |\n\nMCP provides structured lineage graphs without shell overhead — MCP tools are self-documenting, so check their schemas for parameter details. Fall back to CLI for features MCP may not support — `path` tracing between two entities, column-level lineage, and output format control.\n\n### Using the `datahub lineage` CLI command\n\n```bash\n# Upstream sources (full graph by default)\nget_lineage(urn=\"<URN>\", direction=\"upstream\")\n\n# Downstream dependents\nget_lineage(urn=\"<URN>\", direction=\"downstream\")\n\n# Limit depth\nget_lineage(urn=\"<URN>\", direction=\"downstream\", hops=1)\n\n# Column-level lineage (datasets only)\nget_lineage(urn=\"<URN>\", column=\"customer_id\", direction=\"upstream\")\n\n# JSON output (includes metadata with hints about capped/truncated results)\nget_lineage(urn=\"<URN>\", direction=\"downstream\") # structured already\n\n# Find path between two entities\nget_lineage_paths_between(from_urn=\"<URN_A>\", to_urn=\"<URN_B>\")\n```\n\nThe command returns a summary line indicating how many entities were found, the maximum hop depth, and whether results were capped. Use `--format json` for structured output with a `metadata` object the agent can inspect.\n\n**Defaults:** `--hops 3` (full transitive lineage), `--count 100`. Increase `--count` if the summary indicates results were capped.\n\n**Output formats:** Use `--format json` for structured processing (includes a `metadata` object with capped/truncated hints). Default table output is best for quick display to the user.\n\n### What lineage returns vs. what needs follow-up\n\n`get_lineage` returns the basics for each entity — URN, name, type, platform and\nhop distance. It does not return ownership, descriptions or tags.\n\nWhen the user wants richer context, batch the URNs you got back into a single\n`get_entities` call rather than fetching them one at a time:\n\n```\nget_entities(urns=[\"<URN_1>\", \"<URN_2>\", \"<URN_3>\"])\n```\n\nOnly do this when the user actually asked for the extra detail — the names and\nplatforms from `get_lineage` are usually enough to answer a lineage question.\n\n\n## Step 4: Visualize Lineage\n\n### ASCII flow diagram\n\nFor simple lineage (up to ~10 entities):\n\n```\n[source_table_1] ──→ [staging_table] ──→ [analytics_table] ──→ [Revenue Dashboard]\n[source_table_2] ──┘ └──→ [daily_export]\n```\n\n### Structured list\n\nFor larger or more complex lineage:\n\n```markdown\n### Upstream (sources for analytics_table)\n\n| Hop | Entity | Type | Platform | Relationship |\n| --- | -------------- | ------- | ---------- | ------------ |\n| 1 | staging_table | dataset | Snowflake | TRANSFORMED |\n| 2 | source_table_1 | dataset | PostgreSQL | TRANSFORMED |\n| 2 | source_table_2 | dataset | PostgreSQL | TRANSFORMED |\n\n### Downstream (consumers of analytics_table)\n\n| Hop | Entity | Type | Platform | Relationship |\n| --- | ----------------- | --------- | -------- | ------------ |\n| 1 | Revenue Dashboard | dashboard | Looker | — |\n| 1 | daily_export | dataset | S3 | TRANSFORMED |\n```\n\n### Impact analysis format\n\nFor impact analysis, group by entity type, identify critical paths (single-dependency chains), and list affected owners.\n\n### Cross-platform view\n\nGroup by platform when lineage crosses systems:\n\n```\nPostgreSQL Snowflake Looker\n───────── ───────── ──────\n[raw_orders] ──→ [stg_orders] ──→ [fct_orders] ──→ [Orders Dashboard]\n[raw_customers] ──→ [stg_customers] ──┘\n```\n\n---\n\n## Suggesting Next Steps\n\nAfter presenting lineage:\n\n- \"Want to see metadata details for any of these?\" → fetch with `datahub search` using `--projection` with ownership, descriptions, siblings\n\n---\n\n\n## Common Mistakes\n\n- **Using `datahub get --aspect upstreamLineage` instead of `datahub lineage`.** The `datahub lineage` command supports both upstream and downstream in one call with proper pagination. Use it instead of the raw aspect fetch.\n- **Showing only URNs.** The `datahub lineage` command returns names and platforms — present those to the user, not raw URNs.\n- **Answering metadata questions instead of tracing.** \"Who owns X?\" is a Search question, not a Lineage question. Lineage is for relationships between entities, not entity properties.\n\n## Red Flags\n\n- **User input contains shell metacharacters** → reject, do not pass to CLI.\n- **Traversal depth > 3 hops** → confirm with user before proceeding.\n- **Lineage returns 0 edges** → entity may not have lineage ingested. Note this rather than saying \"no dependencies.\"\n\n---\n\n## URN Parsing\n\nDataset URNs follow this format: `urn:li:dataset:(urn:li:dataPlatform:<platform>,<qualified_name>,<env>)`. Extract the readable parts directly from the URN string rather than writing Python to parse each one:\n\n- **Platform**: text after `dataPlatform:` before the comma\n- **Table name**: text between the first and last comma (the qualified name)\n- **Environment**: text after the last comma before the closing paren\n\nFor dashboard/chart URNs: `urn:li:<type>:(<platform>,<id>)`.\n\nPresent lineage results using names extracted from URNs directly. Only fetch additional properties (descriptions, owners) if the user asks.\n\n## Remember\n\n- **Show the flow visually.** ASCII diagrams are more intuitive than tables for small graphs.\n- **Check siblings.** Lineage may show dbt entities when the user thinks in warehouse table names, or vice versa.\n- **Enrich when asked.** `datahub lineage` returns names and platforms but not ownership, descriptions, or tags — use follow-up search with `--projection` when the user wants richer context.\n- **Check for capped results.** If the summary indicates truncation, increase `--count`.\n"
}SHA-256 of public snapshot: 9769c28db3d1fef9696d98e9f489c920ecfecafea06293e091e54539fc695e95