← Plugin catalog
Productivity

MarcoPolo

Immersa, Inc. v3.0.0

Publisher description

From the marketplace listing

MarcoPolo spins up a secure container where ChatGPT can work with your actual data. Connect to your databases, APIs, S3, lakehouses, CRMs, Jira, logs and much more—using scoped credentials that are never exposed to the model. ChatGPT gets DuckDB, Python, a shell, and a set of tools to explore, query, transform, and analyze data across systems. The workspace persists over time, so that you can build on your work. Prep a report, understand your data, debug an issue, or review the latest metrics right here in the conversation.

Language: English · Automatically detected from descriptions.

Files & skills

File archives

Plugin package2 files · 892 BytesBrowse files →
query-and-analyze1 files · 3.2 KBBrowse files →
setup-connection1 files · 2.29 KBBrowse files →
using-connection-cli9 files · 7.37 KBBrowse files →
using-marcopolo-workspace1 files · 2.95 KBBrowse files →
Skill instructions
query-and-analyze7.75 KB

View saved version →

---
name: query-and-analyze
description: Queries connections, joins results across sources through DuckDB, and analyzes workspace data. Use this skill whenever the user wants to look at, count, summarize, filter, group, join, aggregate, compare, or analyze data, or when exploring a connection's schema or doing follow-up analysis on a previous query.
---

# Query and analyze data

Use this skill for all agent-side analytics: query authoring, schema exploration,
DuckDB materialization, joins across sources, and file analysis.

**Tool selection:**
- `workspace_shell("connection query ...")` is the correct tool for all agent
  analytics. The full result is always materialized into DuckDB; `--sample-rows`
  only controls how many rows come into the agent's context window.
- `data_query` is for generated code that re-queries live data at view or load
  time: Remote Artifacts, external web apps, scheduled scripts. Do not use it
  for agent analytics or one-off snapshot visualizations — embed those inline.

**Session compatibility:** Some sessions (e.g. ChatGPT) expose only
`workspace_shell` and do not have `connections_list` or `data_query`. Check
which tools are available before deciding on a path. `workspace_shell` works
in every session and is the primary analytics tool in all cases.

## Required workflow — follow every step

### Step 1 — Orient in the workspace

```text
workspace_shell("cat /workspace/RULES.md")
```

This is the workspace's long-term memory. Read it before touching any connection.

### Step 2 — Discover connections

Use `connections_list` if the current session exposes it. Otherwise:

```text
workspace_shell("connection list --json")
```

Pick the connection(s) needed. Confirm `query` appears in each connection's
`capabilities`.

### Step 3 — Load connection context

For each connection:

```text
workspace_shell("cat connections/<name>/README.md connections/<name>/SYNTAX.md connections/<name>/RULES.md")
```

`RULES.md` is long-term memory for that connection — field quirks, naming
conventions, reliable query patterns, and known limitations accumulated from
prior sessions. Read it before authoring any query.

### Step 4 — Inspect schema and existing queries

```text
workspace_shell("ls connections/<name>/queries/ connections/<name>/metadata/")
```

If `describe` is in capabilities and metadata is missing or stale:

```text
workspace_shell("connection describe <name> --json")
```

Prefer adapting an existing query over writing from scratch.

### Step 5 — Author a query file

Write to a file under `connections/<name>/queries/`. Use a business-readable
filename. The file extension and query format are determined by the connection
type — use `SYNTAX.md` (loaded in Step 3) as the authoritative reference for
both. Do not default to `.sql` unless SYNTAX.md confirms the connection uses SQL.

```text
workspace_shell("""cat > connections/<name>/queries/<filename>.<ext> <<'EOF'
<query content per SYNTAX.md>
EOF""")
```

**Common patterns by connection type:**

| Connection type | Extension | Query format |
|---|---|---|
| SQL databases (Snowflake, BigQuery, DuckDB) | `.sql` | Standard SQL `SELECT` |
| Salesforce (SOQL) | `.json` | `{"soql": "SELECT ... FROM Object WHERE ..."}` |
| Object storage (S3, SFTP) | `.json` | Path or glob pattern per SYNTAX.md |
| Document storage (Google Drive, OneDrive) | `.json` | File path or search spec per SYNTAX.md |
| Other SaaS APIs | `.json` | Endpoint + parameters object per SYNTAX.md |

If SYNTAX.md does not specify an extension, default to `.json` for API-based
connections and `.sql` for SQL-native connections.

### Step 6 — Execute the query

```text
workspace_shell("connection query <name> --file connections/<name>/queries/<filename>.<ext> --sample-rows 10 --json")
```

The full result is always materialized into DuckDB. `--sample-rows <n>` controls
how many rows appear in `preview` (default 10; omitting it truncates silently).
Pass `--sample-rows -1` when you need all rows in the payload. For large result
sets prefer a DuckDB follow-up query instead. See the `using-connection-cli`
skill for full flag reference and timeout guidance.

`preview` in the response envelope is a JSON-encoded **string** — call
`json.loads(resp["preview"])` to get `list[dict]`. `rows` in the envelope is an
int count, not a record list. For group-bys, totals, or joins, skip parsing
`preview` and run a DuckDB query over `relation_name` (Step 7) instead.

### Step 7 — Analyze and join through DuckDB

Each upstream query materializes as a `relation_name` in DuckDB. Use it for
aggregations, joins across connections, and transformations:

```text
workspace_shell("connection query DUCKDB --file connections/DUCKDB/queries/<file>.sql --json")
```

Save reusable joins in `connections/DUCKDB/queries/`.

### Step 8 — Export large results for the user

When the user needs to retrieve a large result set, export from DuckDB to CSV
in `/workspace/data/downloads/` for pickup from the MarcoPolo web UI:

```sql
COPY (SELECT * FROM <relation_name>) TO '/workspace/data/downloads/<filename>.csv' (HEADER, DELIMITER ',');
```

Run via:

```text
workspace_shell("connection query DUCKDB --file connections/DUCKDB/queries/export.sql --json")
```

### Step 9 — Offer to save learnings

After answering the user's question, offer to save any new facts discovered —
schema quirks, reliable query patterns, field naming conventions, known
limitations — to the appropriate RULES.md:

- Connection-specific: `connections/<name>/RULES.md`
- Workspace-wide: `/workspace/RULES.md`

Ask the user to confirm before writing. Saving these enriches the context layer
for future sessions.

## Join across connections through DuckDB

DuckDB is the in-workspace analytical connection.

1. Run each upstream `connection query` first and note each `relation_name`.
2. Write a DuckDB SQL file that joins or transforms those relations.
3. Execute through DuckDB.

   ```text
   workspace_shell("connection query DUCKDB --file connections/DUCKDB/queries/<file>.sql --json")
   ```

## Work with files in the remote workspace

- User-provided files belong in `data/uploads/`.
- Files fetched via `connection download` land in `data/downloads/`.
- Use `data/databases/` for database files when needed.

DuckDB can read CSV, Parquet, and JSON files directly:

```text
workspace_shell("""cat > connections/DUCKDB/queries/<file>.sql <<'SQL'
SELECT * FROM read_csv_auto('data/uploads/<file>.csv') LIMIT 100
SQL""")
workspace_shell("connection query DUCKDB --file connections/DUCKDB/queries/<file>.sql --json")
```

## Common pitfalls

- The `connection` CLI only exists inside the remote workspace — always use
  `workspace_shell` for workspace commands.
- Trust `connection list --json` or `connections_list` for capabilities.
- Query through named files, not inline SQL.
- **`--sample-rows` defaults to 10 and silently truncates.** Omitting it does
  not return all rows — it caps `preview` at 10. If `row_count` exceeds the
  length of `preview`, the result is truncated; use a higher `--sample-rows`
  value to get more rows, or `--sample-rows -1` to get all rows in the payload.
  The full dataset is always in DuckDB regardless of this flag.
- Query file paths in `--file` resolve from `/workspace`, not from your current
  directory — always include the `connections/<name>/` prefix. A bare
  `queries/<file>` resolves to `/workspace/queries/<file>` and fails with
  "No such file or directory" even if you just created the file via
  `cd <connection-dir> && cat > queries/<file>`.

## Pointers

- adding a connection or installing a demo → `setup-connection`
- per-verb flag reference → `using-connection-cli`
- workspace layout → `using-marcopolo-workspace`
- visualizing results → `build-dashboard`
- building a scheduled workflow → `build-scheduled-pipeline`
- managing an existing recurring run → `setup-automation`
setup-connection5.21 KB

View saved version →

---
name: setup-connection
description: Adds a new connection to the MarcoPolo workspace — hosted demo connections (no credentials) and credentialed connections to databases, warehouses, APIs, and storage (Postgres, Snowflake, BigQuery, Salesforce, S3, Google Drive, etc.). Use this skill whenever the user mentions adding, connecting, installing, hooking up, or wiring up a datasource — even when they describe it informally ("connect my Snowflake", "try the demo data", "let me hook up our Salesforce", "I want to look at the data in S3"). Also use when troubleshooting `connection test` failures, expired credentials, or an OAuth flow that didn't finish.
---

# Set up a connection

There are two paths: a **hosted demo connection** (no credentials, installs
in one call) and a **credentialed connection** (the user opens a browser
setup flow). Both end with the same verification steps inside the workspace.

The in-workspace canonical reference is `/workspace/workflows/setup-connection.md`.

## Path A — install a hosted demo connection

Use this when the user wants to try MarcoPolo without bringing their own
credentials, or asked for a specific demo dataset.

Call the MCP tool directly:

```
install_demo_connection(
  demo_connection="<id-or-natural-language>",
  intent_text="<optional free text>",
  display_name="<optional friendly name>",
)
```

Behavior:

- If `demo_connection` matches a known demo id, it installs immediately.
- If ambiguous, the response has `success: false`, `resolution_mode:
  "ambiguous"`, and `available_demo_connections: [{id, label, description, type}, ...]`.
  Show the user the choices and call again with a specific `id`.
- If unknown, the response includes `available_demo_connections` you can
  offer the user.

On success, run the post-install verification (below).

## Path B — add a credentialed connection

Use this when the user has their own database, warehouse, API, or storage
account.

1. Generate the setup URL via the MCP tool.

   ```
   connection_setup(type="<canonical-type>", intent_text="<optional free text>")
   ```

   `type` should be a canonical type value (`pg`, `mysql`, `snowflake`,
   `bigquery`, `s3`, `google_drive`, `salesforce`, `local_file`, etc.). If
   unsure, pass the user's words as `intent_text` and a best-guess `type` —
   the tool will resolve via intent if `type` is non-canonical. If still
   unknown, the response returns `valid_types` and `suggested_types`; pick
   from those and retry.

   On success, the response includes:
   - `url` — open this in a browser; the user signs in and configures
     credentials
   - `workflow_type` — typically `oauth` or `configure`
   - `instructions` and `next_actions` — surface these to the user
   - `configuration_schema` (for `configure` workflows) — the fields the
     setup UI will collect
   - `workspace_ssh_keypair` (for connections that support SSH tunnelling)
     — show the public key so the user can authorize it on their bastion

2. Surface the URL to the user and wait. Do not try to complete setup from
   the session — the user has to click through the browser flow.

3. Once the user says they're done, confirm the connection is visible.

   ```
   workspace_shell("connection list --json")
   ```

   If it doesn't appear yet, wait briefly and retry — provisioning can take
   a moment.

## Post-install verification (both paths)

1. Verify credentials.

   ```
   workspace_shell("connection test <name> --json")
   ```

   On failure, surface `error` and `message` to the user. For credential
   issues, send them back through `connection_setup` to update credentials.

2. Read the seeded connection docs.

   ```
   workspace_shell("cat connections/<name>/README.md connections/<name>/SYNTAX.md connections/<name>/RULES.md")
   ```

   The `README.md` lists the connection's authoritative `capabilities`. Note
   them before doing anything else with the connection.

3. Write initial metadata snapshots.

   ```
   workspace_shell("connection describe <name> --json")
   ```

   This populates `connections/<name>/metadata/`. The snapshot files become
   the default in-workspace reference for query authoring.

4. Confirm the directory shape.

   ```
   workspace_shell("ls connections/<name>/")
   ```

   Expect: `README.md`, `RULES.md`, `SYNTAX.md`, `queries/`, `metadata/`,
   `profile/`, `scratch/`.

## Troubleshooting

- **`connection_setup` returns `Unknown connection type`.** The response
  includes `valid_types` and `suggested_types`. Pick a canonical value and
  retry. If user intent is natural language, also pass `intent_text`.
- **`install_demo_connection` returns `ambiguous`.** Surface
  `available_demo_connections` to the user and retry with a specific id.
- **`connection test` fails after setup.** Likely cause: incomplete browser
  flow, wrong host/port, or missing network access. Surface the error
  message to the user; for credential changes, rerun `connection_setup` to
  reissue the setup URL.
- **Connection not visible in `connection list --json`.** Wait and retry —
  provisioning may still be running. If it persists, surface that to the
  user.

## Pointers

- writing the first query → `query-and-analyze`
- per-verb flag reference → `using-connection-cli`
- workspace layout → `using-marcopolo-workspace`
using-connection-cli3.99 KB

View saved version →

---
name: using-connection-cli
description: Reference for the in-workspace `connection` CLI — verb shape, JSON envelope, the capability rule, and per-verb flag details. Use this skill whenever any `connection` verb (`list`, `add`, `test`, `describe`, `query`, `browse`, `download`, `upload`) is about to run, when looking up flags, when checking whether a verb is allowed by a connection's capabilities, or when reading the JSON response. Consult it even for routine commands — guessing flag names or capability-gating leads to wasted work and surprising failures.
---

# Using the `connection` CLI

The `connection` CLI is the verb surface for all connection work inside
the MarcoPolo workspace. It runs in the workspace pod; from a session,
invoke it through `workspace_shell`:

```
workspace_shell("connection <verb> [args] --json")
```

Always pass `--json`. The output is then a structured envelope you can
parse — without `--json` you get human-formatted text that's harder to
work with programmatically.

## JSON envelope

Every `--json` response has at least:

```json
{ "success": true | false, "operation": "<verb>", ... }
```

On failure: `error`, usually `message`, and often `next_actions` and
(for unknown types) `suggested_types`. On success: verb-specific fields
documented in the per-verb references below.

## Capability rule

`connection list --json` returns each connection's `capabilities` array.
That list is **authoritative** — never call `browse`, `download`, or
`upload` on a connection unless the verb appears in its capabilities.

The reason: capabilities depend on connection type, the user's auth
state, and platform configuration. Calling a non-advertised verb wastes
work, may produce confusing errors, and clutters the workspace's audit
trail. The `connections/<name>/README.md` file mirrors the same
capabilities — both come from the same source.

## Verbs at a glance

| Verb | What it does | Reference |
|---|---|---|
| `list` | Discover connections + capabilities | `references/list.md` |
| `add` | Get a browser setup URL for a credentialed connection | `references/add.md` |
| `test` | Verify stored credentials | `references/test.md` |
| `describe` | Write metadata snapshots into `connections/<name>/metadata/` | `references/describe.md` |
| `query` | Execute a saved query file; materialize result into DuckDB | `references/query.md` |
| `browse` (gated) | List provider-side files for storage connections | `references/browse.md` |
| `download` (gated) | Fetch a provider file into the workspace | `references/download.md` |
| `upload` (gated) | Push a workspace file to the provider | `references/upload.md` |

Read the per-verb reference before running a verb you haven't run
recently, especially for flags. The references include the exact
response shape, common pitfalls, and follow-on commands.

**`connection query` — three facts that cause most retries:**
- **Path:** `--file` resolves from `/workspace`, ignoring cwd. Always pass
  `connections/<name>/queries/<file>`; a bare `queries/<file>` fails with
  "No such file or directory" even if the file was just created.
- **`--sample-rows`:** defaults to 10 — omitting it silently truncates `preview`.
  Use a higher value to get more rows, or `-1` to get all rows in the payload.
- **Response:** `preview` is a JSON-encoded *string* — call `json.loads` on it
  to get records; `rows` in the envelope is an int count, not a record list.
  The full result lives in DuckDB as `relation_name`.

See `references/query.md` for the full flag contract and response shape.

## Self-discovery

When in doubt, ask the CLI directly:

```
workspace_shell("connection --help")
workspace_shell("connection <verb> --help")
```

These are the live source of truth for flags. Prefer them over guessing
or relying on the references if the workspace platform may have moved
ahead of this skill.

## Pointers

- adding a connection end to end → `setup-connection`
- writing and running a query → `query-and-analyze`
- workspace layout and where files belong → `using-marcopolo-workspace`

Referenced files: 8

using-marcopolo-workspace7.51 KB

View saved version →

---
name: using-marcopolo-workspace
description: Orientation for the MarcoPolo remote workspace, what it is, how `/workspace` is laid out, when to use the product MCP data tools versus `workspace_shell`, and how the `connection` CLI fits in. Use this skill whenever MarcoPolo, the marcopolo MCP server, `workspace_shell`, `/workspace`, connections, or `connection` CLI commands come up. Read this first when entering a MarcoPolo session, before reaching for a more specific skill.
---

# Using the MarcoPolo workspace

MarcoPolo is a persistent remote Linux workspace at `/workspace` for working
with company data, building dashboards, scheduling jobs, and keeping a durable
collection of queries, scripts, and artifacts.

## Two execution surfaces

Two execution surfaces coexist in a MarcoPolo session:

- `workspace_shell` for all agent-side work: query authoring, analytics, DuckDB
  joins, workspace files, scripts, git, and cron inside `/workspace`.
- Product MCP data tools (`connections_list`, `data_query`) for generated code
  that re-queries live data at view or load time — Remote Artifacts, external
  web apps, scheduled scripts.

For all agent analytics, use `workspace_shell`. Reserve `data_query` for
programmatic interfaces, not for the agent's own data exploration.

## Session capability detection

Check which tools are available in the current session before choosing a path:

- Sessions with `connections_list` and `data_query` (Claude, Cursor, etc.):
  - Agent analytics → always use `workspace_shell`
  - Programmatic interfaces (web apps, scripts, dashboards) → use `data_query`
- Sessions with only `workspace_shell` (ChatGPT, older sessions):
  - Agent analytics → use `workspace_shell`
  - Generated artifact code → use bounded `workspace_shell("connection query <name> --file <file> --sample-rows <n> --json")`, noting in the code that it can be upgraded to `data_query` if the session gains that tool

`workspace_shell` is the primary analytics tool in every session. `data_query`
is an addition for programmatic interfaces, not a replacement for agent work.

When using `workspace_shell` for queries, treat results as CLI envelopes:

- rows from `data`, otherwise `preview`
- `row_count` from `row_count`, otherwise `len(rows)`
- `run_id` if present
- `relation_name` if present

If `row_count` exceeds the length of `preview`, the preview is truncated —
use a higher `--sample-rows` value to get more rows, or `--sample-rows -1`
to get all rows in the payload.

## Two shell environments

Two shell environments coexist in this session:

- Your built-in shell and filesystem tools act on the client's own environment.
- `workspace_shell` runs commands inside the MarcoPolo remote workspace at
  `/workspace`.

Your built-in tools cannot reach the MarcoPolo workspace. They cannot read or
create files there, run the `connection` CLI or `crontab` that only exist there,
or see git state inside it. Only `workspace_shell` can.

So for all MarcoPolo workspace work, such as reading files, writing queries,
running scripts, or inspecting git, use `workspace_shell`. Reach for your
built-in tools only for things outside MarcoPolo.

User-uploaded files land in `data/uploads/` inside the MarcoPolo workspace.
`workspace_shell` reads them, not the built-in tools.

## Common `workspace_shell` operations

Treat `/workspace` like a checked-out repo. Common shapes:

- read files: `workspace_shell("cat /workspace/RULES.md")`
- list and search: `workspace_shell("ls connections/")`,
  `workspace_shell("rg <pattern> connections/")`
- write and edit files: `workspace_shell` with heredocs, `sed`, or other shell
  tools
- run scripts: `workspace_shell("python scripts/<file>.py")`
- inspect git state: `workspace_shell("git status")`,
  `workspace_shell("git diff")`

Read `RULES.md` and the relevant `workflows/` guide before authoring; use git
as part of normal work.

## MCP tool families

Product data tools:

- `connections_list` for connection discovery when available
- `data_query` for bounded governed query execution when available

Workspace and ext-app tools:

- `workspace_shell(command, timeout=30)` for remote workspace commands
- `connection_setup(type, intent_text=None)` for credentialed connection setup
- `install_demo_connection(demo_connection, display_name=None, intent_text=None)`
  for hosted demo connections

Some sessions may also expose legacy or host-specific tools. Do not rely on
them as the primary dashboard or query path unless a more specific skill tells
you to.

## The `connection` CLI is the workspace verb surface

For full reference see the `using-connection-cli` skill. The shape:

```text
connection <verb> [args] --json
```

Common verbs: `list`, `add`, `test`, `describe`, `query`, `browse`, `download`,
`upload`. Always pass `--json` so output is structured.

`connection list --json` returns each connection's `capabilities` array. That
list is authoritative. Never call `browse`, `download`, or `upload` on a
connection unless that verb appears in its capabilities.

## Workspace layout

```text
/workspace/
  README.md                       workspace overview
  RULES.md                        workspace-wide rules and conventions
  workflows/                      curated guides for recurring tasks
    README.md
    setup-connection.md
    query-and-analyze-data.md
    build-dashboard.md
    setup-automation.md
  connections/                    one subdirectory per visible connection
    <name>/
      README.md
      RULES.md
      SYNTAX.md
      queries/
      metadata/
      profile/
      scratch/
    DUCKDB/
  scripts/
  artifacts/
  data/
    uploads/
    downloads/
    databases/
  .dv/
```

Always read first before authoring:

- `workspace_shell("cat /workspace/RULES.md")`
- `workspace_shell("cat /workspace/workflows/README.md")`
- `workspace_shell("cat connections/<name>/README.md connections/<name>/RULES.md connections/<name>/SYNTAX.md")`
- Before authoring or running any query, also read the `query-and-analyze` and
  `using-connection-cli` skills — they are prerequisites, not optional
  further reading.

`RULES.md` files are long-term memory — the workspace-level one holds general
conventions, and each `connections/<name>/RULES.md` holds connection-specific
facts: field quirks, reliable query patterns, naming conventions accumulated
from prior sessions. Read them before authoring queries and update them when
you discover new facts.

## DUCKDB is a connection

DUCKDB is the in-workspace analytical connection, backed by
`.dv/duckdb/workspace.duckdb`. Query it through the `connection` CLI:

```text
workspace_shell("connection query DUCKDB --file connections/DUCKDB/queries/<file>.sql --json")
```

Use it for joins across connections, intermediate tables, and in-workspace
derived datasets.

## Where to put things

- query files -> `connections/<name>/queries/`
- metadata snapshots -> `connections/<name>/metadata/`
- reusable programs -> `scripts/`
- user-facing outputs -> `artifacts/`
- scheduled jobs -> the user crontab (`crontab -l`), not a workspace file
- user-provided data -> `data/uploads/`
- fetched data -> `data/downloads/`
- database files -> `data/databases/`

Do not write to `.dv/`; it is runtime-managed.

## Pointers

- adding a connection, installing a demo, fixing credentials -> `setup-connection`
- querying data, exploring schemas, joining sources -> `query-and-analyze`
- building a chart or dashboard -> `build-dashboard`
- building a scheduled data or AI workflow -> `build-scheduled-pipeline`
- managing an existing recurring job -> `setup-automation`
- before running any `connection` verb (even routine ones) -> `using-connection-cli`
Package details

Publisher declarations from the archived package. These are separate from our research and the live service's terms.

Package author
Immersa, Inc.

Package observed Sep 30, 2026.

Technical details
First seen
Sep 30, 2026 · 22:02 UTC
Last seen
Oct 1, 2026 · 18:00 UTC
Collection status
Collected

plugin_asdk_app_698429b2c5fc8191bb997f52cb2a413a

Download plugin data (JSON)