← Files SavvyARCHIVED FILE

skills/savvy/references/substrate/system-substrate.md

11.5 KB · Oct 5, 2026 · 18:27 UTC

↓ Download file

# System Substrate

> **Naming.** Savant's UI calls these **systems** (Data → Systems). The app API calls the same
> object a **connection** (`/api/connections`, each with a `connector` type). They are
> interchangeable; this project standardizes on **system** for everything user- and agent-facing,
> and keeps "connection" only where it mirrors the raw API (the `list_connections` helper / the
> endpoint). When talking to a user, always say "system."

## Objective

Use this to discover and bind a connected **system** in the target workspace — a user-set-up system
(OneDrive, Google Drive, …) that a workflow can write to as a **destination**. It is the
system-level companion to `dataset-substrate.md`: datasets are uploaded/static or system-backed
inputs; systems are the live connections those system-backed inputs and file destinations bind to.

## Use When

- A request names an external service to **write to** ("write to OneDrive", "export to Google Drive").
- Builder/Creator/Editor needs to bind a file destination to a real connected system.
- A request asks to **read from** a connected system (handle per "Sources from a system" below).
- You need to confirm a connection exists, is active, or has not expired before relying on it.

## Default Action

- Require `api_enabled`.
- **Read-only.** Connections are created and authenticated by a person in Savant. The skills only
  read existing connections — never create a system, never enter credentials, never run or complete
  an OAuth/SSO flow, never re-authenticate. If no suitable connection exists, stop and ask the user
  to set it up in Savant.
- Discover existing connections and bind by connection `id`. Prefer an unambiguous match; when more
  than one connection of the right type exists, **ask which one** (same discipline as datasets).
- Supported connectors today: **OneDrive (`onedrive`)** and **Google Drive (`googledrive`)**. Others
  (`s3`, `gcs`, `sharepoint`, `sftp`, `box`, …) exist in Savant but are out of scope here for now;
  treat them as not-yet-supported and ask the user rather than guessing their config shape.

## Do Not

- Do not create, authenticate, or re-authenticate a connection, and do not enter any credential or
  OAuth consent — that is always a person-in-the-UI action.
- Do not invent connection ids, folder links/ids, file names, or paths. These are environment
  bindings; take them from a real connection or from the user.
- Do not assume "the connection exists" means "the credentials are still valid" — check `status` and
  `expiresAt` (see below).
- Do not treat a connection as a dataset. A destination binds a connection through
  `fileSystemConfig`; a *source* still binds a dataset id (see "Sources from a system").

For the object model, read `savant-context.md`. For destination config shape, read
`../components/destination.md`. For dataset discovery/binding, read `dataset-substrate.md`.

**Discovery runs off the MCP session and does not need `api_enabled`** (see below). **Binding a
destination** (building + importing the node) still hits the live app API and needs `api_enabled`; if
it is false you can still discover and name a connection but cannot import the change — say so rather
than claim it was applied.

## How do I discover a system?

Use the resources `search` tool (works without `api_enabled`):

1. `search(types=["connection"])` — paginate via `cursor`. Each hit carries `id`, `title` (name),
   `summary` (`<connector> connection (<status>)`), and **`namespace`** (the connection's true owner).
2. **Filter by namespace.** Connection search is **org-wide**: it returns this workspace's own
   connections *and* org-shared ones owned by other workspaces, each tagged with its own `namespace`
   (not your session's). Keep only hits whose `namespace` matches the workspace you'll bind in — a
   destination must bind a connection usable in the target workspace.
3. **Filter by connector** client-side from the `summary` token (`onedrive` / `googledrive` / …);
   there is no server-side connector argument.
4. For one connection's detail — `owner`, `expiresAt` (token expiry, epoch ms), `status` — call
   `fetch(savant://connection/{id})`.

- Turn a connection **name** into its **id** from the search hits; bind by id.
- If several connections of the wanted connector exist (after the namespace + connector filter), ask
  which one by display name — do not pick for the user:
  > I see two Google Drive connections — `Anil Google Drive` and `Finance GDrive`. Which should this output write to?
- If none of the wanted connector exists in the target namespace, stop: the user must create and
  authenticate it in Savant first. Name what's missing.

## How a user finds a system id by hand (no API)

When MCP connection search is unavailable, the user supplies the system id.
Tell them how to get it from the Savant UI:

1. Open Savant → **Data** → **Systems** tab.
2. Find the system, open its row menu (the `⌄`/caret), and click **Edit**.
3. The id is in the browser URL: `app.savantlabs.io/en/app/connection/{id}/edit` — e.g.
   `.../connection/jdvyptumte/edit` means the system id is `jdvyptumte`.

That Edit screen also shows the account, the auth **expiry date**, and a **Re-authenticate** button —
useful if a run fails on expired credentials (re-auth is the user's action, never ours).

## Auth / lifecycle: existence is not validity

A connection can exist and still fail at run time if its token has expired or it was revoked.

- `status` should be `Active`; `expiresAt` is the token expiry (epoch ms). If `expiresAt` is in the
  past or `status` is not active, say so — a run will fail to authenticate, and that surfaces at
  **run time**, not when you wire the node.
- The skills cannot refresh or re-authenticate — re-auth is a person-in-the-UI step.

## Destinations: writing to a connected system (supported)

A file destination binds to a connection and writes a file into it. The node shape lives in
`../components/destination.md` / `../registry/components/destination.json`; this substrate owns the
**discovery and binding** layer around it.

- Resolve the target system **id** via connection `search` (filtered to the target namespace; ask if
  ambiguous), then build the
  destination with `connector`/`type` set to that connector (`onedrive` / `googledrive`) and a
  `fileSystemConfig` describing the write. **The destination binds to the system by `config.id` =
  the system id** (e.g. the Google Drive system `jdvyptumte` → the node's `config.id`), exactly as a
  source binds a dataset id. File destinations carry no `mode` (unlike native CSV). Build with
  `nb.destination_file(name, connector=..., system_id=..., folder_link=..., file_name=...,
  file_type="EXCEL"|"CSV", tab_name_mode=..., tab_name_field=..., subsequent_mode=...)`.
- **Offline (`api_enabled` false):** you can still build the node if the user supplies the system
  **id** (the `config.id` binding), the connector type, and the `folder_link`. You cannot verify the
  id resolves — a wrong/missing id is silently dropped on import — so the id must be correct.
- **Import-drop gotcha (verified).** A file destination whose connector has **no matching connection
  in the target workspace is silently dropped on import** — Savant removes the node, the created flow
  comes back one node short, and there is no import error. This mirrors the AI-provider behavior in
  `ai-provider-substrate.md`. So Creator must confirm the connection exists *before* import; the only
  signal otherwise is a node-count shortfall against the source JSON.
- **Never invent the folder binding.** `folderLink` (OneDrive/SharePoint share URL) and the Google
  Drive folder reference are environment values — take them from the user or a real connection, or
  leave a clearly labelled placeholder for the user to bind in Savant.

Observed OneDrive destination `fileSystemConfig` (reference shape):

```json
{
  "fileType": "EXCEL",            // EXCEL or CSV
  "fileName": "Fixed Asset by Legal Entity",
  "fileNameMode": "static",       // or "field" + "fileNameField" (a real upstream column)
  "tabNameMode": "field",         // one tab per value of "tabNameField"
  "tabNameField": "Legal Entity",
  "folderLink": "https://…sharepoint.com/…",
  "subsequentMode": "replace",    // default replace; "append" accumulates across runs
  "flatFileConfig": {"fileWriterProps": {"delimiter": ",", "qualifier": "\"", "escape": "\\", "charset": "UTF_8"}}
}
```

Google Drive (`googledrive`) is the same `fileSystemConfig` family with its own folder reference
instead of a SharePoint `folderLink`. The exact folder field is environment-bound — read it from the
chosen Google Drive connection / a real node rather than assuming it, and confirm with the user.

## Writing to the same file from several destinations (assembling a workbook)

A file destination **updates an existing file in place — it does not recreate it.** So several
destinations can point at the **same file** (same folder + file name), within one workflow or across
workflows, to build up one workbook:

- **Different tabs → one workbook.** Point destination A at file `X` tab "Summary" and destination B
  at the same file `X` tab "Detail" (or use `tab_name_mode="field"` to fan out tabs); the file ends
  up with all the tabs. This is how you assemble a multi-tab Excel from multiple branches or flows.
- **Updating / adding data.** A later write to file `X` updates it rather than replacing the whole
  file; tabs not touched by a write are left intact. `subsequent_mode="replace"` replaces the
  content of the tab being written; `"append"` adds rows to it.
- **Last write wins.** If two writes target the **same tab** (same file + same tab name), the later
  one overwrites the earlier — there is no merge. Within a single run, ordering is not something you
  control finely, so don't have two destinations write the same tab of the same file expecting both
  to survive; give them distinct tab names (or distinct files). Across separate runs/flows, the most
  recent run's write is what remains.

Design implication: to build one workbook with one tab per legal entity, a *single* destination with
`tab_name_mode="field"` is cleaner than many destinations; use multiple destinations to the same file
when the tabs come from genuinely different branches/flows, and keep their tab names distinct.

## Sources from a system (not yet supported — coming soon)

Reading from OneDrive/Google Drive is **not** a direct source-to-connection bind. In Savant you
create a **dataset from the system** (New Dataset → Connect → Select System → pick the file), which
produces a *system-backed dataset*; the source node then binds that dataset id like any other.

The skills do not create system-backed datasets (creating connector-backed/credentialed datasets is
out of scope — see `dataset-substrate.md`). So today:

- **Treat "read from OneDrive/Google Drive" as not yet supported / coming soon.** Tell the user the
  source must first exist as a dataset created from that system in Savant.
- Once the user has created that dataset, it appears in `search(types=["source"])` and is bound
  exactly like any uploaded dataset via `dataset-substrate.md` — no special handling.

## Enforcement / boundaries

- Connection discovery is read-only; binding a connection (create/import/edit) is a state-changing
  action and follows the usual confirm-before-acting rule.
- The validator does not check connector strings (300+ vary per workspace); validity for a
  destination comes from binding to a real connection that exists in the workspace — which is exactly
  why the existence check here matters before import.

SHA-256: 26fcb9aaa261268ddb7e6dbd54ae0f08bf66e3aba3315945629a126204e7d73f