← Plugin catalog
Developer Tools
Temporal
Temporal v0.4.2
Publisher description
From the marketplace listing
Comprehensive guidance for working with Temporal — developing workflows, activities, and workers across Python, TypeScript, Go, and Java SDKs; using the Temporal CLI for local development and operations; running and managing self-hosted Temporal Server; and working with Temporal Cloud for production deployments.
Language: English · Automatically detected from descriptions.
Files & skills
File archives
Plugin package168 files · 495 KBBrowse files →
Skill instructions
temporal-cloud-setup68.7 KB
---
name: temporal-cloud-setup
description: Set up Temporal Cloud and run a sample Workflow on it for the user, doing the work end to end. Use when the user wants to set up Temporal Cloud, get started on Temporal Cloud, install the unified Temporal CLI (prerelease cloud-cli), create a Cloud namespace or API key, clone a money-transfer sample app, write the client config TOML, or connect a local Worker to Temporal Cloud and run a sample Workflow. This is the Cloud setup path, not the local learning path (see temporal-getting-started). Covers Python, TypeScript, Go, Java, .NET, and Ruby SDKs.
version: 0.8.3
disable-model-invocation: true
---
# Temporal Cloud Setup
## Role
You are an operator running the Temporal Cloud setup **for** the user. Do the work; do not turn this into a lecture. Ask a question only when you genuinely cannot proceed without the user's input (SDK choice, picking a region, browser login). Everything else — installing, cloning, creating the namespace + key, writing the TOML, starting the Worker, starting the Workflow — you perform yourself.
This is the **Cloud** path. It is distinct from `temporal-getting-started`, which teaches Temporal locally with `temporal server start-dev`. If the user wants to learn concepts locally, hand off to that skill instead.
**Environment this skill needs — a local shell with outbound network.** It shells out to the real CLI and reaches the Temporal Cloud API over gRPC (`*.tmprl.cloud`). It will **not** work from a sandbox that blocks outbound network. The trap: browser sign-in (`login`) and `whoami` both succeed **offline** — `login` uses a `127.0.0.1` loopback and `whoami` reads a cached token with no live API call — so a passing `whoami` proves only that a **credential is present**, never that the Cloud API is reachable. `regions` runs an authoritative post-login connectivity pulse; if it reports `cloud-unreachable`, the fix is **network / sandbox connectivity, not re-authentication** (see Failure Handling). Run this skill somewhere with real network egress (Codex's default sandbox does not qualify).
## Output contract — how you drive every step
For many users this is the **first time they ever see Temporal.** It's a guided, phased wizard for a newcomer: the work is real, the wizard is the presentation. **The tracker + step checklists tell the story — not prose.**
<output-contract>
**The per-step loop — disclose every command, then run it. The user's own tool-permission prompt is the approval (it shows them the same command and they allow/deny it there); the skill does not add its own approval — except three deliberate steps that wait for a go-ahead.**
1. **Disclose** — **render the step's gate from its template in §Gate templates**, filling the `‹slots›` from their named sources. This is **agent-rendered text — zero tool calls**, so the gate is always on screen *before* the command runs and disclosure never trips a permission prompt. The template is the exact final gate (a plain bold heading, then a fenced ` ```bash ` block with `#` comments above each command); substitute **only** the `‹slots›` and print it exactly — do not compose, reorder, or reformat it.
2. **Run it** — run the real `scripts/provision.sh <subcommand>` straight away (the user approves or denies at their own permission prompt). **Parse the `=== RESULT ===`** on stdout. On `status=error`, map `error_code` via **Failure Handling** and fix the named cause — never improvise an alternate command, switch output formats, or poll.
3. **Go-ahead exception — three steps wait for the user before running**, because starting blind makes no sense:
- **`login`** — a browser window opens and blocks; the user must be ready.
- **`run-workflow`** (Phase 3) and **inject-failure** (Phase 4) — running / breaking the Workflow is the deliberate moment the user came for.
For these, after rendering the gate, append two choices and **wait** — `1. <action> / 2. Chat about this`, where the action verb is step-specific:
```
1. <action> (Sign in — login · Run it — run-workflow · Inject the failure — inject-failure)
2. Chat about this
```
`1` → run it. `2. Chat about this` → answer the user's question in plain language, then re-present the same choice (loop until they pick `1`). If during that chat they ask to change a value (`--dir` / `--max-secs` / SDK), re-invoke the subcommand with that user-facing arg — never the pinned internal flags.
**A state-changing command that isn't a `provision.sh` subcommand** (so it has no template in §Gate templates — e.g. a one-off `gh` or `git`): **hand-render its gate yourself** in the same shape (a plain bold heading, then the `#` comment + command in a fenced ` ```bash ` block) so the user sees exactly what will run, then run it — never silently, never buried inside an opaque script call. (This skill's normal flow has none: all `git` runs inside `provision.sh scaffold`, and it uses no `gh`.)
**Give every real `scripts/provision.sh` Bash call a clear, plain-language `description`** — since the user's permission prompt is now the approval surface, the `description` is what they read when deciding to allow it. Never a bare "Run script", and **name material side effects**: e.g. `Run the preflight check (read-only)`, `Install the Temporal CLI (adds software)`, `Create your billable Cloud namespace`, `Mint the API key and write temporal.toml`. (Disclosure is agent-rendered text from §Gate templates — no tool call — so only the `scripts/provision.sh` runs need a description.)
**Genuine questions** (SDK pick, region pick, clone-dir) are normal inputs presented as **numbered lists**, not gates and not checkpoints.
**Everything you print is a template — fill the slots, add nothing else.** Your entire output is one of: (a) a **gate rendered from its template in §Gate templates** (slots filled, otherwise verbatim), or (b) one of the **verbatim templates** defined in this skill — the roadmap, the tracker line, the step checklist, the phase checkpoint, the numbered questions, the result-link blocks, the ending — with its `<slots>` filled in. **Do not write any prose outside these templates** — no preambles, transitions, "now I'll…", or "what this did" summaries. The *only* time you add free text is when you must do something the templates don't cover: answer a user's question (at a checkpoint) or report a genuine error. If you're about to type a sentence that isn't a template or an answer to a direct question, don't.
**Exceptions / hard limits — the only "don'ts":**
- **Gate before run — never call a `provision.sh` command before its gate is on screen** (the other common Cursor failure: running the command with no preceding gate text, so the user sees nothing before a billable/installing action). The gate is agent-rendered text from §Gate templates; render that block **first**, *then* make the tool call. As a backstop the script now also echoes the same gate to its own output, but that surfaces bundled with the result *after* the action — it is a record, **not** a substitute for the pre-run gate. Order is always: render the gate, then run.
- **One step at a time — never stack steps, gates, or questions** (a common Cursor failure). Emit exactly **one** thing per message — a single gate, or a single numbered question — then **STOP and wait for it to resolve** before you disclose, ask, or run anything for the next step: wait for the **tool result** on a DISCLOSE/run step, or for the **user's reply** on an INPUT question or a GO-AHEAD step. Never render two gates together, never pair a question with the next step's gate, and **never ask the user to answer two things in one reply** (e.g. *"reply with your manager choice **and** whether to sign in"*). Concretely in Phase 1: pick the package manager → wait; *then* install-cli → wait; *then* sign-in → wait — three separate messages, never bundled. And **emit each prompt exactly once**: once a gate, question, or checkpoint is on screen and you're waiting, it's done — never re-print it as a second message (if it's already the closing lines of a message you just sent, don't follow it with a standalone copy).
- **Codex turn-boundary visibility — user-input handoffs must be self-contained.** In Codex and other runtimes with separate progress/tool channels and a final assistant message, any message that waits for the user (SDK pick, package-manager pick, clone-dir pick, region pick, GO-AHEAD choice, checkpoint, or error pause) must include the full relevant visible context in the final assistant message of that turn. Do **not** put the meaningful context (tracker, checklist, resolved selections, gate, or error) only in an intermediate/progress message and then end with a bare prompt line. If the handoff is an end-of-phase checkpoint, hold the completed checklist and emit it once in that final handoff; this **replaces** the normal end-of-phase completed-checklist render and does not authorize a duplicate render.
- **No prose narration** (the #1 historical failure, esp. on Codex). Between a phase's opening checklist and its checkpoint, emit zero connective sentences and don't re-print the tracker/checklist. Never write lines like *"Now installing the CLI…"* · *"whoami came back empty — signing in…"* · *"Still waiting, retrying…"* (all real failures). Retries / polls / readiness-waits inside one confirmed call are **silent**. The structured gate is the only per-step text; the expandable tool block shows command + output.
- **Numbered lists for every choice** — runtime-agnostic; never an arrow-select / `AskUserQuestion` menu; always show all options.
- **No Skip** — a go-ahead step's choices are only `1. <action> / 2. Chat about this` (every step is required; "Chat about this" never skips it — it answers a question, then re-presents). Don't print a "no skip" note.
- **Disclose in full.** A bundled subcommand (e.g. `scaffold` = clone + deps) gets **one** gate, but its GATE block shows **all** its commands. Don't unbundle into per-`temporal` gates; don't hide what it runs.
- **Never edit this skill's files** — invoke `scripts/provision.sh` as shipped; it's pinned to run unchanged on every platform (macOS bash 3.2). Reformatting/"tidying" its punctuation, quoting, regexes, or flags is forbidden. The only file you change on disk is the user's `temporal.toml`, via the script. If a flag has genuinely drifted (script returns `status=error`), stop and report it as a one-line maintenance note — don't fix it mid-run.
- **Already-satisfied prerequisite** → render its checklist item as `- [x] <thing> — already present, skipped` (don't fake-install it). **Exception: the Temporal CLI.** When the CLI is already present the Install-CLI step *updates* it to the latest (PE-79), so render that step as updated/up-to-date, never "skipped" — see the Install-CLI flow step.
- **Secret carve-out** (below) overrides disclosure for the API-key token.
</output-contract>
## Steps — the flow (the spine)
<steps>
The whole run in order. Tier legend (full mechanics in the output contract above): **DISCLOSE** = render the gate, then run (the user's permission prompt is the approval); **GO-AHEAD** = render, then append `1. <action> / 2. Chat about this` and wait (only the three deliberate steps); **INPUT** = a numbered question (no script). Each step is one `scripts/provision.sh` subcommand unless noted. "On-error" lists the `error_code`s to map via Failure Handling.
| # | Phase | Step | Tier | Subcommand | Emits | On-error |
|---|-------|------|------|------------|-------|----------|
| 1 | 1 | Choose SDK | INPUT | — (numbered list) | sdk | — |
| 2 | 1 | Preflight | DISCLOSE | `preflight --sdk` | `config_path`,`warnings`,`stray_env` | `config-dir-unwritable` |
| 3 | 1 | Detect tools + pick manager | DISCLOSE (+ INPUT if >1 manager) | `detect-tools --sdk` | `default`,`managers`,`discrepancies` | `version-too-old` (advisory) |
| 4 | 1 | Install / update CLI | DISCLOSE | `install-cli` | `status` (`ok`); `update` (`updated`/`up-to-date`/`skipped`/`failed`) | `brew-missing`,`manual-install` |
| 5 | 1 | Sign in | **GO-AHEAD** | `login` | `identity` | `login-failed`,`not-authenticated` |
| 6 | 1 | List + pick region | DISCLOSE + INPUT | `regions` | region list | `cloud-unreachable` |
| 7 | 2 | Start namespace (async) | DISCLOSE | `start-namespace --sdk --region` | `namespace_name` | `create-rejected` |
| 8 | 2 | Choose clone dir | INPUT | — (1=default / 2=Edit) | dir | — |
| 9 | 2 | Scaffold the app | DISCLOSE | `scaffold --sdk [--manager] [--dir]` | `repo_path`,`manager` | `clone-failed`,`unknown-sdk`,`manager-not-found`,`unsupported-manager` |
| 10 | 2 | Await namespace (join) | DISCLOSE | `await-namespace --name` | `namespace_handle`,`address` | `namespace-timeout`,`namespace-not-provisioning`,`handle-not-found` |
| 11 | 2 | Create key + save config | DISCLOSE | `create-key --handle --address` | `key_id` (token never printed) | `key-empty`,`key-limit-reached`,`no-json-parser`,`manual-key-needed` |
| 12 | 2 | Verify config | DISCLOSE | `verify-config` | — | `profile-missing` |
| 13 | 3 | Await auth | DISCLOSE | `await-auth` | `auth_ready` | `auth-timeout`,`key-expired` |
| 14 | 3 | Run the Workflow | **GO-AHEAD** | `run-workflow --sdk --dir` | `workflow_status`,`workflow_id`,`run_id` | `worker-unauthorized`,`worker-not-polling`,`worker-start-failed`,`workflow-failed`,`workflow-not-submitted`,`workflow-timeout`,`precompile-failed` |
| 15 | 4 | Inject failure + recover | **GO-AHEAD** | `run-workflow … --demo-failure transient` | same as 14 | same as 14 |
Phase bodies below add only the human nuance the table can't (region-pick guardrails, KeyId-vs-secret labeling, the result links). The exact command of any step comes from its gate — render it from the `‹sub›` template in §Gate templates (slots filled), don't hand-write it.
</steps>
### Secret-handling carve-out (overrides command disclosure)
The output contract says disclose the real command. **The API-key steps are the exception.** The `eyJ…` token must never be reprinted, logged, rendered in a diff, or passed as an argv (a rendered diff is the one exposure that leaves the local machine). For the key-capture and TOML-write actions:
- Show the friendly label and a **redacted** form of the command — e.g. `api_key = "eyJ…(captured, not shown)"`.
- Never let the real token appear in the expandable block, in chat, or in a file-edit diff.
- The **KeyId** (e.g. `JW4LO…`) is *not* secret and may be shown. See Phase 2 for the KeyId-vs-secret distinction.
- **Never read, `cat`, `grep`, or open `temporal.toml` (or any key-capture file) with the Read/Edit/Update tool.** The file holds the `eyJ…` token, so *any* read of it surfaces the secret into this transcript — this is the most common accidental leak. To confirm the profile, use **only** `scripts/provision.sh verify-config` (it lists profile *names*, never the key value).
- **Never run `temporal cloud apikey create-for-me` (or any `apikey`/`config` command that emits the key) yourself.** Only `scripts/provision.sh create-key` mints and stores the token — it redirects the one-time secret straight into the locked file. Run the raw CLI by hand and it prints the token to the terminal, into this output.
## Execution model — drive the bundled script, don't hand-roll the CLI
The variance-prone work — installing the CLI, signing in, listing regions, creating the namespace, minting the API key, and writing the client-config TOML — is owned by a bundled script: **`scripts/provision.sh`**. **Invoke it and parse its result block; do not reassemble these `temporal cloud` commands yourself.** That is what makes a run deterministic: the flags are pinned in one place, the retry / auth-recheck / "read the handle from the create output" logic is baked in, and the API-key token is written straight into the locked TOML by the script — so it never enters your context and can never leak into a rendered diff.
Each operation prints one delimited block on **stdout** — parse *that*, not the prose:
```
=== RESULT ===
status=ok # ok | error | skipped
<key>=<value> # operation-specific, e.g. namespace_handle=…, address=…, key_id=…
=== END ===
```
Human-readable progress goes to **stderr** (it shows in the expandable tool block — the teaching surface). On `status=error` the block carries `error_code` + `message` — map it via **Failure Handling** and fix the named cause (never improvise, switch output formats, or poll — as the output contract requires).
The flow steps — subcommand, tier, and error codes — are the **Steps spine table above** (single source of truth). The RESULT keys each emits:
- `preflight` → `os`, `config_path`, `cli_installed` (drives Install-CLI: install if absent, update if present), `warnings`, `stray_env`
- `detect-tools` → `default`, `managers`, `versions`, `discrepancies`
- `install-cli` → `status` (`ok`) + `update` (`updated`/`up-to-date`/`skipped`/`failed` when present; `skipped`/`failed` still proceed with the working CLI) + `reason` (`brew-missing`/`unsupported-os`, present only alongside `update=skipped`) · `login` → `identity` · `regions` → raw list on stderr (you recommend, user picks)
- `start-namespace` → `namespace_name` · `scaffold` → `repo_path`, `manager` · `await-namespace` → `namespace_handle`, `address`
- `create-key` → `key_id` (token never printed) · `verify-config` → profile names only · `await-auth` → `auth_ready`
- `run-workflow` → `workflow_status` (`COMPLETED`), `workflow_id`, `run_id`, `task_queue` (add `--demo-failure transient` for Phase 4)
**Utility subcommands (not in the main flow):** `preview <sub> [args]` (emits the `=== GATE ===` block + `cmd_N` + resolved params; side-effect-free — **maintenance/testing only, not used in the flow**: gates are rendered from §Gate templates, never by calling a script) · `provision-and-scaffold` (the older bundled namespace+clone+deps call — superseded by start/scaffold/await, kept as a fallback) · `install-deps` (re-install / switch manager) · `clone` (clone only) · `repair-config` (strip duplicate `cloud-setup` blocks; keeps `default`) · `cleanup-info` (prints the teardown commands, never runs them).
**Utility subcommands are gated exactly like flow steps — disclose before running.** "Not in the main flow" means *don't run them as routine steps*, **not** that they skip disclosure: if you ever invoke one (`install-deps` to switch a manager, `repair-config` to fix a duplicate profile, `clone`, etc.), render its gate first (derive it from the `scaffold`/`install-cmd` shapes in §Gate templates — these utilities have no dedicated template). **And don't improvise them into the flow:** the main steps already cover the work (`scaffold` clones *and* installs dependencies — never add an extra `install-deps` "to confirm deps resolve", and never narrate doing so).
The script is the **single source of truth for CLI flags** and is **read-only during a run** — invoke it as shipped, never edit it; if a prerelease flag has genuinely drifted (`status=error`), stop and report it as a one-line maintenance note (per the output contract), never fix it mid-setup. The wizard layer (tracker, checklists, checkpoints, the no-narration rules above) is still yours; only the imperative CLI work lives in the script.
## Per-command gate — disclose, then run
This setup runs real commands that **create billable Cloud resources and install software on the user's machine** — which is why the disclose-then-run loop and the three **go-ahead** steps (`login`, `run-workflow`, inject-failure), both defined in the output contract above, matter here. This section pins the **exact shape** of the gate you render.
**Deterministic backstop (don't rely on it):** every effectful `provision.sh` subcommand now also echoes its own gate to stderr (the tool block) *before* it acts, so the run is self-documenting even if you forget the chat-side gate. This is a safety net — it surfaces bundled with the result, *after* the action — so it never replaces rendering the gate first. Always render the §Gate-template gate, then run. (Disable only for tests via `TCLOUD_DISCLOSE=0`.)
**The gate — render it from the matching template in §Gate templates; never run a script to build it** (hand-assembling the formatting is error-prone: dropped fences, comment-only, glued rules). Fill **only** its `‹slots›` and print it exactly; rendering is **agent text — zero tool calls**. For a **go-ahead** step, append the numbered choices **stacked one per line** and wait. The gate's shape — a plain bold heading, then a fenced ` ```bash ` block with `#` comments above each command — looks like this (a disclose step — render it, then run):
````
**Installing the Temporal CLI**
```bash
# install the prerelease Temporal CLI via Homebrew (adds software to your machine)
brew install temporalio/prerelease/temporal-cloud
```
````
A multi-step subcommand shows each underlying command, one per step — a terse `#` note **followed by the actual command** (never a comment on its own). Still a disclose step — render, then run:
````
**Creating your namespace & downloading the sample app**
```bash
# 1. create your Cloud namespace — billable; provisions on Temporal's servers (~a few min)
temporal cloud namespace create --name <name> --region aws-us-east-1 \
--api-key-auth-enabled --retention-days 30 --auto-confirm
# 2. clone the Cloud-ready sample
git clone --branch money-transfer-project-cloud-setup --single-branch \
<repo-url> money-transfer-project-template-python
# 3. install dependencies (pip)
cd money-transfer-project-template-python \
&& python3 -m venv env && source env/bin/activate \
&& python -m pip install -q temporalio
```
````
The run step is a **go-ahead** step — it shows the real Worker + starter commands (**the commands themselves, not just their `#` labels**) and waits for the choice:
````
**Run your first Workflow**
```bash
# Worker — runs in the background, polls the task queue, stopped when done
cd money-transfer-project-template-python && source env/bin/activate && python run_worker.py
# starter — submits the Workflow, waits for it to reach COMPLETED, then exits
cd money-transfer-project-template-python && WORKFLOW_ID=money-transfer-demo python run_workflow.py
```
1. Run it
2. Chat about this
````
**A few rules these examples encode** (everything else is in the output contract above — don't restate it):
- **A `#` comment above every command — never a comment alone, never a bare command.** The real commands must appear (never just `# Worker` / `# starter` with the commands missing). Keep each comment to a few words; it's both a label and a one-line lesson for a newcomer.
- **Name material side effects in the relevant comment** — `# … - billable`, `# … (adds software to your machine)`, `# mint key + write the cloud-setup profile to temporal.toml`. The §Gate templates already encode this; render them verbatim.
- **`create-key` secret carve-out:** its GATE block shows the mint command **without** the token (captured straight into the locked TOML) — render as-is; never a token, never a redacted diff.
## Gate templates
These are the **verbatim source** for every step's gate — agent-rendered text (zero tool calls), so only the effectful `scripts/provision.sh <sub>` run ever prompts.
**Hard rule — render the matching template verbatim.** Substitute **only** the `‹slots›`, never add/drop/reorder/reformat lines or fences; keep every static character (headings, `#` comments, the ` ```bash ` fence, spacing) byte-for-byte. Each `‹slot›`'s value comes **only** from its named source in "Filling the slots" below — never from memory, never improvised. The result is exactly what `scripts/provision.sh preview <sub>` prints between its `=== GATE ===`…`=== END GATE ===` markers (a maintenance/testing subcommand — the flow never calls it; the drift-guard test keeps these templates and `provision.sh`'s runtime commands in sync).
### Filling the slots
| Slot | Source |
|------|--------|
| `‹sdk›` | the user's SDK pick (Phase 1) |
| `‹region›` | the user's region pick (Phase 1, from the `regions` list) |
| `‹manager›` | `detect-tools` RESULT `default`, or the user's override when they pick a non-default manager |
| `‹clone-dir›` | the user's clone-dir pick (Phase 2); default = the repo basename for `‹sdk›` (see SDK command reference) |
| `‹namespace-name›` | `start-namespace` RESULT `namespace_name` |
| `‹namespace-handle›` | `await-namespace` RESULT `namespace_handle` |
| `‹address›` | `await-namespace` RESULT `address` |
| `‹repo-url›` | SDK command reference, keyed by `‹sdk›` |
| `‹install-cmd›` | SDK command reference, keyed by `‹sdk›`/`‹manager›` |
| `‹worker-cmd›` | SDK command reference, keyed by `‹sdk›` |
| `‹starter-cmd›` | SDK command reference, keyed by `‹sdk›` |
| `‹task-queue›` | SDK command reference, keyed by `‹sdk›` |
| `‹runtime›` | SDK command reference, keyed by `‹sdk›` (the version-probe binary) |
| `‹probe-bins›` | SDK command reference, keyed by `‹sdk›` (the manager binaries to look for) |
`‹key-id›` is **never** shown in any gate — it appears only in the final summary (Phase 2 checklist). There is **no token slot**: the `eyJ…` token never appears in a template.
### SDK command reference
Source of truth = `scripts/provision.sh`. Keyed by `‹sdk›` (and `‹manager›` where it varies):
| `‹sdk›` | `‹repo-url›` | `‹runtime›` | `‹task-queue›` |
|---------|-------------|-------------|----------------|
| python | https://github.com/temporalio/money-transfer-project-template-python | python3 | TRANSFER_MONEY_TASK_QUEUE |
| go | https://github.com/temporalio/money-transfer-project-template-go | go | TRANSFER_MONEY_TASK_QUEUE |
| ts | https://github.com/temporalio/money-transfer-project-template-ts | node | money-transfer |
| java | https://github.com/temporalio/money-transfer-project-java | java | MONEY_TRANSFER_TASK_QUEUE |
| dotnet | https://github.com/temporalio/money-transfer-project-template-dotnet | dotnet | MONEY_TRANSFER_TASK_QUEUE |
| ruby | https://github.com/temporalio/money-transfer-project-template-ruby | ruby | money-transfer |
`‹clone-dir›` default (repo basename) = the trailing path segment of `‹repo-url›` (e.g. python → `money-transfer-project-template-python`, java → `money-transfer-project-java`).
**Managers (`default` first) + `‹probe-bins›` (the binaries `detect-tools` looks for):**
| `‹sdk›` | managers | `‹probe-bins›` |
|---------|----------|----------------|
| python | pip (default), uv | python3 uv |
| ts | npm (default), pnpm, yarn | npm pnpm yarn |
| go | go | go |
| java | maven | mvn |
| dotnet | dotnet | dotnet |
| ruby | bundler | bundle |
(manager → its probe binary: pip→python3, uv→uv, npm→npm, pnpm→pnpm, yarn→yarn, go→go, maven→mvn, dotnet→dotnet, bundler→bundle. `‹probe-bins›` for an SDK is the space-joined probe binaries of all its managers, in the order above.)
**`‹install-cmd›`** (keyed by `‹sdk›`/`‹manager›`):
| `‹sdk›`/`‹manager›` | `‹install-cmd›` |
|---------------------|-----------------|
| python/pip | `python3 -m venv env && . env/bin/activate && python -m pip install -q temporalio` |
| python/uv | `uv venv env && . env/bin/activate && uv pip install -q temporalio` |
| ts/npm | `npm install` |
| ts/pnpm | `pnpm install` |
| ts/yarn | `yarn install` |
| go/go | `go mod download` |
| java/maven | `mvn -q -DskipTests dependency:resolve` |
| dotnet/dotnet | `dotnet restore` |
| ruby/bundler | `bundle install` |
**`‹worker-cmd›` / `‹starter-cmd›`** (keyed by `‹sdk›`):
| `‹sdk›` | `‹worker-cmd›` | `‹starter-cmd›` |
|---------|----------------|-----------------|
| python | `source env/bin/activate && python run_worker.py` | `source env/bin/activate && python run_workflow.py` |
| go | `go run worker/main.go` | `go run start/main.go` |
| ts | `npm run worker` | `npm run client` |
| java | `mvn -q compile exec:java -Dexec.mainClass=moneytransferapp.MoneyTransferWorker -Dorg.slf4j.simpleLogger.defaultLogLevel=warn` | `mvn -q compile exec:java -Dexec.mainClass=moneytransferapp.TransferApp -Dorg.slf4j.simpleLogger.defaultLogLevel=warn` |
| dotnet | `dotnet run --project MoneyTransferWorker` | `dotnet run --project MoneyTransferClient` |
| ruby | `ruby worker.rb` | `ruby starter.rb` |
The scaffold install-comment also varies by where deps land — keep the comment exactly as the template shows for that `‹sdk›`/`‹manager›` (python/ts say "inside the repo"; go/java/dotnet/ruby say "GLOBAL, outside the repo …"). The worked example and templates below carry the right wording per SDK; for non-python SDKs use the install comment from `scripts/provision.sh preview scaffold --sdk ‹sdk›` if you ever need to re-verify it (maintenance only).
### Templates
**Phase 1 — `preflight`** (static):
````
**Checking your environment**
```bash
# check git / jq / brew are available (read-only, local)
command -v git jq brew
# check the Temporal config directory is writable (so temporal.toml can be saved)
touch "$(dirname "<config_path>")/.probe" && rm -f "$(dirname "<config_path>")/.probe"
# flag any stray TEMPORAL_* env vars that would override your saved config
env | grep '^TEMPORAL_' || true
```
````
**Phase 1 — `detect-tools`:**
````
**Detecting your local tools**
```bash
# detect which package managers are installed for ‹sdk› (read-only, local)
command -v ‹probe-bins›
# read each tool's version to flag anything below the minimum
‹runtime› --version
```
````
**Phase 1 — `install-cli` (not installed):**
````
**Installing the Temporal CLI**
```bash
# install the Temporal CLI via Homebrew (adds software)
brew install temporalio/prerelease/temporal-cloud
```
````
**Phase 1 — `install-cli` (already installed — render this variant instead when the CLI is present):**
````
**Updating the Temporal CLI to the latest**
```bash
# BETA: the prerelease CLI has no real versions yet, so always pull the latest (adds/updates software)
brew upgrade temporalio/prerelease/temporal-cloud
```
````
**Phase 1 — `install-cli` (already installed, non-macOS — render this variant when the CLI is present and you're not on macOS; there's no prerelease tap to upgrade from):**
````
**Updating the Temporal CLI to the latest**
```bash
# already installed; update temporal-cloud manually from https://github.com/temporalio/cloud-cli/releases/latest
```
````
**Phase 1 — `login`** (static):
````
**Sign in to Temporal Cloud**
```bash
# open a browser to sign in (blocks until you finish)
temporal cloud login
# confirm the signed-in identity
temporal cloud whoami
```
````
**Phase 1 — `regions`** (static):
````
**Listing your Cloud regions**
```bash
# list the regions your account can use
temporal cloud region list
```
````
**Phase 2 — `start-namespace`:**
````
**Creating your Cloud namespace**
```bash
# create your Cloud namespace - billable; submits async, provisions server-side (~a few min)
temporal cloud namespace create --name <name> --region ‹region› --api-key-auth-enabled --retention-days 30 --auto-confirm --async
```
````
**Phase 2 — `create-namespace`** (the synchronous variant — rare; the flow uses `start-namespace` + `await-namespace`):
````
**Creating your Cloud namespace**
```bash
# create your Cloud namespace - billable; provisions server-side (~a few min)
temporal cloud namespace create --name <name> --region ‹region› --api-key-auth-enabled --retention-days 30 --auto-confirm
```
````
**Phase 2 — `scaffold` (clone + deps):**
````
**Downloading the sample app (clone + dependencies)**
```bash
# clone the Cloud-ready sample
git clone --branch money-transfer-project-cloud-setup --single-branch ‹repo-url› ‹clone-dir›
# install dependencies with ‹manager› ‹install-location-note›
(cd ‹clone-dir› && ‹install-cmd›)
```
````
`‹install-location-note›` per SDK (keep verbatim): python = `into a local venv (env/) inside the repo` · ts = `into node_modules inside the repo` · go = `into the shared Go module cache - GLOBAL, outside the repo (~/go/pkg/mod)` · java = `into the shared Maven cache - GLOBAL, outside the repo (~/.m2)` · dotnet = `into the global NuGet cache - GLOBAL, outside the repo (~/.nuget/packages)` · ruby = `into globally-installed gems - GLOBAL, outside the repo`.
**Phase 2 — `await-namespace`** (static):
````
**Waiting for the namespace to provision**
```bash
# poll until the namespace is ACTIVE — provisioning usually takes ~a few minutes (bounded)
temporal cloud namespace list --name <namespace-name> -o json
```
````
**Phase 2 — `create-key`** (secret carve-out — the token is captured to a 0600 file, never shown; **never** add a token slot):
````
**Creating your API key and saving the config**
```bash
# mint the key + write the cloud-setup profile to temporal.toml (token captured to a 0600 file, never printed)
temporal cloud apikey create-for-me --display-name money-transfer-cloud-setup-<random> --expiry-duration 25h --auto-confirm -o json
```
````
**Phase 2 — `verify-config`** (static):
````
**Verifying the saved config**
```bash
# list the cloud-setup profile fields (api_key redacted, never shown)
temporal --profile cloud-setup config list
```
````
**Phase 3 — `await-auth`** (static):
````
**Waiting for the API key to be accepted**
```bash
# poll an authorized call until the new key is accepted (bounded ~90s)
temporal --profile cloud-setup workflow list --limit 1
```
````
**Phase 3 — `run-workflow` (clean run):**
````
**Run your first Workflow**
```bash
# Worker - runs in the background, polls the task queue, stopped when done
(cd ‹clone-dir› && ‹worker-cmd›)
# starter - submits the Workflow, waits for COMPLETED, exits (the run can take a minute or two)
(cd ‹clone-dir› && WORKFLOW_ID=money-transfer-demo ‹starter-cmd›)
```
````
**Phase 4 — `run-workflow … --demo-failure transient` (inject + recover):**
````
**Run the recovery Workflow (inject a failure)**
```bash
# Worker - runs in the background, polls the task queue, stopped when done
(cd ‹clone-dir› && DEMO_FAILURE=transient ‹worker-cmd›)
# starter - submits the Workflow, waits for COMPLETED, exits (the run can take a minute or two)
(cd ‹clone-dir› && WORKFLOW_ID=money-transfer-demo-recovery ‹starter-cmd›)
```
````
(Utility subcommands — `clone`, `install-deps`, `provision-and-scaffold`, `repair-config` — are **rare** and have no dedicated template; derive their gate from the `scaffold`/`install-cmd` shapes above if you ever invoke one.)
### Worked example — the `scaffold` gate for Java, fully filled
`‹sdk›` = java → `‹repo-url›` = `https://github.com/temporalio/money-transfer-project-java`, `‹clone-dir›` (default) = `money-transfer-project-java`, `‹manager›` = `maven`, `‹install-cmd›` = `mvn -q -DskipTests dependency:resolve`, `‹install-location-note›` = `into the shared Maven cache - GLOBAL, outside the repo (~/.m2)`. The rendered gate:
````
**Downloading the sample app (clone + dependencies)**
```bash
# clone the Cloud-ready sample
git clone --branch money-transfer-project-cloud-setup --single-branch https://github.com/temporalio/money-transfer-project-java money-transfer-project-java
# install dependencies with maven into the shared Maven cache - GLOBAL, outside the repo (~/.m2)
(cd money-transfer-project-java && mvn -q -DskipTests dependency:resolve)
```
````
## Start — show the plan, then begin Phase 1
This skill begins when the user invokes it (e.g. `/temporal-cloud-setup`) or asks to set up Temporal Cloud. On start, print the roadmap, then go straight into Phase 1 (whose first step is choosing the SDK). Do **not** narrate.
Print this opening block verbatim — the ⚠️ notice first, then the plan and the profile line:
```
> ⚠️ **Heads-up:** this creates real resources in your Temporal Cloud account — a namespace and an API key — which may incur cost.
**Let's get you set up on Temporal Cloud.** Four phases, end to end.
1. **Get set up** — install the CLI, sign in, and choose your SDK + region
2. **Download the sample app & create your API key** — create your Cloud namespace and clone the money-transfer app (already wired for Cloud) in parallel, then mint your API key
3. **Run your first Workflow** — start the Worker and run the money-transfer Workflow
4. **See Durable Execution** — break the transfer on purpose and watch Temporal recover it
I'll show each command before running it — and depending on your setup, your tool may ask you to approve it first.
Setting up for: **<OS> · SDK pending**
```
The ⚠️ Heads-up in that block is a notice, not a blocking gate — continue unless the user objects.
**Selections are collected up front, in Phase 1, while the user is most engaged** — SDK first (no prerequisite), then region right after sign-in (the live region list requires being logged in). Do not begin cloning before the SDK answer — the SDK selects which repo is cloned.
## Phase output envelope
This is the **single source of truth for per-phase formatting** — the phase bodies below supply only *content* (the intent sentence and the steps); this envelope supplies the *format*. Render every phase in this fixed order:
1. **Tracker line** at the top, marker advanced (see below).
2. **Intent sentence** — one short line on what this phase sets up and why it matters (gloss any Temporal term). No more than one line.
3. **Step checklist — once, all unchecked** (the phase plan). Then run each step (real tool call with a friendly `description`, or a genuine question) **without re-printing the checklist or tracker between steps**. Render an already-present prerequisite as `[x] … already present, skipped` (except the Temporal CLI, which updates when present — render it updated/up-to-date, not "skipped"). **No** `**What this did:**` summary and **no** `↪ Learn more:` link during the run.
4. **End-of-phase: completed checklist + checkpoint** — print the checklist once more with **every box checked**, show `**Phase N complete ✅**`, and **close that same message** with the numbered checkpoint prompt (its content is detailed in the next section — it belongs to *this* message, it is not a second message). In Codex-style runtimes, this entire block must be the final assistant message of the turn when you pause for the user; do not print it earlier as progress and then repeat or fragment it in the final response. No checkpoint after Phase 4 — go straight to the Ending.
The tracker is just the four phase markers and the phase counter — **no leading label** (don't prefix it with "Setup" or anything before the first marker). Legend: completed = `✅`, current = `🔵`, upcoming = `⚪` (a white dot — same filled-circle style as the blue current dot). **Bold the current step's name** (the one with the blue dot). Reprint it at the top of each phase, advancing one marker:
```
🔵 **Set up** · ⚪ App & API key · ⚪ Run · ⚪ Recover · phase 1/4
✅ Set up · 🔵 **App & API key** · ⚪ Run · ⚪ Recover · phase 2/4
✅ Set up · ✅ App & API key · 🔵 **Run** · ⚪ Recover · phase 3/4
✅ Set up · ✅ App & API key · ✅ Run · 🔵 **Recover** · phase 4/4
```
(Glyph notes: `⚪` white dot = upcoming and `🔵` blue dot = current — same filled-circle style, so the row reads as one consistent set. `✅` = the reliably-green completion mark; a green *circle*-with-check isn't a dependable cross-platform glyph, so stick with `✅`.)
## After-phase checkpoint (after Phases 1–3)
This documents the checkpoint prompt that the envelope's step 4 already emits as the closing lines of the completed-checklist message — **it is described here, not a second message to send.** It hands control back via a numbered prompt (not an arrow-select — numbered works on every runtime). **Do not explain anything proactively**; the checklist already told the story. At a turn boundary, the checkpoint handoff must be self-contained and emitted exactly once: tracker/checklist context + `**Phase N complete ✅**` + this prompt together in the final assistant message. The prompt is:
```
1. Continue
2. I have a question about this phase
Choose a number, or write your response.
```
- If they pick **2 (question):** give a short, plain-language explanation of what the phase just did, answer their typed question in newcomer-friendly language, then re-present the numbered prompt — `1. Continue` / `2. I have another question` / `Choose a number, or write your response.` — and loop until they proceed.
- Explanation is **on-demand only** — it appears solely when they pick the question option.
- A checkpoint follows Phases 1–3. **No checkpoint after Phase 4** — go straight to the Ending.
- Genuine blocking input (SDK pick, browser sign-in, region pick) is normal work, **not** this checkpoint (genuine input, not a checkpoint).
---
# The phases
```text
Phase 1 — Get set up: choose SDK · install CLI · sign in · choose region
Phase 2 — App & API key: create namespace + clone sample + install deps (parallel) · create API key · write config TOML
Phase 3 — Run your first Workflow: start the Worker · run the Workflow (success)
Phase 4 — See Durable Execution: inject a failure · watch Temporal retry & recover ──► then END
```
<phases>
## Phase 1 — Get set up
*Intent: get the tools on your machine, sign you in, and capture your choices — everything the rest of the run needs.*
Step checklist: `SDK chosen` · `Tools detected` · `CLI installed` · `Signed in` · `Region chosen`.
**Step — Choose your SDK** (genuine input, not a checkpoint): present the six as a **numbered list** (numbered list, runtime-agnostic) — Python, Go, TypeScript, Java, .NET, Ruby — and take a typed name/number. This selects which repo is cloned and the language of the local app. Echo the resolved profile line (`Setting up for: macOS · Python SDK`) once answered.
Right after the SDK pick, the **preflight** check (DISCLOSE — render, then run): render its gate from the `preflight` template (§Gate templates), then run `scripts/provision.sh preflight --sdk <sdk>`. Note its `cli_installed` flag — it drives whether the Install-CLI step below installs (absent) or updates the existing CLI (present). If its `stray_env` lists any `TEMPORAL_*` vars, tell the user they override the saved profile and ask them to unset them before continuing; surface other `warnings` (e.g. `brew-missing`, `no-json-parser`) only if they block a later step.
**Step — Detect local tools + choose your package manager** (DISCLOSE — render, then run). Render its gate from the `detect-tools` template (§Gate templates), then run `scripts/provision.sh detect-tools --sdk <sdk>` and read its RESULT. This adapts the setup to the user's machine, and it surfaces tooling problems **early** (here in Phase 1) instead of deep in Phase 2/3.
- **Package manager (Python & TypeScript only).** `managers` is the set of package managers **this sample supports** for the chosen SDK. When it lists more than one, ask **which of the sample-supported managers to use** — frame it that way ("the sample supports these — which should we use?"), **not** as "here's what's installed on your machine." Present them as a **numbered list** (runtime-agnostic) with the `default` marked — e.g. Python `1. pip (default)` / `2. uv`; TypeScript `1. npm (default)` / `2. pnpm` / `3. yarn`. The user confirms the default or overrides; carry the choice into Phase 2 as `scaffold --manager <m>`. **This is a question — ask it alone and wait for the answer; do not disclose install-cli or sign-in in the same message** (see "One step at a time"). For **Go / Java / .NET / Ruby** the sample has a single toolchain — state it (`Using: maven`) and **skip the prompt**.
- **Surface `discrepancies` HERE (fail-early), plain language:**
- `version-too-old:<tool>@<have>(min<min>)` — an **advisory** warning with remediation (e.g. *"Node 16 detected; 18+ recommended — consider upgrading, but I can proceed"*). Not a hard stop.
- `tool-missing:<runtime>` / `manager-not-found:<m>` — the runtime or chosen manager isn't installed. Offer another **sample-supported** manager from `managers`, or ask the user to install the tool, then re-run `detect-tools`.
- The default proposal is **deterministic** — the same machine yields the same default every run; nothing is persisted (no state file).
**Step — Install or update the unified Temporal CLI** (DISCLOSE — render, then run). It ships the `temporal cloud` command group (binary `temporal-cloud`). **Branch on preflight's `cli_installed`** (it used the same `temporal cloud` probe the installer does, so it's authoritative — don't re-check by calling `install-cli` just to confirm):
- **`cli_installed=true` → update it to the latest.** Render its gate from the `install-cli` (already-installed) template (§Gate templates), then run `scripts/provision.sh install-cli`. **BETA stopgap:** the prerelease CLI has no meaningful version numbers yet, so instead of a real "is it out of date?" check we always try to pull the latest from the `temporalio/prerelease` Homebrew tap; `install-cli` returns `status=ok` with `update`=`updated`/`up-to-date`. When it can't update — Homebrew missing, or a non-macOS host — it emits `update=skipped` and proceeds with the working CLI rather than failing (never yank a working install). Once the CLI ships real versions this becomes a genuine version check; we no longer skip a present CLI.
- **`cli_installed=false` → install it.** Render its gate from the `install-cli` (not-installed) template (§Gate templates), then run `scripts/provision.sh install-cli`. It installs via the `temporalio/prerelease` Homebrew tap on macOS, and returns `error_code=brew-missing` / `manual-install` with the fallback URL if it can't — do not auto-install Homebrew; relay the message and wait. (`install-cli` self-checks presence, so if it's actually already there it updates instead of installing.)
- (No prerelease disclaimer here — the ⚠️ notice at the top of the run already covers that.)
This CLI is separate from any local `temporal server start-dev`. The Cloud path does not start a local server.
**Step — Sign in** (**GO-AHEAD** — a browser opens, so wait for the user). Render its gate from the `login` template (§Gate templates), then present `1. Sign in / 2. Chat about this` and wait. On `2`, answer the question and re-present. On `1`, run `scripts/provision.sh login` — it runs `temporal cloud login` (which opens a browser on the user's machine and blocks until they finish) and then confirms with `whoami`. **Do not ask the user to run the command themselves** (no `! temporal cloud login` hand-off); you run the script, the user just completes the browser prompt. Tell them to complete the sign-in in the browser. On success the result block carries `identity`; on `error_code=login-failed`/`not-authenticated`, ask them to finish the browser login and re-run. If they're not part of a Cloud account yet, point them to https://temporal.io/get-cloud and pause.
**Step — Choose your region** (DISCLOSE — render, then run): render its gate from the `regions` template (§Gate templates), then run `scripts/provision.sh regions` (you're authenticated now). It prints the live region list; present it as a **numbered list** (runtime-agnostic), then apply these picking rules:
- **Recommend the nearest** — infer it from the system timezone/locale (e.g. `America/New_York` → suggest `aws-us-east-1`) and mark it `(recommended)`, but the user still picks.
- **Never accept a region from memory and never auto-select** — a valid-but-wrong region creates a persistent, billable namespace in the wrong place (and has set off internal alerts). Use the exact identifier the user picks (shape `<provider>-<region>`); it's passed to the create step next.
- **Fail fast — only accept a region that appears in the `regions` output** (it reflects what your account can actually use). If the user names one that isn't listed, don't pass it to create (a region your account lacks access to is the classic "it spun for minutes" trap) — re-show the list and have them pick a listed value.
- **Steer away from `unsupported_regions`** — the `regions` RESULT may list `unsupported_regions` (regions whose provider reads `UNKNOWN`, e.g. `azure-centralus`). On such accounts those **accept a namespace create but never provision it** — a phantom that stalls the next phase. Mark any listed there as `(may not be available on your account)`, **don't recommend it**, and steer to an AWS/GCP region. If the user insists, warn it may never provision (you'll catch it fast — see `namespace-not-provisioning`).
- **If the create is still rejected** for the region, see Failure Handling (`create-rejected`) — re-list and re-pick, never silently retry the same region.
**If the user asks at the checkpoint, explain (plain language):** installed the Temporal CLI, signed you in, and recorded your SDK + region. Your namespace itself gets created in the next phase — in parallel with setting up your app. (Docs: https://docs.temporal.io/cli)
## Phase 2 — App & API key
*Intent: create your Cloud namespace and download the sample app (clone + dependencies) in parallel, then mint your API key and save the connection config — everything needed to run Workflows on Cloud.*
Step checklist: `Namespace active` · `Sample cloned` · `Dependencies installed` · `API key created` · `Config saved`.
The namespace and the app are set up as **three separate, individually-gated steps** — so the **billable** namespace gets its own explicit approval, distinct from the benign clone. The namespace provisions on Temporal's servers (`--async`) **while** the app is cloned, so the parallelism is preserved without any background process having to survive across steps (identical on Claude Code, Codex, Cursor).
**Step — Start your namespace** (`start-namespace`, DISCLOSE — its own gate, then run; the user's permission prompt approves this **billable** create — make the `description` say so):
- **Fire-and-forget:** submits the create with `--async`, returns immediately, and provisions server-side while you download the sample app.
- **Read `namespace_name`** from the result — you pass it to `await-namespace` below. The handle isn't known yet (it resolves once provisioning completes).
- **Errors:** `create-rejected` (region/name).
**Step — Ask where to clone** (genuine input — **always ask, don't silently default**). Present a two-option numbered choice where **option 1 is the default clone path itself** (so the user sees and confirms the real path) and **option 2 is "Somewhere else"**:
```
Where should I clone the sample app?
1. ./money-transfer-project-template-python
2. Somewhere else
```
- **1** → clone to that path (omit `--dir`, or pass it explicitly — same result).
- **2 (Somewhere else)** → ask for the path, then pass it as `--dir`.
Show the concrete default for the chosen SDK in option 1 (e.g. `./money-transfer-project-template-ts` for TypeScript, `./money-transfer-project-template-go` for Go). Pass the chosen path as `--dir` to `scaffold` below.
**Step — Download the sample app** (`scaffold`, DISCLOSE — its own gate, then run, separate from the namespace):
- **Clones the cloud-ready sample + installs dependencies** — runs **while the namespace provisions**, threading the Phase 1 manager choice and the clone dir just chosen.
- **`--manager`/`--dir` are validated fail-fast before the clone** — a bad/uninstalled manager errors immediately (`manager-not-found` / `unsupported-manager`).
- **Read `repo_path` + `manager`.** Other errors: `clone-failed`, `unknown-sdk`.
- The cloned branch ships wired for Cloud — **no connection code to edit.**
**Step — Wait for the namespace** (`await-namespace --name <namespace_name>`, DISCLOSE — render, then run): render its gate from the `await-namespace` template (§Gate templates), then run `scripts/provision.sh await-namespace --name <namespace_name>`. Polls (via the exact `namespace list --name` filter) until the namespace is **ACTIVE**, then reads the handle + endpoint from that result. Read **`namespace_handle`** + **`address`**; carry both into the key step below. Errors: `namespace-timeout` (appeared but slow — re-run), `namespace-not-provisioning` (never appeared — bad region, re-run `start-namespace` elsewhere), `handle-not-found`.
For reference, the per-SDK repo mapping (the script selects the right one):
| SDK | Repository (branch `money-transfer-project-cloud-setup`) |
|------------|------------|
| Python | `https://github.com/temporalio/money-transfer-project-template-python` |
| Go | `https://github.com/temporalio/money-transfer-project-template-go` |
| TypeScript | `https://github.com/temporalio/money-transfer-project-template-ts` |
| Java | `https://github.com/temporalio/money-transfer-project-java` |
| .NET | `https://github.com/temporalio/money-transfer-project-template-dotnet` |
| Ruby | `https://github.com/temporalio/money-transfer-project-template-ruby` |
**How it connects (no edit needed):** all six SDKs load the **`cloud-setup`** profile from `temporal.toml` (env-config). The key step below writes it; nothing else is needed at run time.
**After `await-namespace` returns, go straight to the `create-key` step — no *phase checkpoint* between them, and don't re-print the checklist or tracker. But `create-key` still gets its own gate — its command is disclosed, then run (the user's permission prompt approves it, like every step); it mints your key *and* writes/replaces the `cloud-setup` profile in `temporal.toml`. The next checklist render is the completed one at the end of the phase.**
**Step — Create the key and save the config** (`create-key --handle <namespace-handle> --address <address>`, DISCLOSE — render, then run, using the values from the `await-namespace` step above). In one deterministic, secret-safe operation the script:
- **re-verifies auth** (`whoami`, re-prompting login if expired);
- **mints the key** with the pinned flags;
- **captures it via `-o json` into a `0600` temp file** so the `eyJ…` token never reaches stdout, your context, or a rendered diff;
- **writes a named `cloud-setup` profile** into the shared `temporal.toml` (address, namespace, api_key) **without touching `[profile.default]`** or its login session;
- **`chmod 600`s the file** and deletes the temp capture.
It returns only the **non-secret** `key_id` and the `config_path` — never the token. Then verify the profile (DISCLOSE — render, then run): render its gate from the `verify-config` template (§Gate templates), then run `scripts/provision.sh verify-config` to confirm the profile loads (it never prints the api_key value).
This is the **secret-handling carve-out** (§ above): render the action's label without the token; never reprint, log, argv-pass, or commit it.
- On `error_code=key-empty`/`not-authenticated`: an expired login — let the script re-prompt and retry once; don't switch output formats or poll.
- On `error_code=key-limit-reached`: the account is at its API-key cap — the mint was rejected at create time. Have the user delete stale keys (`temporal cloud apikey list`, then `temporal cloud apikey delete --key-id <id>` on old `money-transfer-cloud-setup-*` keys), then re-run `create-key`. Don't re-run login.
- On `error_code=no-json-parser`: install `jq` or `python3`, then re-run (the safe capture needs one).
- On `error_code=manual-key-needed`: automatic capture failed and there was no terminal to paste into. Ask the **user** to paste the one-time key and re-run `create-key` from a context with a terminal — the script reads the paste *hidden*, straight into the locked file. **Never** have the user paste the key into the chat, and never paste it yourself.
- The secret is shown only once and cannot be retrieved later; if it's truly lost, mint a new key (re-run this step) — don't try to recover the old value.
**KeyId vs. secret — not the same thing.** The `key_id` the script returns (e.g. `JW4LO…`) is a non-secret *identifier* — the handle used to revoke the key later (`temporal cloud apikey delete --key-id <KeyId>`); fine to show. **Only the `eyJ…` token is the credential** — never reprint, log, or commit it.
**When you check off "API key created," label the `key_id` so a newcomer can't mistake it for the secret.** A bare 32-char `key_id` on its own reads like a leaked key and alarms people. Render it with an explicit non-secret tag, and make the config line state the secret was stored (not shown) — e.g.:
```
- [x] API key created — key id JW4LO… (non-secret identifier; used to revoke the key)
- [x] Config saved — ~/Library/Application Support/temporalio/temporal.toml (chmod 600; secret stored, not shown)
```
Never render the `key_id` bare and unlabeled, and never put the `eyJ…` token on either line.
**Env vars override the profile.** Env-config gives `TEMPORAL_*` env vars **higher precedence** than the TOML profile — the Phase 1 preflight flags any stray `TEMPORAL_ADDRESS`/`TEMPORAL_NAMESPACE`/`TEMPORAL_API_KEY`. If `stray_env` was non-empty, make sure they're unset in the run shell or they'll override the saved profile at run time.
**If the user asks at the checkpoint, explain (plain language):** created your Cloud namespace and, in parallel, cloned a small money-transfer app in your language and installed its dependencies. Then minted an API key for your namespace (`<namespace-handle>`) and saved the connection — address, namespace, key — into a locked `cloud-setup` profile your app reads at run time. (Docs: https://docs.temporal.io/develop · https://docs.temporal.io/cloud/api-keys)
## Phase 3 — Run your first Workflow
*Intent: prove the whole setup works by running a real Workflow on Temporal Cloud.*
Step checklist: `Worker running` · `Workflow completed`.
The app connects from config and Phase 2 supplied the credentials — this phase is **run-only**. All SDKs read the `cloud-setup` profile from `temporal.toml`; confirm it's present with `scripts/provision.sh verify-config` if you haven't already.
**Connect-readiness is handled for you — do not hand-roll the Worker.** The Worker start, the wait-until-it's-polling, the starter, and the Worker teardown all live in **one deterministic call** (`run-workflow`), so you never launch `nohup … &` / `ps` / `pgrep` yourself — that is exactly the flaky, noisy step this replaces. Two things still matter:
1. A freshly-minted API key isn't accepted by the data plane *immediately* — wait on `await-auth` first.
2. If the connect still fails, the namespace endpoint (`<handle>.tmprl.cloud:7233`) is correct — do **not** switch to a regional endpoint, re-mint the key, or edit the profile (per temporalio/documentation#4733). The failure is readiness, not the endpoint.
Run two script calls, both using `repo_path` from Phase 2 (the script handles per-SDK run commands and Python venv activation internally — you don't):
1. **Wait for auth** (`await-auth`, DISCLOSE — render, then run): render its gate from the `await-auth` template (§Gate templates), then run `scripts/provision.sh await-auth` and wait for `auth_ready=true`. On `error_code=key-expired`, the key is permanently rejected (commonly a next-day re-test against a key that auto-expired in ~25h) — re-run `create-key` (see Failure Handling); on `auth-timeout`, read the appended redacted CLI stderr, wait, and re-run.
2. **Run the Workflow** (`run-workflow --sdk <sdk> --dir <repo_path>`, **GO-AHEAD** — the deliberate moment): render its gate from the `run-workflow` (clean run) template (§Gate templates), append `1. Run it / 2. Chat about this` (on `2`, answer, then re-present), and on `Run it` one synchronous call starts the Worker, waits until it's polling (Temporal's API, not the OS process table), runs the starter, and stops the Worker. Wait for `workflow_status=COMPLETED`; it emits the run's **`workflow_id`** + **`run_id`** (don't run `workflow list` yourself). On `worker-unauthorized`, re-run `await-auth` then `run-workflow` (see Failure Handling).
3. **Show the success link, then confirm the win.** `run-workflow` already verified `COMPLETED`. Surface the run's **timeline** page on its own bare line (bare URL so the terminal auto-linkifies it — no backticks/fence), using the `workflow_id` and `run_id` from the RESULT block:
**View your completed Workflow on Temporal Cloud:**
https://cloud.temporal.io/namespaces/<namespace-handle>/workflows/<workflow-id>/<run-id>/timeline
That's your first Workflow on the Cloud — no separate `workflow describe` needed.
*Fallback — only if `run_id` came back `unknown` (the data plane was briefly lagging): show the Workflow **detail** page instead, rendered exactly like the link above — bold lead-in, then a bare URL alone on its own line (no backticks) so it auto-linkifies:*
**View your completed Workflow on Temporal Cloud:**
https://cloud.temporal.io/namespaces/<namespace-handle>/workflows/<workflow-id>
*Never fall back to the bare `…/workflows` namespace list.*
At the **Phase 3 checkpoint**, lead with the win — `**Phase 3 complete ✅ — your first Workflow ran clean.**` — then the standard numbered prompt (`1. Continue` / `2. I have a question about this phase`). Phase 4 itself frames and triggers the failure injection, so this checkpoint stays a plain **Continue**.
**If the user asks at the checkpoint, explain (plain language):** started a Worker that polls your Cloud namespace, ran the money-transfer Workflow against Temporal Cloud, and confirmed it reached `COMPLETED` — your first Workflow on the Cloud. (Docs: https://docs.temporal.io/workflows)
## Phase 4 — See Durable Execution (inject a failure)
*Intent: break the transfer on purpose and watch Temporal retry and recover it — the durable-execution payoff. This phase is **required**; the user reaches it via **Continue** at the Phase 3 checkpoint, then triggers the break via the prompt below.*
Step checklist: `Failure injected` · `Recovered & completed`.
**Open by framing the break, then let the user trigger it.** After the tracker + intent + unchecked checklist, explain in one or two plain sentences what's about to happen — *we'll run the same transfer again, but force the deposit to fail on its first attempts, then watch Temporal automatically retry until it succeeds* — then present a numbered prompt so the user actively triggers it:
```
1. Inject the failure
2. Chat about this
```
On `1`, run the inject step below; on `2`, answer the question and re-present. This is a genuine engagement point (genuine input, not a checkpoint), not a checkpoint.
> The `DEMO_FAILURE` toggle is shipped on the `money-transfer-project-cloud-setup` branches (the deposit activity reads it), verified end-to-end on real Cloud for all six supported SDKs.
**Step — Inject the failure.** The `1. Inject the failure` choice above is this step's go-ahead — so disclose, then run (don't add a second go-ahead prompt):
- **Disclose:** render its gate from the `run-workflow` (inject + recover) template (§Gate templates).
- **Run:** `scripts/provision.sh run-workflow --sdk <sdk> --dir <repo_path> --demo-failure transient`. Starts the Worker with `DEMO_FAILURE=transient` (the deposit activity fails its first attempts, then succeeds), runs **the same starter command — the sample's source is never edited** (it reads its Workflow ID from the environment), and stops the Worker when done.
- **Distinct Workflow ID:** this run is named `money-transfer-demo-recovery` (vs the clean run's `money-transfer-demo`), so it appears as a **separate Workflow** in Cloud whose history shows the failure-and-recovery.
- **Wait for `workflow_status=COMPLETED`** — the retry recovered it. **Keep this call's `workflow_id` + `run_id`**; they identify the recovery run the Ending link points at.
**Step — Show the recovery.** Keep it to **one line**: the deposit failed on purpose, Temporal retried it automatically, and the Workflow still reached `COMPLETED` (the withdrawal never re-ran) — then send them to the dashboard to see it. Don't walk through the worker logs or event history line-by-line; the dashboard CTA carries the detail.
*(Variant — advanced, manual.)* `DEMO_FAILURE=permanent` makes the deposit fail **non-retryably**, so the **`refund`** compensation runs — the saga/rollback story. Two caveats:
- **Run it by hand, not through `run-workflow`** — `run-workflow` expects `COMPLETED` and would report a non-zero starter as `workflow-failed`.
- **Terminal state differs by SDK** — narrate what the history actually shows, don't assert one outcome:
- Python / Go / TypeScript / .NET → **`FAILED`** (the original error propagates after the refund).
- Java / Ruby → **`COMPLETED`** (their saga returns after compensating).
There is **no checkpoint after Phase 4** — go straight to the Ending, whose Cloud-UI link is the recovered run's timeline page (`money-transfer-demo-recovery`).
**If the user asks (plain language):** we made the deposit fail on purpose; Temporal retried it automatically and the Workflow still finished correctly — no lost state, no manual recovery. (Docs: https://docs.temporal.io/encyclopedia/retry-policies)
</phases>
## Ending the Skill
Once Phase 4 has shown the recovery, the setup is done. **Close with a short summary and the recovery link — keep it tight, no big recap table.** Do exactly this:
1. **No Worker should still be running** — `run-workflow` stops its Worker when it returns, so Phases 3 and 4 leave nothing polling Cloud. Only stop a process by hand if you ran the advanced manual `DEMO_FAILURE=permanent` variant.
2. **Show the recovered Workflow, then close with the summary.** First the Cloud UI link (bold lead-in, bare URL alone on its own line, no backticks) — the **recovery run's timeline** page, using the `workflow_id`/`run_id` Phase 4 emitted:
**View your recovered Workflow on Temporal Cloud:**
https://cloud.temporal.io/namespaces/<namespace-handle>/workflows/<workflow-id>/<run-id>/timeline
*Fallback — only if `run_id` is `unknown`: show the Workflow **detail** page the same way — bold lead-in, bare URL on its own line:*
**View your recovered Workflow on Temporal Cloud:**
https://cloud.temporal.io/namespaces/<namespace-handle>/workflows/<workflow-id>
*— never the bare `…/workflows` list.*
Then close with a short summary (a few plain sentences, no table) that makes the durable-execution payoff concrete — for example:
> 🎉 You're set up on Temporal Cloud — and you just watched **Durable Execution** in action. Your money-transfer Workflow ran on real Cloud infrastructure, and when the deposit failed on purpose, Temporal automatically retried it until it succeeded — the transfer still completed, with no lost state and not a line of retry code from you. That's the whole idea: you write the business logic; Temporal makes it survive failures and run to completion.
3. **One-line note:** the API key is saved in `temporal.toml` (give the path, `chmod 600`) and **auto-expires in ~25 hours** — don't share raw terminal logs and never commit the TOML.
The summary above is the payoff — keep it to that one short recap. The only calls-to-action are the **two Workflow timeline links** — the completed run (Phase 3) and the recovered run (here at the end). Don't loop back, re-run, keep teaching, offer teardown, or suggest other next steps — a successful completion is the terminal state.
## Failure Handling
On `status=error`, map the `error_code` via **`references/failure-handling.md`** and fix the exact cause it names — never improvise an alternate command, switch output formats, or poll. Read that file only when an error fires (progressive disclosure); the Steps spine's On-error column indexes which codes each step can emit.
## Files
- `scripts/provision.sh` — **the deterministic executor.** Owns preflight / **detect-tools** / **preview** / install / login / regions / namespace-create / **install-deps (manager-parameterized)** / key-mint+config-write / verify / await-auth / **run-workflow (Worker + starter)** / clone / repair-config / cleanup-info. Invoke it and parse its `=== RESULT ===` block (see "Execution model" above); it is the single source of truth for the pinned CLI flags, the per-SDK run commands, **and the per-(SDK,manager) install matrix + minimum-version table**. Pure bash, portable across Claude Code, Codex, and Cursor. **Read-only during a run — invoke it, never edit it (read-only script).**
- `references/unified-cli.md` — background on the prerelease CLI and the client-config TOML: `login`/`whoami`, `region list`, `namespace create`, `apikey create-for-me`, file locations, and the auth-override gotcha. The script encodes these; read the reference when a flag drifts and you need to update the script.
- `references/sdk-cloud.md` — per-SDK table: repo + cloud branch, task-queue name, how each connects (`cloud-setup` profile), and worker/starter run commands. No connection edits — the branch is pre-wired.
- `references/failure-handling.md` — the `error_code` → remediation map. Read it **only when a subcommand returns `status=error`** (progressive disclosure — the happy path never opens it); the Failure Handling section above is a one-line pointer to it, and the Steps spine's On-error column is the index.
Referenced files: 5
temporal-developer7.44 KB
---
name: temporal-developer
description: Develop, debug, and manage Temporal applications across Python, TypeScript, Go, Java, .NET, Ruby, and Rust. Use when the user is building workflows, activities, or workers with a Temporal SDK, debugging issues like non-determinism errors, stuck workflows, or activity retries, using Temporal CLI, Temporal Server, or Temporal Cloud, or working with durable execution concepts like signals, queries, heartbeats, versioning, continue-as-new, child workflows, or saga patterns. Also use when the user mentions "run a Temporal workflow from the CLI", "start a dev server", "run temporal server start-dev", "temporal workflow start", "temporal workflow execute", "temporal workflow signal", "temporal workflow query", "temporal workflow update".
version: 0.6.2
---
# Skill: temporal-developer
## Overview
Temporal is a durable execution platform that makes workflows survive failures automatically. This skill provides guidance for building Temporal applications in Python, TypeScript, Go, Java, .NET, Ruby, and Rust.
## Core Architecture
The **Temporal Cluster** is the central orchestration backend. It maintains three key subsystems: the **Event History** (a durable log of all workflow state), **Task Queues** (which route work to the right workers), and a **Visibility** store (for searching and listing workflows). There are three ways to run a Cluster:
- **Temporal CLI dev server** — a local, single-process server started with `temporal server start-dev`. Suitable for development and testing only, not production.
- **Self-hosted** — you deploy and manage the Temporal server and its dependencies (e.g., database) in your own infrastructure for production use.
- **Temporal Cloud** — a fully managed production service operated by Temporal. No cluster infrastructure to manage.
**Workers** are long-running processes that you run and manage. They poll Task Queues for work and execute your code. You might run a single Worker process on one machine during development, or run many Worker processes across a large fleet of machines in production. Each Worker hosts two types of code:
- **Workflow Definitions** — durable, deterministic functions that orchestrate work. These must not have side effects.
- **Activity Implementations** — non-deterministic operations (API calls, file I/O, etc.) that can fail and be retried.
Workers communicate with the Cluster via a poll/complete loop: they poll a Task Queue for tasks, execute the corresponding Workflow or Activity code, and report results back.
## History Replay: Why Determinism Matters
Temporal achieves durability through **history replay**:
1. **Initial Execution** - Worker runs workflow, generates Commands, stored as Events in history
2. **Recovery** - On restart/failure, Worker re-executes workflow from beginning
3. **Matching** - SDK compares generated Commands against stored Events
4. **Restoration** - Uses stored Activity results instead of re-executing
**If Commands don't match Events = Non-determinism Error = Workflow blocked**
| Workflow Code | Command | Event |
|--------------|---------|-------|
| Execute activity | `ScheduleActivityTask` | `ActivityTaskScheduled` |
| Sleep/timer | `StartTimer` | `TimerStarted` |
| Child workflow | `StartChildWorkflowExecution` | `ChildWorkflowExecutionStarted` |
See `references/core/determinism.md` for detailed explanation.
## Getting Started
### Ensure Temporal CLI is installed
Check if `temporal` CLI is installed. If not, follow the instructions at `references/core/install_cli.md` to install it for your platform.
### Read All Relevant References
1. First, read the getting started guide for the language you are working in:
- Python -> read `references/python/python.md`
- TypeScript -> read `references/typescript/typescript.md`
- Go -> read `references/go/go.md`
- Java -> read `references/java/java.md`
- .NET (C#) -> read `references/dotnet/dotnet.md`
- Ruby -> read `references/ruby/ruby.md`
- Rust -> read `references/rust/rust.md` (in Public Preview)
2. Second, read appropriate `core` and language-specific references for the task at hand.
## Primary References
- **`references/core/determinism.md`** - Why determinism matters, replay mechanics, basic concepts of activities
- Language-specific info at `references/{your_language}/determinism.md`
- **`references/core/patterns.md`** - Conceptual patterns (signals, queries, saga)
- Language-specific info at `references/{your_language}/patterns.md`
- **`references/core/gotchas.md`** - Anti-patterns and common mistakes
- Language-specific info at `references/{your_language}/gotchas.md`
- **`references/core/versioning.md`** - Versioning strategies and concepts - how to safely change workflow code while workflows are running
- Language-specific info at `references/{your_language}/versioning.md`
- **`references/core/standalone-activities.md`** - Standalone Activities: run an Activity directly from a Client without a Workflow (Public Preview)
- Language-specific info at `references/{your_language}/standalone-activities.md`
- **`references/core/troubleshooting.md`** - Decision trees, recovery procedures
- **`references/core/error-reference.md`** - Common error types, workflow status reference
- **`references/core/interactive-workflows.md`** - Testing signals, updates, queries
- **`references/core/dev-management.md`** - Dev cycle & management of server and workers
- **`references/core/cli-workflow-commands.md`** - Developer-facing CLI commands for workflow interaction (start, execute, signal, query, update)
- **`references/core/ai-patterns.md`** - AI/LLM pattern concepts
- Language-specific info at `references/{your_language}/ai-patterns.md`, if available. Currently Python only.
## Task Queue Priority and Fairness
If the developer is building a **multi-tenant application**, proactively recommend Task Queue Fairness. Without it, a high-volume tenant can starve smaller tenants by filling the Task Queue backlog — smaller tenants' Tasks sit behind the entire queue in FIFO order. Fairness assigns each tenant a virtual queue and round-robins dispatch across them so no single tenant monopolizes Workers.
Priority and Fairness also apply to tiered workloads (batch vs. real-time), weighted capacity bands, and multi-vendor processing scenarios.
- **`references/core/priority-fairness.md`** - Priority keys, fairness keys and weights, rate limiting, SDK examples, and limitations
## Additional Topics
- **`references/{your_language}/observability.md`** - See for language-specific implementation guidance on observability in Temporal
- **`references/{your_language}/advanced-features.md`** - See for language-specific guidance on advanced Temporal features and language-specific features
## Third-Party Integrations
For Temporal plugins and integrations with third-party frameworks and SDKs (Spring Boot, Spring AI, OpenAI Agents SDK, Google ADK, etc.), see **`references/integrations.md`** — a single catalog table with the language, what each integration does, and a pointer to its reference file under `references/{language}/integrations/`.
## Feedback
### Reporting Issues in This Skill
If you (the AI) find this skill's explanations are unclear, misleading, or missing important information—or if Temporal concepts are proving unexpectedly difficult to work with—draft a GitHub issue body describing the problem encountered and what would have helped, then ask the user to file it at https://github.com/temporalio/skill-temporal-developer/issues/new. Do not file the issue autonomously.
Referenced files: 105
temporal-ops33.3 KB
--- name: temporal-ops description: 'Administer and diagnose running Temporal Cloud or self-hosted Temporal Server environments via CLI (temporal, tcld) — not SDK code. Operations: namespace CRUD, Cloud capacity/APS, API-key rotation, mTLS cert rotation, workflow health, batch cancel/terminate/reset, export, search attributes, Ops API, billing, audit logs, Terraform, SAML/SCIM, migration. Diagnosis: bottom-up triage of stuck workflows, non-determinism, worker-health, task-queue problems, HA failover, payload size limits, performance bottlenecks, missed schedules. Do NOT trigger for generic TLS/gRPC errors unrelated to Temporal, writing application code (temporal-developer), or worker tuning/sizing (temporal-workertuning).' version: 0.2.2 disable-model-invocation: true --- # Skill: temporal-ops ## Overview This skill operates and diagnoses Temporal environments. It has two modes: - **Operations:** the user wants to do something — create a namespace, rotate a key, check capacity, find unhealthy workflows, cancel a batch, set up export. The skill executes the right commands and interprets the output. - **Diagnosis:** the user arrives with a symptom — a stuck workflow, a cert error, a connection timeout, a non-determinism panic. The skill routes the investigation through a layered, bottom-up diagnosis until a root cause is identified with a confidence score. It does not teach how to write workflows or activities (use `skill-temporal-developer` for that), and it does not reproduce exhaustive CLI flag tables — run `temporal <cmd> --help` for those, and see [cli-conventions.md](references/ops/cli-conventions.md) for cross-command CLI conventions. The boundary is: if the user needs to administer or troubleshoot a running Temporal environment, this skill applies. ## Out of scope - **Writing workflows, activities, or SDK code** → `skill-temporal-developer`. - **Exhaustive CLI flags / command reference** → run `temporal <cmd> --help`; **cross-command CLI conventions** → [cli-conventions.md](references/ops/cli-conventions.md). - **Worker performance tuning, sizing, capacity planning** → `skill-temporal-workertuning`. - **Helm, Kubernetes, database admin, monitoring stack config** for self-hosted — beyond the CLI surface. If the conversation drifts into one of these areas, hand off to the relevant sibling skill rather than improvising. ## Philosophy ### Operator discipline When the user wants to perform an operational task: 1. **Identify the intent and backend.** Is this a Cloud operation (`tcld`) or a self-hosted operation (`temporal operator`)? Data-plane operations (`temporal workflow`, `temporal batch`, etc.) work on both. **If the backend is ambiguous, ask before proceeding — do not assume Cloud or self-hosted and do not output environment-specific commands until you know.** 2. **Execute commands and interpret output.** Run the documented command, read the result, and report what it means — or act on it if the user asked for an action. Read-only commands (`get`, `list`, `describe`, `count`, `show`) run freely. Anything listed under [Destructive operations](#destructive-operations) is proposed to the user first. 3. **Verify the result.** After a mutating operation, confirm the new state matches the user's intent. ### Destructive operations An operation belongs to this tier if it is **irreversible** (`tcld namespace delete`), **revokes access for a live identity** (`tcld apikey delete`), **moves production traffic** (`tcld namespace failover`), or **fans out to every match** (any `--query` form). Apply the test to the operation in front of you — this is a rule, not a list, and a command's absence from any list in this skill does not place it outside the tier. For anything in the tier: gather the evidence and **propose**. Do not run it on your own initiative, and do not run one to find out what it would do. The reference file for each command states its specific blast radius; read that before proposing, not after. 1. **Blast radius as a number, not a description.** For any `--query` form, run `temporal workflow count --query '<query>'` with the byte-identical query first and carry the result into the proposal. A filter with no narrowing predicate beyond `ExecutionStatus="Running"` matches every open Execution in the Namespace. 2. **Name the target.** State the exact command, the target, and the Namespace it resolves to. Connection settings can come from `TEMPORAL_*` env vars or a config-file profile, so the target is frequently not visible in the command text. If the backend or Namespace was inferred from context rather than stated by the user, say so — a destructive command aimed at the wrong Namespace is the most common way this goes wrong. 3. **Ask explicitly, then run it so it completes.** Put the command, the target, and — for any `--query` form — the count from step 1 to the user as a direct question, and wait for an answer. Once they approve, run it with `--yes` on the `--query` batch forms; that flag is what lets an approved batch finish, since the interactive prompt needs a terminal and without one the command reports `user denied confirmation` and does nothing. `--yes` belongs in a command the user approved, never in a retry of one that failed its prompt. Do not substitute a loop over single-target `workflow terminate --workflow-id`. Approval covers one command against one target; it does not carry to the next command, a widened query, or a second Namespace. 4. **Verify, and know the abort path.** Re-run the corresponding `get`, `describe`, or `count`. A batch job drains asynchronously: `temporal batch describe --job-id <id>` shows how far it has gotten and `temporal batch terminate --job-id <id>` stops it before it reaches the rest of its matches. When a reversible sibling reaches the same goal, propose it alongside: `cancel` lets Workflow cleanup code run where `terminate` does not; `apikey disable` is reversible where `delete` is not; `accepted-client-ca add` appends where `set` replaces. Assume nothing in the environment will stop a destructive command on your behalf. Credential scope, command denylists, and confirmation prompts may or may not be configured, and their possible presence is not a reason to skip any step above — you are the safeguard the user is relying on. ### Diagnostic discipline When the user arrives with a failure or anomaly: 1. **Bottom-up diagnosis.** Verify the lower layer before blaming the upper one. The layers, from bottom to top: 1. DNS / network path 2. TCP / port reachability 3. TLS handshake 4. Authentication (API key or mTLS client cert) 5. gRPC health and Temporal frontend reachability 6. Temporal namespace, task queues, workers 7. Workflow code (determinism, signals, timers, child workflows) The full ladder lives in [diagnostic-ladder.md](references/triage/diagnostic-ladder.md). 2. **Always verify the next layer up** rather than prescribing a speculative fix. If TLS works, prove auth works before blaming the workflow. If pollers are present, prove the workflow's last event before blaming the worker. 3. **Attach a confidence score** (1-10) to every proposed diagnosis: - 9-10: symptoms, operation, and confirming signals line up cleanly. - 6-8: evidence is good but at least one alternative remains plausible. - 1-5: the issue is still ambiguous; the "fix" is the next discriminating check, not a root cause. 4. **Name ambiguity explicitly.** Errors like `context deadline exceeded` are not self-describing, and a single code such as `RESOURCE_EXHAUSTED` can mean more than one condition (account-limit throttling vs. per-Workflow lock contention). Surface that, gather more context, and scope the next step narrowly. These are skill conventions, not docs-derived facts. ## Intent routing ### Operations Find the row that matches the user's intent. The reference file contains the commands and procedures. | Intent | Category | Reference | |---|---|---| | Create, get, list, delete a Cloud namespace | Cloud namespace admin | [cloud-namespace-admin.md](references/ops/cloud-namespace-admin.md) | | Add/remove region, failover, HA config | Cloud namespace admin | [cloud-namespace-admin.md](references/ops/cloud-namespace-admin.md) | | Set retention, tags, codec-server, connectivity rules | Cloud namespace admin | [cloud-namespace-admin.md](references/ops/cloud-namespace-admin.md) | | Add or rename Cloud search attributes | Cloud namespace admin | [cloud-namespace-admin.md](references/ops/cloud-namespace-admin.md) | | Check current APS / capacity mode | Cloud capacity | [cloud-capacity.md](references/ops/cloud-capacity.md) | | Switch On-Demand ↔ Provisioned, set TRUs | Cloud capacity | [cloud-capacity.md](references/ops/cloud-capacity.md) | | Understand APS / RPS / OPS limits and throttling | Cloud capacity | [cloud-capacity.md](references/ops/cloud-capacity.md) | | Create, disable, enable, delete an API key | Cloud IAM | [cloud-iam.md](references/ops/cloud-iam.md) | | Invite, list, remove users; set roles/permissions | Cloud IAM | [cloud-iam.md](references/ops/cloud-iam.md) | | Manage user groups and service accounts | Cloud IAM | [cloud-iam.md](references/ops/cloud-iam.md) | | Generate mTLS certs, upload CA, set cert filters | Cloud certs | [cloud-certs.md](references/ops/cloud-certs.md) | | Rotate mTLS certificates | Cloud certs | [cloud-certs.md](references/ops/cloud-certs.md) | | Set up Workflow History Export (S3 / GCS) | Cloud namespace admin | [cloud-namespace-admin.md](references/ops/cloud-namespace-admin.md) | | Set up PrivateLink / PSC, manage connectivity rules | Cloud connectivity | [cloud-connectivity.md](references/ops/cloud-connectivity.md) | | Self-hosted cluster health, describe, namespace CRUD | Self-hosted admin | [self-hosted-admin.md](references/ops/self-hosted-admin.md) | | Self-hosted search attributes, Nexus endpoints | Self-hosted admin | [self-hosted-admin.md](references/ops/self-hosted-admin.md) | | Check or manage a Cloud Nexus Endpoint's caller-Namespace allowlist; the 1,000-caller default | Cloud namespace admin | [cloud-namespace-admin.md#tcld-nexus-endpoint-allowed-namespace](references/ops/cloud-namespace-admin.md#tcld-nexus-endpoint-allowed-namespace) | | Find stuck/hung/unhealthy workflows via list queries | Workflow health | [workflow-health.md](references/ops/workflow-health.md) | | Task queue poller status, workflow counts | Workflow health | [workflow-health.md](references/ops/workflow-health.md) | | Cancel, terminate, or reset workflows | Workflow recovery | [workflow-stuck.md#recovery-commands](references/triage/workflow-stuck.md#recovery-commands) | | Bulk / batch operations on workflows (`--query`) | CLI conventions | [cli-conventions.md](references/ops/cli-conventions.md#the---query--batch-job-bridge) | | Schedule CRUD, time-spec, and operations | CLI conventions | [cli-conventions.md](references/ops/cli-conventions.md#schedule-time-spec-forms) | | Complete or fail an activity externally | CLI conventions | [cli-conventions.md](references/ops/cli-conventions.md#operation--command-index) | | Cloud Ops API access, rate limits, Go SDK | Cloud Ops API | [cloud-ops-api.md](references/ops/cloud-ops-api.md) | | View billing, generate billing report, cost attribution | Cloud billing | [cloud-billing.md](references/ops/cloud-billing.md) | | Audit Logs: view, query via API, configure sink (AWS/GCP) | Cloud audit logs | [cloud-audit-logs.md](references/ops/cloud-audit-logs.md) | | Terraform provider: Namespace/User/SA/API Key/Nexus CRUD | Cloud Terraform | [cloud-terraform.md](references/ops/cloud-terraform.md) | | Expiry alerts (cert, API key, credit), status page | Cloud notifications | [cloud-notifications.md](references/ops/cloud-notifications.md) | | SAML SSO, SCIM provisioning, IdP integration | Cloud SAML/SCIM | [cloud-saml-scim.md](references/ops/cloud-saml-scim.md) | | Migrate self-hosted to Cloud (automated or manual), migrate between Cloud regions | Cloud migration | [cloud-migration.md](references/ops/cloud-migration.md) | | End-to-end ops playbook (setup, rotation, audit, billing, Terraform) | Ops recipes | [ops/recipes.md](references/ops/recipes.md) | ### Diagnosis Find the row that matches the user's symptom. Start the investigation at the first check, then read the linked reference. | Symptom | Category | First check | Reference | |---|---|---|---| | `connection refused`, cannot reach frontend | Connectivity | `nc -zvw10 <host> 7233` | [connectivity.md#connection-refused](references/triage/connectivity.md#connection-refused) | | `no such host`, DNS resolution fails | Connectivity | `dig +short <host>` or `nslookup <host>` | [connectivity.md#dns](references/triage/connectivity.md#dns) | | `tls: handshake failure`, server rejects handshake | Certificates | `openssl s_client -connect <host>:7233 -servername <host> </dev/null` | [certificates.md#handshake-failure](references/triage/certificates.md#handshake-failure) | | `x509: certificate has expired` or `not yet valid` | Certificates | `openssl x509 -enddate -noout -in cert.pem` | [certificates.md#expired-or-not-yet-valid](references/triage/certificates.md#expired-or-not-yet-valid) | | `x509: certificate signed by unknown authority` | Certificates | `openssl verify -CAfile ca.pem client.pem` | [certificates.md#unknown-authority](references/triage/certificates.md#unknown-authority) | | `tcld` session / auth fails, Cloud role unclear | Authentication | `tcld account get` | [authentication.md#cloud-role-and-permission-model](references/triage/authentication.md#cloud-role-and-permission-model) | | `UNAUTHENTICATED`, API key rejected | Authentication | `env \| grep -i TEMPORAL_API_KEY`, then `tcld apikey get --id <apikey_id>` | [authentication.md#things-to-check-when-unauthenticated-is-returned-with-an-api-key](references/triage/authentication.md#things-to-check-when-unauthenticated-is-returned-with-an-api-key) | | `namespace not found` / wrong namespace string with an API key | Authentication | Confirm Regional Endpoint form `<region>.<cloud_provider>.api.temporal.io:7233` | [authentication.md#address-form-for-api-key-connections](references/triage/authentication.md#address-form-for-api-key-connections) | | `RESOURCE_EXHAUSTED` gRPC status | Rate limits | Identify which limit fired: throttle metrics on Cloud v1, the `resource_exhausted_cause` label on v0 / self-hosted | [rate-limits.md#identifying-which-limit-was-hit](references/triage/rate-limits.md#identifying-which-limit-was-hit) | | Task queue shows no pollers | Worker health | `temporal task-queue describe --task-queue <q>` | [worker-health.md#what-no-pollers-looks-like](references/triage/worker-health.md#what-no-pollers-looks-like) | | Workflow stuck on a pending activity / timer / child / signal | Workflow stuck | `temporal workflow describe --workflow-id <id>` | [workflow-stuck.md#the-primary-inspection-command-temporal-workflow-describe](references/triage/workflow-stuck.md#the-primary-inspection-command-temporal-workflow-describe) | | `NondeterminismError`, repeating `WorkflowTaskFailed` | Non-determinism | Identify the last `WorkflowTaskFailed` cause in the Event History | [non-determinism.md#the-wft-failure-signature-of-non-determinism](references/triage/non-determinism.md#the-wft-failure-signature-of-non-determinism) | | Replay fails locally but prod workflow was running | Non-determinism | Fetch the history and run the SDK replayer in one test | [replay.md#step-2--run-the-sdk-replayer-all-supported-sdks](references/triage/replay.md#step-2--run-the-sdk-replayer-all-supported-sdks) | | HA failover did not route traffic to failover region | HA failover | `tcld namespace get --namespace <ns>.<acct>` vs. DNS CNAME | [ha-failover.md#start-here-establish-ground-truth](references/triage/ha-failover.md#start-here-establish-ground-truth) | | Serverless Worker (AWS Lambda) stopped processing work after a Namespace failover | HA failover | Confirm a `FailoverNamespace` audit event, then compare the new active region against the Lambda ARN on the Worker Deployment Version | [ha-failover.md#symptom-serverless-workers-kept-running-in-the-old-region-after-failover](references/triage/ha-failover.md#symptom-serverless-workers-kept-running-in-the-old-region-after-failover) | | `context deadline exceeded` (unknown layer) | Runtime errors | Identify which operation and SDK emitted it | [runtime-errors.md#deadline-exceeded](references/triage/runtime-errors.md#deadline-exceeded) | | `Workflow is busy` / `ResourceExhausted` on signal/update/query to one Workflow (BusyWorkflow) | Runtime errors | Rule out account-limit throttling, then split `temporal_cloud_v1_resource_exhausted_error_count` by `operation` | [runtime-errors.md#workflow-lock-contention-busyworkflow](references/triage/runtime-errors.md#workflow-lock-contention-busyworkflow) | | `PAYLOADS_TOO_LARGE`, `exceeds size limit`, payload/gRPC blob size error | Blob size limits | Check whether the issue is payload (2 MB) or gRPC message (4 MB) | [blob-size-limits.md](references/triage/blob-size-limits.md) | | Workflow stuck in invisible retry loop (gRPC message too large) | Blob size limits | Check Worker logs for `ResourceExhausted`, reduce batch size | [blob-size-limits.md](references/triage/blob-size-limits.md) | | High schedule-to-start latency, task slot depletion, slow Workflow Tasks | Performance bottlenecks | Check `temporal_workflow_task_schedule_to_start_latency` P95 | [performance-bottlenecks.md](references/triage/performance-bottlenecks.md) | | High replay latency, cache evictions, deadlock detected | Performance bottlenecks | Check `workflow_task_replay_latency` and sticky cache metrics | [performance-bottlenecks.md](references/triage/performance-bottlenecks.md) | | Schedule did not fire, missed catchup window | Missed Schedule Actions | Alert on `temporal_cloud_v1_schedule_missed_catchup_window_count` | [schedule-missed.md](references/triage/schedule-missed.md) | If a symptom does not map to a row, start at [diagnostic-ladder.md](references/triage/diagnostic-ladder.md) and work up from whichever layer was last known healthy. ## The process ### Operations path #### Step 1: Identify intent and backend Determine what the user wants to do and whether it targets: - **Temporal Cloud** → use `tcld` commands (requires `tcld login`) - **Self-hosted cluster** → use `temporal operator` commands - **Data plane (either backend)** → use `temporal workflow`, `temporal batch`, `temporal schedule`, etc. If the backend is unambiguous from context — `.tmprl.cloud` address, `tcld` command, Cloud namespace format `ns.account` → Cloud; Kubernetes/Helm, `docker-compose`, self-hosted config files → self-hosted — proceed without asking. **Otherwise, stop and ask: "Are you on Temporal Cloud or self-hosted?" before outputting any environment-specific commands.** Do not default to either environment. Once known, save the answer to memory so you don't ask again in future conversations. #### Step 2: Execute and interpret Look up the intent in the Operations table above. Read the linked reference file for the exact commands, flags, and expected output. Run the command and interpret the result for the user. #### Step 3: Verify After a mutating operation (create, update, delete, rotate), confirm the new state: - Re-run the corresponding `get` or `describe` command - Confirm the output matches the user's intent - Report the result ### Diagnosis path #### Step 1: Identify the symptom Ask the user for the exact, copy-pasted error text. Do not accept paraphrases — the exact string often encodes the layer (e.g., `x509:` prefix means TLS/cert layer, `RESOURCE_EXHAUSTED:` prefix means gRPC rate limit, `NondeterminismError` means workflow replay layer). Note that `RESOURCE_EXHAUSTED` alone does not tell you which condition fired — account-limit throttling and per-Workflow lock contention (`Workflow is busy`, i.e. BusyWorkflow) share the code. Split the resource-exhausted metric by its label (`operation` on the Cloud v1 family, `resource_exhausted_cause` on v0 and self-hosted) rather than parsing the free-text message. Confirm three things before continuing: - What command was run, or what SDK call produced the error? - What environment produced it (local dev server, self-hosted cluster, Temporal Cloud)? If clear from context (addresses, commands, namespace format), don't ask — **but if uncertain, ask now before proceeding with any diagnosis.** Save the answer to memory for future conversations. - What changed recently (new deploy, new certs, new namespace, new region)? #### Step 2: Gather context The context the investigation needs depends on the category. At minimum: - **For any Cloud auth / connectivity issue:** auth method (API key vs mTLS), exact address, exact namespace, SDK + version. The endpoint family differs by auth method — see [connectivity.md#endpoint-formats](references/triage/connectivity.md#endpoint-formats). For private connectivity (PrivateLink / PSC), TLS server name overrides also vary by auth method — see [cloud-connectivity.md](references/ops/cloud-connectivity.md). - **For a stuck workflow:** namespace, workflow ID, run ID, and the output of `temporal workflow describe --workflow-id <id>` (pending-operation state lives here, not in the Event History alone). Event History via `temporal workflow show` is the companion view. - **For a worker health issue:** worker logs (registration errors, auth errors, panics), the output of `temporal task-queue describe --task-queue <q>`, and `temporal worker describe --task-queue <q>` for per-worker details. - **For a non-determinism error:** the worker log line containing the error, the workflow type name, and access to the history JSON for replay. #### Step 3: Validate pasted SDK config (if any) If the user has pasted SDK connection code — even just the address/namespace/auth fields — review it against [sdk-snippet-review.md](references/triage/sdk-snippet-review.md) **before** descending the ladder. Wrong endpoint family, short namespace, or mismatched auth method will make every network-layer probe below look broken when nothing lower actually is. Skip this step when the user has an established, previously-working config and the symptom is new — the snippet is not the culprit, the environment changed. Otherwise treat snippet validation as Layer 0. #### Step 4: Descend the ladder Use [diagnostic-ladder.md](references/triage/diagnostic-ladder.md) to pick the right starting layer. As a rule of thumb: - Auth / connectivity / cert symptom → start at layer 1 (DNS) and walk up. - Worker / task-queue symptom → start at layer 6 (namespace + pollers). - Stuck workflow / determinism symptom → start at layer 7 (workflow code), but confirm layer 6 (worker is actually polling) first. Each layer has a command that proves it healthy and a failure signature that tells you whether the problem lives at that layer or higher. #### Step 5: Fix and verify Prescribe the fix scoped to the root cause. Then verify by re-running the layer's healthy-check command and, if possible, the original user operation. Attach the confidence score to the diagnosis. If the layer above the fix is still failing, return to step 4 and continue walking upward — the first broken layer is rarely the only one. ## Prerequisites - **Temporal CLI** (`temporal`) — required for data-plane operations and self-hosted admin. Install: `brew install temporal` or see [Temporal CLI docs](https://docs.temporal.io/cli). - **tcld** — required for Cloud operations. Install: `brew install temporal-cloud-cli` or see [tcld docs](https://docs.temporal.io/cloud/tcld). Authenticate with `tcld login` before use. ## Reference files ### Operations - [cloud-namespace-admin.md](references/ops/cloud-namespace-admin.md) — Cloud namespace lifecycle via `tcld`: create, get, list, delete, failover, add-region, retention, tags, codec-server, HA config, connectivity rules, search attributes, accepted-client-ca, certificate filters, export, and the `tcld nexus endpoint allowed-namespace` caller allowlist (1,000-caller Access Policy ceiling). - [cloud-capacity.md](references/ops/cloud-capacity.md) — Capacity modes (On-Demand / Provisioned), APS/RPS/OPS definitions, TRUs, `tcld namespace capacity update`, default limits, throttling, APS management best practices. - [cloud-iam.md](references/ops/cloud-iam.md) — API key lifecycle (`tcld apikey`), users (`tcld user`), user groups (`tcld user-group`), service accounts, account operations (`tcld account`), roles, namespace permissions. - [cloud-certs.md](references/ops/cloud-certs.md) — mTLS cert management: generating certs with `tcld generate-certificates`, uploading CAs, certificate filters, cert rotation, switching mTLS ↔ API keys. - [cloud-connectivity.md](references/ops/cloud-connectivity.md) — Private connectivity (AWS PrivateLink / GCP PSC), connectivity rules: setup, rule parameters, tcld commands, attaching rules to namespaces. - [cloud-migration.md](references/ops/cloud-migration.md) — Migration paths: automated self-hosted→Cloud (S2S proxy, `tcld migration` commands, 5 phases), manual self-hosted→Cloud (client changes, workflow strategies), within-Cloud region-to-region (HA add-region/failover). - [cloud-ops-api.md](references/ops/cloud-ops-api.md) — Cloud Ops API: HTTP and gRPC endpoints (`saas-api.tmprl.cloud`), Go SDK, protobuf compilation, rate limits (160 RPS account, 40 user, 80 SA, 10 concurrent async), API version header, use cases. - [cloud-billing.md](references/ops/cloud-billing.md) — Cloud billing: Billing Center (invoices, credits, plans, cost by namespace), Usage Dashboards, Billing API: async CSV report generation, FOCUS-friendly format, 27-column report schema, date range constraints. - [cloud-audit-logs.md](references/ops/cloud-audit-logs.md) — Cloud Audit Logs: supported control plane events (Account, API Keys, Connectivity Rules, Namespace, Export, Nexus, Service Accounts, User, User Groups), JSON format, API access (30-day retention), AWS Kinesis and GCP Pub/Sub sink configuration. - [cloud-terraform.md](references/ops/cloud-terraform.md) — Terraform provider: setup (`TEMPORAL_CLOUD_API_KEY`), `temporalcloud_namespace`/`temporalcloud_user`/`temporalcloud_service_account`/`temporalcloud_apikey`/`temporalcloud_nexus_endpoint` CRUD, import, data sources (regions, namespaces), limitations (API keys not importable, cannot manage Account Owner). - [cloud-notifications.md](references/ops/cloud-notifications.md) — Cloud notifications: certificate expiry (15/10/5 days), API key expiry (30/20/10 days), credit consumption/expiry alerts, plan changes, failover events, recipient roles, `noreply@temporal.io` sender. - [cloud-saml-scim.md](references/ops/cloud-saml-scim.md) — SAML SSO (Entra ID, Okta): entity identifier (`urn:auth0:prod-tmprl:ACCOUNT_ID-saml`), callback URL (`login.tmprl.cloud`), IdP configuration steps, support ticket workflow. SCIM: supported vendors, prerequisites (SAML first), 10-minute sync window, group-to-role mapping. - [self-hosted-admin.md](references/ops/self-hosted-admin.md) — Self-hosted control plane via `temporal operator`: cluster health/describe, namespace CRUD, search-attribute create/list/remove, Nexus endpoint CRUD. - [workflow-health.md](references/ops/workflow-health.md) — Data-plane health queries: `temporal workflow list` with List Filters, `temporal workflow describe`/`show`/`count`, `temporal task-queue describe` for poller status. - [cli-conventions.md](references/ops/cli-conventions.md) — Cross-command `temporal` CLI conventions: connection/identity (`TEMPORAL_*` env vars ↔ `--address`/`--namespace`/`--api-key`, `--identity`), output/formatting (`--output`, `--time-format`, payload shorthand), the `--query` ⇒ batch-job bridge (with `temporal batch describe/list/terminate`), and schedule time-spec forms. Ends with an operation→command index that routes each data-plane operation to its owner file. Delegates exhaustive flags to `temporal <cmd> --help`. - [ops/recipes.md](references/ops/recipes.md) — End-to-end ops playbooks: set up new namespace, check APS, switch capacity mode, find hung workflows, rotate API key, audit access, rotate mTLS certs, check self-hosted health, view billing / generate billing report, configure audit log sink, provision resources with Terraform, set up SAML SSO. ### Diagnosis - [sdk-snippet-review.md](references/triage/sdk-snippet-review.md) — Layer-0 config check for pasted SDK connection snippets: endpoint form per auth method, namespace format, auth / TLS expectations, `TEMPORAL_*` env vars, common misconfigurations. Run before the diagnostic ladder. - [diagnostic-ladder.md](references/triage/diagnostic-ladder.md) — the seven-layer bottom-up model, with one canonical command per layer and cross-links into the topical leaves. - [connectivity.md](references/triage/connectivity.md) — DNS, TCP, endpoint families (Namespace Endpoint for mTLS vs. Regional Endpoint for API keys), firewall/proxy shapes, PrivateLink/PSC, quick diagnostic scripts. - [certificates.md](references/triage/certificates.md) — x509 and TLS alert strings, expiry / unknown-authority / hostname-mismatch / key-mismatch diagnosis, Cloud accepted-client-CA set via `tcld namespace accepted-client-ca`, Cloud mTLS certificate requirements, rotation and expiry notifications, openssl recipes. - [authentication.md](references/triage/authentication.md) — `UNAUTHENTICATED` vs `PERMISSION_DENIED`, API-key lifecycle (`tcld apikey` commands, env var propagation, required Regional Endpoint form), mTLS after TLS (certificate filters, identity-to-role mapping), Cloud account-level roles and namespace-level permissions. - [workflow-stuck.md](references/triage/workflow-stuck.md) — Workflow Execution Status values, `temporal workflow describe` as the primary inspection command, Event History via `temporal workflow show`, pending activities / child workflows / signals / Nexus operations / Workflow Tasks, WorkflowTaskFailed retry loops, recovery commands (signal, terminate, cancel, reset, pause/unpause). - [non-determinism.md](references/triage/non-determinism.md) — determinism definition, WFT-failure signature, ND-inducing code patterns, per-SDK error shapes, identifying ND from Event History, local replay reproduction, remediation via Worker Versioning / patching / reset. - [worker-health.md](references/triage/worker-health.md) — no-pollers runbook via `temporal task-queue describe`, reachability and versioning, worker-level describe, schedule-to-start latency, worker task slots, sticky execution and sticky cache, worker heartbeating, Cloud namespace-level poller limits, worker log signatures. - [rate-limits.md](references/triage/rate-limits.md) — what `RESOURCE_EXHAUSTED` means (and does not), Cloud APS / RPS / OPS under On-Demand and Provisioned capacity modes, self-hosted `frontend.rps` / `frontend.namespaceRPS` dynamic config, identifying which limit fired via the throttle metrics (Cloud v1) or the `resource_exhausted_cause` label (v0 / self-hosted), and separating account-limit throttling from single-resource exhaustion. - [ha-failover.md](references/triage/ha-failover.md) — Cloud HA routing via the Namespace Endpoint CNAME, verifying the active region (control-plane `tcld namespace get` vs. DNS view), clients that did not follow the failover, PrivateLink after failover, failover-not-executing, handover-window errors, platform limits, RPO/RTO semantics, and Serverless Workers (AWS Lambda) not following a failover because compute-provider configuration is region-scoped. - [runtime-errors.md](references/triage/runtime-errors.md) — deadline-exceeded disambiguated by operation and by where the call was made, Workflow lock contention (BusyWorkflow) separated from account-limit throttling and confirmed via the `operation` breakdown, routing for `no pollers` / `INVALID_ARGUMENT` / unspecified `UNAVAILABLE`. - [replay.md](references/triage/replay.md) — fetching Event History with the SDK client (CLI export as fallback), running the SDK replayer in every supported SDK (Go, Python, TypeScript, Java, .NET, Ruby, PHP), `TEMPORAL_DEBUG` and the deadlock detector, interpreting divergent and successful replays, and the TypeScript-only VS Code extension. - [blob-size-limits.md](references/triage/blob-size-limits.md) — Payload size limit (2 MB) and gRPC message size limit (4 MB): error messages, per-SDK behavior (Python 1.23.0+ vs. others), claim check pattern, External Storage (Pre-release), batch-size reduction. - [performance-bottlenecks.md](references/triage/performance-bottlenecks.md) — Latency and throughput diagnosis via SDK metrics: schedule-to-start latency, workflow task execution latency, replay latency, activity execution latency, task slot depletion, network request metrics, sticky cache metrics. - [schedule-missed.md](references/triage/schedule-missed.md) — Missed Schedule Actions: alerting via `temporal_cloud_v1_schedule_missed_catchup_window_count` / `schedule_missed_catchup_window`, investigation via `temporal schedule list` + `temporal schedule describe`, DescribeSchedule fields (`missedCatchupWindow`, `overlapSkipped`, `bufferDropped`), default catchup window (one year), root causes, overlap policies (6 values), backfill remediation. - [recipes.md](references/triage/recipes.md) — four end-to-end triage walkthroughs: stuck workflow at 3am, cert expired with workers offline, task-queue backlog mystery, non-determinism caught in prod. ## Reporting Issues in This Skill If you (the AI) find this skill's explanations are unclear, misleading, or missing important information, draft a GitHub issue body describing the problem encountered and what would have helped, then ask the user to file it at https://github.com/temporalio/skill-temporal-ops/issues/new. Do not file the issue autonomously.
Referenced files: 33
temporal-serverless38.6 KB
--- name: temporal-serverless description: 'Deploy and operate Temporal Workers on serverless compute (AWS Lambda) driven by the Worker Controller Instance (WCI). Use when the user mentions: "serverless worker", "Temporal serverless", "Worker Controller Instance", "WCI", "deploy Temporal worker on Lambda", "Lambda packaging", "Lambda timeout", "WCI inspection", "CloudFormation Temporal".' version: 0.6.2 disable-model-invocation: true --- # Skill: temporal-serverless ## Overview This skill helps users deploy and operate Temporal Workers on serverless compute. Instead of a long-lived process, Temporal invokes the Worker on demand through the Worker Controller Instance (WCI); the Worker processes available Tasks and shuts down, scaling to zero when idle. The skill produces Worker code, deployment configuration, connection configs, and packaging steps for the chosen SDK, and walks users through troubleshooting when serverless Workers aren't picking up Tasks. ## Supported compute providers | Cloud provider | Compute service | Support | Reference directory | |---|---|---|---| | AWS | Lambda | Supported — Public Preview, open to all Temporal Cloud customers | `references/aws-lambda/` | | GCP | Cloud Run | Not supported | — | Only a provider marked Supported is covered. If a request names another, say it is not supported and stop; do not adapt a supported provider's material to it. **Never let the provider be an unstated assumption:** when the request does not name one, it is confirmed in the step 1 questions, not silently defaulted. Every supported provider's directory carries the same shared layout — `setup.md`, `iam.md`, `versioning.md`, `diagnostics.md`, `observability.md`, `self-hosted.md` — plus one `sdk-<language>.md` file for each supported SDK. Paths below are written `references/<provider>/…`; substitute the directory from the table. Provider-specific commands, templates, permissions, SDK APIs, and defaults live there — this file stays at the workflow level. When a step needs concrete commands or SDK details, go to the reference file named at the end of that step. | SDK language | AWS Lambda reference | |---|---| | Go | `references/aws-lambda/sdk-go.md` | | Python | `references/aws-lambda/sdk-python.md` | | TypeScript | `references/aws-lambda/sdk-typescript.md` | | Java | `references/aws-lambda/sdk-java.md` | | .NET | `references/aws-lambda/sdk-dotnet.md` | **Public Preview is not GA.** The APIs are still evolving and may change: pin SDK and CLI versions for anything long-lived, and read the installed package's actual API surface rather than writing from memory. ## Deployment workflow Follow these steps in order. Each step is provider-neutral; the concrete commands, templates, and options live in the reference file named at the end of the step. **Open a new deployment with a plain-language summary of the run.** Before the step 1 questions, tell the user in a few sentences what is about to happen: that this creates real resources in their cloud account which cost money for as long as they exist; that you will ask about a handful of things, then show an exact list of what you are about to create and wait for approval, and that nothing is created before that approval; that the middle of the run is unattended; and that it ends with a Workflow they can watch execute, an inventory of everything created, and an offer to remove it all. Name the five stages below in ordinary words. Do not explain Temporal or serverless compute; keep it short enough to read at a glance. **Lay it out as bullets, with the five stages as sub-bullets under "How it goes" — one stage per line, never chained into a single run-on bullet.** Follow this shape: > Here's what's about to happen, before I ask anything: > > - This creates real resources in your cloud account — the compute unit that runs your Worker, roles, an infrastructure stack, logs. They're live and billable for as long as they exist. > - **How it goes.** Five stages: > - **Scope** — a handful of questions, below. > - **Access** — check credentials and permissions on both sides, then show you an exact list of what I'm about to create and wait for your approval. > - **Build** — write, package, deploy the Worker. > - **Connect** — bind the Task Queue, set the version current. > - **Verify and hand back.** > - Nothing gets created before you approve that list. After approval the middle stretch runs unattended. > - At the end you get a Workflow you can watch execute, a full inventory of everything created, and an offer to remove it all. **Write the summary provider-neutral, because at that point you do not know the provider.** It is one of the things step 1 asks. Say "your cloud account", never the name of a provider you have not been told. The same applies to the account, Namespace, and region: if a cheap read-only call has already told you (see step 1), name what you actually found; otherwise leave it out rather than filling it in with a plausible guess. Skip the summary for troubleshooting, inspection, and configuration-change tasks. Someone whose Worker is not being invoked does not need an overview of a deployment they have already done. **Then track the run on a checklist, and reprint it every time a step completes.** The eight steps group into the five stages below. Create one item per step, grouped under its stage, and build the checklist as soon as step 1's answers land, so items can name the confirmed provider and the agreed prefix instead of hedging. **Reprint the whole checklist at each step boundary — not just the item that changed, and not a sentence saying the stage is done.** Mark finished items ✅, the one you are starting ⏳, and the rest ⬜. Use the bare marker with nothing in front of it — `✅ Confirm SDK`, not `- [x] Confirm SDK` — and put each item on its own line. A narrated "Access complete, now Build" is not a substitute: it says where you are but not what remains, and the user cannot see it without scrolling back to a checklist printed twenty commands ago. Reprint during Scope and Access too — those stages end in a user decision, and the reprint is what shows the decision landed and what it unblocked. Where the harness has a todo list, use it *in addition to* the printed checklist, not instead of it. It is not part of the transcript the user reads back. **Word each item as plain language about what happens, not as a compressed step title,** and name both sides concretely — the confirmed compute provider and Temporal, never "both sides." Follow this shape: > **Scope** > ✅ Confirm SDK (Go), compute provider (AWS Lambda), Namespace (`<ns>`), and naming prefix (`<prefix>`) > > **Access** > ⏳ Check credentials and permissions on AWS and on Temporal, then show the exact list of resources to be created and wait for your approval > > **Build** > ⬜ Write the Worker against the installed package's real API > ⬜ Cross-compile, package, deploy the compute unit, wait for it to report ready > > **Connect** > ⬜ Create the role Temporal assumes to invoke the Worker > ⬜ Register the Worker Deployment Version, confirm the validation invocation bound the Task Queue, set it current > > **Verify and hand back** > ⬜ Start a Workflow and confirm it executes, from both the Temporal side and the provider's logs > ⬜ Deliver the inventory of everything created, then offer teardown | Stage | Steps | Complete when | |---|---|---| | Scope | 1 | SDK, compute provider, Namespace, and naming prefix are all confirmed by the user. | | Access | 2 | Compute provider and Temporal both authenticated, permissions confirmed, and the list of resources to create approved. | | Build | 3–4 | The compute unit is deployed and reports ready, built for the architecture it runs on. | | Connect | 5–6 | The Task Queue is bound and the version is current. | | Verify and hand back | 7–8 | A Workflow completed, two independent signals agree, the inventory is delivered, and teardown has been offered. | **A step is complete when its verification passed — not when its command exited zero.** Several commands in this workflow exit clean having done nothing: the traffic-shifting and key-revocation commands no-op when their confirmation prompt goes unanswered, and providers return from create and update calls while the resource is still settling. Check an item off against state you read back, not against an exit code. When a step's verification fails, say which step you are on and what it is blocked on rather than moving down the list. 1. **Scope the task.** Identify the SDK language (Go, Python, TypeScript, Java, or .NET), the deployment target (Temporal Cloud or self-hosted — self-hosted has its own server prerequisites), the compute provider, and whether this is a new setup, a configuration change, or troubleshooting. Confirm the deployment target is compatible with the chosen provider — see "A Namespace on the target cloud provider is required" under Provider-neutral principles. Ensure a Temporal client/CLI is available and authenticated to the target. Each changes the specifics. → `references/concepts.md` for what the user is building; `references/<provider>/setup.md` for the compatibility and client-setup details. **Put the compute provider in that batch of questions as a confirmable default, not a free choice.** Pre-select the supported provider from the table above and carry its support status in the option's description. The user confirms rather than chooses, so it costs no extra turn, but the provider is never something they were assumed into. Skip the question only when the request already names a provider. Do not restate any of this in a paragraph before the questions; the option description is where it belongs. **Let the user pick the Namespace from a list; never make them retype one.** Namespace names are long and error-prone — a generated suffix on an account ID, `<name>-<suffix>.<account>`. Where control-plane access is available, `tcld namespace list` returns the full Namespace objects, so one call gives every name with its region — and a region ID is provider-prefixed (`aws-…`, `gcp-…`), so the same response tells you each Namespace's provider. Only the prefix carries meaning; the region itself imposes no constraint. Present it like this: - **Offer the eligible Namespaces as the options**, each labelled with its region. - **Summarize the ineligible ones in a single line** — "you also have 2 Namespaces on \<provider\>, which this skill does not support" — rather than listing them individually or hiding them. A user who knows they have a Namespace and cannot find it in the list concludes the tool is broken; one line keeps them informed and explains the constraint. - **Name the account you are listing from and confirm it is the intended one** before showing anything. A stale credential lists a real account that is not the one the user means to deploy into, and every option under it looks authoritative. - **If more Namespaces are eligible than the question format can hold, print the labelled list and ask the user to name one.** Do not silently show only the first few. This also settles the compute-provider answer, since a Namespace can only be served by compute on its own cloud provider — so a mismatch is caught here rather than at connection time, several steps later. **Degrade gracefully if `tcld` is not authenticated.** Ask the user for the Namespace name rather than stopping to fix the login — they can copy it from the Cloud UI, where it appears on the Namespace page and in the URL. Ask for its region in the same batch of questions: the name alone does not tell you the provider, and a mismatch missed here surfaces at connection time instead. **Never source a Namespace, account, or resource identifier from shell history.** History is stale by construction — it is full of last quarter's accounts — and reading it to guess a deployment target produces confident, wrong answers. Take identifiers from the user or from an authenticated API call, and nowhere else. **Agree a resource-naming prefix in this same batch of questions, and propose a default so the user can accept without thinking about it.** Assume the account and the Namespace are shared — unprefixed names like `temporal-serverless-worker` collide with, or quietly shadow, another team's deployment. Naming is not a late cosmetic choice you can patch on the provider side: the deployment name, build ID, and Task Queue are compiled into the Worker binary, so changing them after step 3 means editing code, rebuilding, repackaging, and cleaning up whatever was already created under the old names. Once agreed, apply the prefix to everything you create on both sides — compute unit, roles, infrastructure stacks, log groups, deployment name, and Task Queue. **The prefix you propose must be identifying** — derived from the user, their team, or the project. A generic word like `demo`, `test`, or `temporal` collides about as readily as no prefix at all, so never offer one as the safe choice. Offer exactly two options plus the free-text escape: the identifying prefix, and "no prefix" — some users genuinely own the account. Do not offer a second prefix string; the consequential choice is prefix versus none, and anything else goes in free text. When you offer "no prefix," say what it risks in the same breath: unprefixed names can collide with or shadow an existing deployment, and that surfaces as another team's Worker behaving oddly rather than as an error you will see. 2. **Confirm you can make the required changes — before making any.** Determine which credentials are available (for the compute provider and for Temporal) and confirm the active identity actually has permission to make the changes the task needs — creating or updating compute resources, creating roles, registering deployment versions. Verify *both* sides: the compute provider AND Temporal access. Do not run account-mutating commands and let them fail partway. **If access is missing or unconfirmed, stop and ask the user how they want to proceed** — extend their identity's permissions, have an administrator make the change and hand back the result, or generate the commands for the user to run under a privileged identity. Changing a user's cloud account is consequential; confirm authorization and the preferred method first. → `references/<provider>/iam.md` (exact permissions, compute-provider preflight) and `references/<provider>/setup.md` (Temporal connection preflight). **Classify an authentication failure before acting on it — "not signed in" and "not permitted" have different fixes.** A failed preflight does not automatically mean the bottom row of the table below. An absent or expired credential is usually recoverable in this session, in under a minute. A caller that resolves but is denied a specific action is a real permissions problem. Never collect credentials in the conversation: no interactive credential-configuration wizards, and never ask the user to paste access keys, API keys, or session tokens. → `references/<provider>/iam.md` (credential recovery). **Then ask the user which way they want to go, and do not choose for them:** - **Fix the CLI** — you run the login flow, surface the verification URL for them to open, wait for it to complete, re-run the preflight, and continue with full automation. - **Work in the browser** — the user makes the changes in the Cloud UI and their cloud provider's console while you give the instructions step by step, naming the exact path for each one, and they report the result back. Both paths reach the same end state, so present them as equals rather than as a preference and a fallback — every control-plane step in this workflow exists in the Cloud UI (see `references/<provider>/setup.md`). Ask once, then commit to the answer — do not re-offer the login at every subsequent step, and never start an identity-provider login as a silent side effect of a preflight. Adapt to what is available — the skill is valuable at every level: | Compute-provider access | Temporal access | Behavior | |---|---|---| | Authenticated | Authenticated | Full workflow — run commands, verify results, register deployment versions. | | Authenticated | None | Deploy compute infrastructure; walk the user through the Temporal steps in the Cloud UI, or generate the commands for them to run. | | None | Authenticated | Write Worker code and configs; walk the user through the compute steps in their provider's console, or generate the deploy commands; run Temporal commands and verify WCI state. | | None | None | Write Worker code, deploy templates, permission policies, connection configs, packaging scripts; provide all commands with placeholder values. | **Do not self-select a row.** Drop to a lower one only after the choice above has been put to the user and the browser path chosen, or the login attempted and failed. When you hand off a runbook, say the offer stands — if the user authenticates and comes back, take the work over rather than leaving them to run the steps by hand. **Before the first account-mutating command, list what you are about to create — with final names — and get approval.** Name the target account and region, then every resource: compute unit, execution role, infrastructure stack, log group, deployment name, and Task Queue. Say plainly that they are live and billable. This is the mirror of the inventory in step 8, and it is worth more here than there: it makes the naming prefix concrete while changing it is still free, and the deployment name, build ID, and Task Queue become expensive to change once step 3 compiles them into the Worker. Skip it only when nothing will be created — a troubleshooting or inspection task. 3. **Author the Worker.** *Install the SDK's serverless Worker package before writing any code* — it is usually shipped separately from the main SDK — sometimes on its own version line, sometimes in lockstep with it, and in one SDK not separately at all — so having the base SDK installed does not mean it is importable. Then read the installed package's actual API surface and write against that; these are Public Preview APIs that drift between versions, and generating code from memory costs a build cycle. Entry-point names are not consistent between SDKs, so inspect first rather than pattern-matching from another language. Every Workflow must declare a versioning behavior (`Pinned` or `AutoUpgrade`), per-Workflow or as a Worker-level default — code without it fails at runtime. → `references/<provider>/sdk-<language>.md` (package, install, API inspection, entry point, handler shape, versioning behavior, tuned defaults). 4. **Package and deploy the compute unit.** Build and package per SDK, deploy the compute unit, and set the invocation deadline high enough for the Worker to start, connect, register the Task Queue, and shut down gracefully. Match the build's target architecture to the deployed compute unit's — a mismatch fails only at invocation time, not at build time. After a create or update, wait for the compute unit to reach a ready state before the next step; providers return from these calls while the unit is still settling. → `references/<provider>/sdk-<language>.md` (build, packaging, runtime, handler, architecture, and SDK-specific deployment values) and `references/<provider>/setup.md` (shared deployment lifecycle). 5. **Grant Temporal permission to invoke the Worker.** Configure the compute provider's access so Temporal can invoke and inspect the Worker. This access is separate from the compute unit's own execution role — do not confuse the two. Two things to get right before you create anything: (a) this grant is **shared, account-wide infrastructure** that a previous deployment may already have created — look for an existing one and extend it to cover your new Worker rather than creating a parallel copy, and never delete or repurpose one you did not create without asking; (b) scope the grant so that *future* immutable builds are covered, not just today's — a grant pinned to one build breaks the next release in a way that surfaces later as an unrelated-looking invocation failure. → `references/<provider>/iam.md`. 6. **Register the Worker Deployment Version, verify the validation invocation, then set it current.** Create the Worker Deployment Version with the compute provider configured; the deployment name and build ID must exactly match the values in the Worker code. Creating it triggers one validation invocation — **check that it bound the Task Queue before going further.** If the Task Queue is bound, the permission grant, package, config, and deadline are all provably correct, and any later failure is downstream; if it is not, setting the version current will not fix it. Then set it current: through the UI this happens automatically, through the CLI it is a separate step, without which Tasks never route to the version. → `references/<provider>/setup.md`. 7. **Verify.** Start a Workflow on the Task Queue and confirm Temporal invokes the Worker — check the Workflow history in the Temporal UI and the compute provider's logs. If it does not progress, → `references/<provider>/diagnostics.md`. 8. **Hand back the inventory first; offer teardown as the closing note.** The order is inventory → offer, never the reverse. Close with what now exists — compute unit and published build identifiers, roles, infrastructure stacks, region, deployment name and build ID — and what the run actually did, including anything you worked around or deviated from. Say plainly that it is live and billable. These names are only knowable from the run that created them, and reconstructing them later means scanning the user's account. **Do not write a teardown script before the user asks for one.** Generating it unprompted buries the inventory under a file they did not request, and the inventory is what they need in order to decide. End with a single line — *"Let me know if you want a teardown script to remove these resources"* — and stop there. Write the script, or run the teardown, when they take you up on it. → `references/<provider>/setup.md` (Teardown). ## Working practices How to move through the workflow above. - **Say what a command will do in your own text, above the command.** The user sees a collapsed "Ran 6 shell commands" in the transcript, not the commands themselves, so an unannounced batch is opaque at exactly the moments that matter. State it in one line before the tool call: what it does and to what — the resource, and the account or Namespace it touches. **The tool's own description field does not count.** It renders at the bottom of the command block, underneath the command it is describing, where the user has to go looking for it. The summary belongs above the block, as ordinary message text. **Set it apart on its own line so it is visibly not narrative prose** — bold, and nothing else on the line: > **Checking for the Temporal CLIs, AWS credentials, and which AWS account and Temporal account they resolve to** > > **Creating the invocation role Temporal assumes, with a generated External ID** For anything that creates, updates, or deletes, name the resource and the target account or Namespace explicitly — an approval prompt should arrive with its justification already on screen, not after it. - **Read the current state instead of recalling it.** Check the installed package's API, the CLI's own `--help` for the flags you are about to pass, the compute unit's reported state, and the CLI version. Each of these has drifted in practice: a Public Preview SDK whose fields moved, a CLI too old to have the serverless subcommand at all, a resource that reports success while still settling. - **Do not chain `cd` with commands that create or modify files.** A compound `cd <dir> && <write>` triggers a manual approval prompt no matter how the user's permissions are configured, so scaffolding a project this way asks for approval on every run. Use absolute paths, or the tool's own directory flag (`go -C <dir> …`), and rely on the shell's working directory persisting between calls — the `cd` buys nothing and costs a prompt. Keep the command count down for the same reason: one `go get` covering both packages beats two. - **Verify each step before building the next on top of it.** Compile the Worker before packaging it, confirm the package's target architecture before uploading, wait for the compute unit to be ready before publishing a build, and confirm the Task Queue is bound before shifting traffic. Deployment failures here surface far from their cause — an architecture or dependency mismatch appears only at first invocation, and a first-invocation failure appears as "the Worker is never invoked", several steps later. - **When something fails, read the actual error before changing anything.** Fetch the failure reason from the provider (deployment events, logs, status fields) and fix that. Do not retry the same command with variations, and do not start editing permissions or trust policies on the theory that the problem might be access — most first-invocation failures are not permission problems, and some failures are on Temporal's side and will reproduce no matter what you change. - **Treat the user's account as shared and pre-existing.** Assume other deployments, roles, and stacks are already there. Look before creating, extend rather than duplicate, and never delete or repurpose something you did not create without asking. When you do work around existing infrastructure — a different name, a reused role — say so explicitly in your summary rather than leaving it as a silent deviation. - **Confirm the end state from two independent signals.** A Workflow that completes in the Temporal UI *and* the Worker's own logs showing startup, Task Queue registration, and Task execution. One signal alone can mislead: a system Workflow that exists and is running proves nothing about invocation health, and a command that exits zero may have done nothing at all if it was waiting on a confirmation prompt. - **Account for what you created.** Keep the inventory as you go rather than reconstructing it at the end, say plainly that the resources are live and billable, and offer to tear them down (step 8). ## Never create or manage the WCI Temporal creates the WCI automatically once a Worker Deployment Version has a compute provider. You never create, start, or manage it. A WCI that exists or is running is *not* evidence that invocation works — it continue-as-news and keeps running even while its Activities fail. Diagnose from Temporal's own signals: read the WCI Workflow history and look for Activity failures. Do not enumerate compute resources across regions or scan the account to reverse-engineer state. → `references/concepts.md`, `references/<provider>/diagnostics.md`. ## Provider-neutral principles Surface these early — they apply regardless of compute provider: - **A Namespace on the target cloud provider is required.** A Serverless Worker runs only on the cloud provider that hosts its Temporal Cloud Namespace — there is no cross-cloud pairing. Confirm the user has a Namespace on the provider they intend to run compute on *before* building anything; without one, the work stops there and they need either a Namespace on that provider or a different provider. A mismatch is not caught at deploy time — it fails later, at connection time. **Regions do not have to match:** a Namespace in one region can drive a compute unit in another, so never tell a user to move or re-create a Namespace to line up regions. - **Use `tcld` for every Temporal Cloud control-plane operation** — accounts, Namespaces, API keys, users, service accounts. Do not use the unified CLI's `temporal cloud …` subcommands for them. Worker Deployments and Workflows are *not* control-plane operations: they live on the Namespace frontend, have no `tcld` equivalent, and use `temporal worker deployment …`. → `references/<provider>/setup.md`. - **Versioning behavior is mandatory.** Every Workflow needs `Pinned` or `AutoUpgrade`, or the Worker sets a default. - **Deployment name and build ID must match exactly** between the Worker code and the Worker Deployment Version. A mismatch causes an invocation loop (Temporal invokes → Worker polls with the wrong version → Task not processed → invoke again). Signature: rapid repeated invocations with no Workflow progress. - **Set the invocation deadline high enough.** Providers often default to a very short timeout. If the first invocation times out before the Worker registers the Task Queue, the binding is never created and the Worker is never invoked again. → `references/<provider>/setup.md` for the exact default. - **Use an immutable, versioned build per Build ID in production.** Pointing the provider at a mutable "latest" target lets code change under in-flight Workflows and cause non-determinism errors, even for Pinned Workflows. Keep a 1-to-1 mapping between each Build ID and one immutable build. → `references/<provider>/versioning.md`. - **Tune the timeout triple together for long-running Activities:** (1) worker stop timeout > longest Activity runtime, (2) shutdown deadline buffer > worker stop timeout + shutdown hook time, (3) invocation deadline > longest Activity runtime + shutdown deadline buffer. Raising one alone does not help. If the longest Activity exceeds half the maximum invocation deadline, recommend Activity Heartbeats. → `references/concepts.md`, `references/<provider>/sdk-<language>.md`. - **Eager Activities are always disabled** — serverless invocations don't maintain persistent connections. Don't suggest them as an optimization. - **Activities are bounded by the invocation limit** (minus the shutdown deadline buffer); Workflow duration is unbounded and can span many invocations. Flag Activities that approach the provider's limit early. → `references/concepts.md`. - **Mixed serverless + long-lived Workers on one Task Queue:** do not enable dynamic scaling on the long-lived Workers — the two groups can't coordinate scaling and will cause unnecessary invocations. - **Secrets belong in a secret store**, not plaintext environment variables. Provider docs and quickstarts commonly pass the API key or TLS key as a plaintext environment variable; that is acceptable in a throwaway development walkthrough *only if you say so explicitly at the time*. Anything the user describes as production, shared, or long-lived gets the secret store, loaded at cold start. Either way, keep key material out of shell history and command echoes. - **Both CLIs prompt for confirmation before mutating state, and their flags differ.** Setting the current or ramping version, and revoking an API key, all ask interactively; run non-interactively without the flag, the command exits having done nothing, which reads as success. `temporal worker deployment …` takes `--yes`; `tcld` takes the global `--auto_confirm`. Pass the right one in scripts, CI, and agent shells, and confirm the resulting state rather than trusting the exit code. → `references/<provider>/setup.md`. ## Troubleshooting Start by determining whether the Worker is being invoked at all. Then, in priority order: (1) **Validate Connection** in the Temporal UI (Workers > Deployments > select > Actions > Validate Connection) — checks credentials, role assumption, and reachability in one step; (2) check whether the version's **Task Queue is bound** — if it is, invocation and Worker startup provably work and the fault is downstream, which rules out most of the surface in one command; (3) confirm the version is **current** (CLI-created versions are not automatic, and a confirmation-prompted command may have silently done nothing); (4) check the compute provider's logs for connection, auth, or TLS errors; (5) if rapid repeated invocations show no progress, check the deployment name/build ID match. Distinguish a Temporal-side failure (reproduces no matter what you change on the provider side) from a genuine user-permission problem before editing anything. → `references/<provider>/diagnostics.md`, `references/concepts.md`. ## Common Pitfalls High-impact mistakes — warn the user proactively. Each is a symptom → cause → fix. 1. **Deployment name / build ID mismatch → invocation loop.** *Symptom:* rapid, repeated invocations with no Workflow progress. *Cause:* the name or build ID in the Worker code doesn't match the Worker Deployment Version, so the Worker polls with the wrong version, the Task isn't processed, and Temporal invokes again. *Fix:* make the values in code exactly match the version configuration. 2. **Version not set as current.** A version created through the CLI is not automatically current; without it, Tasks don't route to the version and the Worker is never invoked. *Fix:* set it current as a separate step (the UI does this automatically). 3. **Failed first invocation.** When a version is created, the WCI invokes the Worker once to validate. If that invocation fails — missing env vars, bad TLS/auth config, missing dependencies, or an invocation deadline too short for the Worker to start and register the Task Queue — the Worker never connects, never polls, the binding is never created, and the Worker is never automatically invoked again. *Fix:* diagnose by manually invoking the compute unit, and confirm the invocation deadline is set high. 4. **Confusing the two roles.** The compute unit's execution role (grants the function permission to run) is separate from the access Temporal uses to invoke it. Never describe one as the other. → `references/<provider>/iam.md`. 5. **Timeout tuning mismatch.** Raising only the shutdown deadline buffer makes the Worker stop polling earlier but gives in-flight Activities no more time; raising only the worker stop timeout doesn't make it stop polling earlier, so the provider may terminate the Worker first. *Fix:* tune the three values together (see the timeout triple above). 6. **Mutable "latest" build reference in production.** Pointing the provider at a mutable/unqualified target means the code changes on every redeploy; deploying replay-unsafe code then causes non-determinism errors for in-flight Workflows, even Pinned ones. *Fix:* publish an immutable versioned build and keep a 1-to-1 mapping between each Build ID and one build. → `references/<provider>/versioning.md`. 7. **Re-creating shared permission infrastructure that already exists.** *Symptom:* the infrastructure deployment fails outright and rolls back, or it succeeds and leaves a second, redundant grant behind. *Cause:* the permission grant Temporal assumes is account-wide with a fixed default name, so a previous serverless deployment already owns it. *Fix:* check whether it exists and what owns it *before* creating; extend the existing one to cover the new Worker, and fall back to a distinctly named parallel one only when the existing infrastructure is not yours to change — saying why when you do. A failed-and-rolled-back deployment must be deleted before the name can be reused; a successful one is live infrastructure and must not be. → `references/<provider>/iam.md`. 8. **Invoke permission scoped to a single build.** *Symptom:* the deployment works, then the *next* release cannot be invoked, with an error that looks like a connection or configuration problem rather than a permissions one. *Cause:* the grant named one immutable build, and the new release is a different resource. *Fix:* scope the grant to cover the base resource and all its published builds. → `references/<provider>/iam.md`. ## Routing to reference files Most questions need 2–3 reference files. | User intent | Reference file(s) | |---|---| | What is a Serverless Worker / the WCI? How do invocation and autoscaling work? What are the constraints? Serverless vs long-lived Workers? | `references/concepts.md` | | Deploy a Serverless Worker (happy path): write code, package, deploy, register + set-current version, verify, tear down. | `references/<provider>/setup.md` + the selected `references/<provider>/sdk-<language>.md` (+ `references/concepts.md`) | | Operator permissions and preflight; execution role vs Temporal invocation role; CloudFormation (Cloud + self-hosted). | `references/<provider>/iam.md` | | Update or redeploy; version the build, use a qualified ARN, roll back. | `references/<provider>/versioning.md` (+ `references/concepts.md`) | | Self-hosted server enablement (dynamic config, WCI, server AWS credentials). | `references/<provider>/self-hosted.md` (+ `references/<provider>/iam.md`) | | Go SDK-specific options and tuned defaults, package and import, API inspection, handler, build and packaging, runtime and deployment values, versioning-behavior configuration, connection config, OpenTelemetry integration. | `references/<provider>/sdk-go.md` | | Python SDK-specific options and tuned defaults, package and import, API inspection, handler, build and packaging, runtime and deployment values, versioning-behavior configuration, connection config, OpenTelemetry integration, diagnostic signatures. | `references/<provider>/sdk-python.md` | | TypeScript SDK-specific options and tuned defaults, package and import, API inspection, handler, build and packaging, runtime and deployment values, versioning-behavior configuration, connection config, pre-bundled Workflow code, OpenTelemetry integration. | `references/<provider>/sdk-typescript.md` | | Java SDK-specific options and tuned defaults, artifact and imports, API inspection, handler, build and packaging, runtime and deployment values, versioning-behavior configuration, connection config, OpenTelemetry integration, logging and diagnostic signatures. | `references/<provider>/sdk-java.md` | | .NET SDK-specific options and tuned defaults, package and imports, API inspection, handler, RID-specific publish and packaging, runtime and deployment values, versioning-behavior configuration, connection config and `SSL_CERT_FILE`, OpenTelemetry integration, logging and diagnostic signatures. | `references/<provider>/sdk-dotnet.md` | | Add OpenTelemetry observability, Collector config, X-Ray, and IAM. | `references/<provider>/observability.md` + the selected `references/<provider>/sdk-<language>.md` | | Worker not invoked, Workflows not progressing, inspect the WCI. | `references/<provider>/diagnostics.md` + the selected `references/<provider>/sdk-<language>.md` (+ `references/concepts.md`) | | Long-running Activities and timeout relationships. Isolate Activities from resource exhaustion. | `references/concepts.md` (+ the selected `references/<provider>/sdk-<language>.md`) | ## Out of Scope - **General SDK development patterns** (Workflows, Activities, signals, queries, Worker Versioning concepts): see `skill-temporal-developer`. - **Traditional Worker tuning** (slot suppliers, tuners, poller autoscaling, resource-based tuning): see `skill-temporal-workertuning`. - **Temporal Cloud administration** (Namespaces, users, certificates, billing): see `skill-temporal-ops`. - **CLI command reference** (beyond the serverless-specific flags): see `skill-temporal-cli`.
Referenced files: 15
Package details
Publisher declarations from the archived package. These are separate from our research and the live service's terms.
- Package license
- MIT
- Package author
- Temporal
- Keywords
- temporal, workflow, durable-execution, python, typescript, go, java, microservices, distributed-systems
Package observed Oct 2, 2026.
Technical details
- First seen
- Sep 30, 2026 · 22:02 UTC
- Last seen
- Oct 2, 2026 · 00:00 UTC
- Collection status
- Collected
Plugin_3e13d15b5ae4819196b61eb770c858e8
Download plugin data (JSON)