{"id":11248,"plugin_id":"plugin_asdk_app_6a79985b0ef881918e389c82400ea95c","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:59:03.338Z","digest":"95306006c80dd111b230068d71183a6e9e977eadec6837eee8bacc20c69a5443","against":null,"payload":{"description":"Runs Flyte 2 workflows, interacts with runs and actions, retrieves logs and data, and manages run lifecycle. Use when the user wants to run a workflow, check run status, view logs, get run outputs, re-run a workflow, or manage runs programmatically. Trigger words: \"run\", \"execute\", \"logs\", \"status\", \"output\", \"input\", \"watch\", \"rerun\", \"cancel\", \"abort\", \"run metadata\", \"action\".","included_files":[],"name":"flyte-sdk-run","skill_md_contents":"---\nname: flyte-sdk-run\ndescription: 'Runs Flyte 2 workflows, interacts with runs and actions, retrieves logs and data, and manages run lifecycle. Use when the user wants to run a workflow, check run status, view logs, get run outputs, re-run a workflow, or manage runs programmatically. Trigger words: \"run\", \"execute\", \"logs\", \"status\", \"output\", \"input\", \"watch\", \"rerun\", \"cancel\", \"abort\", \"run metadata\", \"action\".'\n---\n\n# Flyte 2 SDK Run Skill\n\nRun workflows, interact with runs, and manage the execution lifecycle.\n\n## Grounding References\n\n| Resource | URL |\n|---|---|\n| Official docs | https://www.union.ai/docs/v2/flyte |\n| Docs index (LLMs) | https://www.union.ai/docs/v2/flyte/llms.txt |\n| SDK API reference | https://www.union.ai/docs/v2/union/api-reference/flyte-sdk/ |\n| CLI API reference | https://www.union.ai/docs/v2/union/api-reference/flyte-cli/ |\n| flyte-sdk source | https://github.com/flyteorg/flyte-sdk |\n| Example code | https://github.com/unionai/unionai-examples |\n| Flyte MCP tools | Available via the `flyte-cluster` and `flyte-docs` MCP servers |\n\n## Tool Priority\n\n1. **Flyte MCP** — if the harness has Flyte MCP tools, prefer them over shelling out to\n   the CLI. They cover listing runs, fetching run details and inputs/outputs, polling to\n   completion, executing a task, and aborting a run, and they return structured data\n   instead of text you have to parse.\n2. **`flyte` CLI** — for local run commands, and anything MCP does not expose\n3. **Python SDK** — for programmatic run control\n\n## Running Workflows\n\n### Via Python SDK\n\n```python\nimport flyte\n\nif __name__ == \"__main__\":\n    # Run with defaults from config\n    result = flyte.run(main, inputs={\"data\": [\"a\", \"b\", \"c\"]})\n    print(f\"Run name: {result.name}\")\n    print(f\"Status: {result.status}\")\n```\n\n### Via CLI\n\n```bash\n# Run with local config\nflyte run pipeline.py main --data '[1,2,3]'\n\n# Run with specific project/domain\nflyte run pipeline.py main --data '[1,2,3]' --project flytesnacks --domain development\n\n# Run with custom run name\nflyte run pipeline.py main --data '[1,2,3]' --name my-custom-run\n\n# Run with specific image\nflyte run pipeline.py main --data '[1,2,3]' --image ghcr.io/myorg/task:v1.0\n\n# Run with local mode (in-process, no remote)\nflyte run --local pipeline.py main --data '[1,2,3]'\n\n# Run with TUI\nflyte run --tui --local pipeline.py main --data '[1,2,3]'\n\n# Pass arguments by type\nflyte run pipeline.py main \\\n  --data '[1,2,3]' \\\n  --learning-rate 0.001 \\\n  --batch-size 32 \\\n  --train-data s3://bucket/train.parquet \\\n  --flag true\n```\n\n### Run command options\n\n| Flag | Description |\n|---|---|\n| `--project` / `--domain` | Target project and domain |\n| `--run-project` / `--run-domain` | Override run project/domain |\n| `--local` | Run locally (in-process) |\n| `--tui` | Terminal UI for local runs |\n| `--name` | Custom run name |\n| `--image` | Image mapping (named or default) |\n| `--copy-style` | `loaded_modules` (default), `all`, `none` |\n| `--root-dir` | Set root directory for code bundling |\n| `--raw-data-path` | Override raw data path |\n| `--service-account` | K8s service account |\n| `--follow` | Follow run progress |\n| `--no-sync-local-sys-paths` | Skip local sys path sync |\n\n### Passing inputs by type\n\n```bash\n# List\nflyte run pipeline.py main --data '[1,2,3]'\n\n# Dict\nflyte run pipeline.py main --config '{\"lr\": 0.001, \"epochs\": 10}'\n\n# Boolean\nflyte run pipeline.py main --flag true\n\n# Datetime\nflyte run pipeline.py main --date '2025-01-01T00:00:00'\n\n# Duration\nflyte run pipeline.py main --timeout '1h'\n\n# File\nflyte run pipeline.py main --input-file s3://bucket/data.parquet\n\n# DataFrame (via file path)\nflyte run pipeline.py main --data-file /path/to/data.parquet\n```\n\n## Interacting with Runs\n\n### Using Flyte MCP\n\nIf Flyte MCP tools are available, prefer them for all of the above — listing runs,\nfetching a run's details, polling until it completes, and reading its inputs and outputs.\n\n\n### Using CLI\n\n```bash\n# List runs\nflyte get run --project flytesnacks --domain development\n\n# Get run info\nflyte get run <run_name> --project flytesnacks --domain development\n\n# Watch run progress\nflyte get run <run_name> --project flytesnacks --domain development\n\n# Get run outputs\nflyte get io <run_name> --project flytesnacks --domain development\n\n# Download run artifacts\nflyte get io <run_name> --outputs-only --project flytesnacks --domain development\n```\n\n### Using Python SDK\n\n```python\nimport flyte\n\n# Run and get handle\nresult = flyte.run(main, inputs={\"data\": [\"a\", \"b\"]})\n\n# Check status\nprint(result.status)  # RUNNING, SUCCEEDED, FAILED, CANCELED\n\n# Wait for completion\nresult.wait()\n\n# Get outputs\nprint(result.outputs)\n\n# Get URL in console\nprint(result.url)\n```\n\n## Viewing Logs\n\n### Using CLI\n\n```bash\n# Stream logs\nflyte get logs <run_name> --project flytesnacks --domain development\n\n# View logs for a specific attempt\nflyte get logs <run_name> --attempt 0\n\n# Filter system logs\nflyte get logs <run_name> --filter-system\n\n# Scope to project/domain\nflyte get logs <run_name> --project flytesnacks --domain development\n```\n\n### CLI log options\n\n| Flag | Description |\n|---|---|\n| `--attempt` / `-a` | View specific attempt logs |\n| `--filter-system` | Filter out system logs |\n| `--pretty` | Auto-scrolling box (limited to `--lines`) |\n| `--project` / `--domain` | Scope logs |\n\n### Using Python SDK\n\n```python\nimport flyte\n\nresult = flyte.run(main, inputs={\"data\": [\"a\"]})\n\n# Logs are retrieved via the CLI: `flyte get logs <run_name>`\nprint(result.url)  # open the run in the UI to view logs\n```\n\n## Re-running Runs\n\n### CLI\n\n```bash\n# Re-run with original code and inputs\nflyte rerun <run_name> --project flytesnacks --domain development\n\n# Re-run with new local code\nflyte run --rerun-from <run_name> pipeline.py main --data '[4,5,6]'\n```\n\n### Python SDK\n\n```python\nimport flyte\n\n# Re-run with new inputs\nresult = flyte.run(\n    main,\n    inputs={\"data\": [4, 5, 6]},\n    run_context=flyte.with_runcontext(run_name=\"rerun-of-abc123\"),\n)\n```\n\n## Running Tasks (vs Workflows)\n\n### Run a single task\n\n```bash\n# Ephemeral run (deploy + run in one command)\nflyte run pipeline.py preprocess --data '[1,2,3]'\n\n# Run a deployed task\nflyte run --task-name preprocess --project flytesnacks --domain development \\\n  --inputs '{\"data\": \"[1,2,3]\"}'\n```\n\n### Using Flyte MCP\n\nExecuting a registered task is available as an MCP tool, taking project, domain, task name,\nversion, and inputs.\n\n\n## Run Context Configuration\n\n### Programmatic run context\n\n```python\nimport flyte\n\n# Configure a run programmatically\nresult = flyte.run(\n    main,\n    inputs={\"data\": [\"a\", \"b\"]},\n    run_context=flyte.with_runcontext(\n        project=\"flytesnacks\",\n        domain=\"development\",\n        raw_data_path=\"s3://my-bucket/{run_id}/\",\n        service_account=\"my-sa\",\n    ),\n)\n```\n\n### Reading run context inside a task\n\n```python\n@env.task\nasync def my_task(data: str) -> str:\n    # Access run metadata inside the task\n    ctx = flyte.ctx()\n    print(f\"Run: {ctx.run_id}\")\n    print(f\"Project: {ctx.project}\")\n    print(f\"Domain: {ctx.domain}\")\n    print(f\"Version: {ctx.version}\")\n    return data\n```\n\n## Abort and Cancel Runs\n\n### CLI\n\n```bash\n# Abort a run\nflyte abort run <run_name> --project flytesnacks --domain development\n```\n\n### Python SDK\n\n```python\nimport flyte\n\nresult = flyte.run(main, inputs={\"data\": [\"a\"]})\nresult.abort()\n```\n\n### Using Flyte MCP\n\nAborting a run is available as an MCP tool, taking the run name.\n\n\n## Programmatic Abort from Within a Task\n\n```python\n@env.task\nasync def long_task(data: str) -> str:\n    import asyncio\n    import signal\n\n    async def check_abort():\n        while True:\n            if asyncio.current_task().cancelled():\n                raise asyncio.CancelledError(\"Run was aborted\")\n            await asyncio.sleep(1)\n\n    # Start abort watcher\n    watcher = asyncio.create_task(check_abort())\n\n    try:\n        # Long-running work\n        await asyncio.sleep(3600)\n    finally:\n        watcher.cancel()\n        await watcher\n\n    return data\n```\n\n## Run Data Access\n\n### Accessing large data from cloud storage\n\n```python\nimport flyte\nimport flyte.io\n\n@env.task\nasync def get_run_data(run_name: str) -> flyte.io.File:\n    \"\"\"Download artifacts from a past run.\"\"\"\n    # Flyte stores outputs in the metadata bucket\n    # Access via the SDK's data retrieval methods\n    ...\n\n@env.task\nasync def upload_local_data(file_path: str) -> flyte.io.File:\n    \"\"\"Upload local file to remote storage for a run.\"\"\"\n    return flyte.io.File(path=file_path)\n```\n\n### S3 / GCS / Azure access\n\n```python\n# S3\nimport boto3\ns3 = boto3.client(\"s3\")\nobj = s3.get_object(Bucket=\"my-bucket\", Key=\"run-artifacts/output.parquet\")\n\n# GCS\nfrom google.cloud import storage\nclient = storage.Client()\nbucket = client.bucket(\"my-bucket\")\nblob = bucket.blob(\"run-artifacts/output.parquet\")\n\n# Azure\nfrom azure.storage.blob import BlobServiceClient\nclient = BlobServiceClient(account_url=\"https://myacct.blob.core.windows.net/\")\nblob = client.get_blob_client(container=\"my-container\", blob=\"output.parquet\")\n```\n\n## Run Modes\n\n### Local execution\n\n```bash\n# In-process (no remote backend needed)\nflyte run --local pipeline.py main --data '[1,2,3]'\n\n# With TUI\nflyte run --tui --local pipeline.py main --data '[1,2,3]'\n```\n\n### Devbox\n\n```bash\n# Start local dev environment\nflyte start devbox\n\n# Create config for devbox\nflyte create config \\\n    --endpoint localhost:30080 \\\n    --project flytesnacks \\\n    --domain development \\\n    --builder local \\\n    --insecure\n\n# Run on devbox\nflyte run pipeline.py main --data '[1,2,3]'\n```\n\n### Remote execution\n\n```bash\n# Create config for remote backend\nflyte create config \\\n    --endpoint <host> \\\n    --project flytesnacks \\\n    --domain development \\\n    --builder local \\\n    --insecure\n\n# Run on remote backend\nflyte run pipeline.py main --data '[1,2,3]'\n```\n\n## Anti-Patterns\n\n1. **Don't confuse `flyte run` (workflow) with `flyte run --task-name` (single task)** — use the right command for your intent.\n2. **Don't skip `--follow`** when running long workflows — you won't see progress.\n3. **Don't hardcode run names** — let Flyte generate them, or use meaningful prefixes.\n4. **Don't access run data directly from S3/GCS** — use Flyte's data retrieval methods when possible.\n5. **Don't use Union-only features** — avoid `ReusePolicy` and other Union-specific APIs.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}