{"id":23765,"plugin_id":"plugins_6ab2f25e4928819184294ebadcbe38ab","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T23:17:47.057Z","digest":"d3539dfc306d1bba6840ea8ce24c01aa8955f71e29b1aaaed3ce21185dec1144","against":null,"payload":{"name":"migrating-ai-sdk-to-common-ai","description":"Migrates Airflow projects from airflow-ai-sdk to apache-airflow-providers-common-ai 0.4.0+. Use when replacing airflow-ai-sdk with the official Airflow AI provider - migrating LLM decorators (@task.llm, @task.agent, @task.llm_branch, @task.embed), switching from model strings/objects to connection-based LLM configuration, updating imports from airflow_ai_sdk to the new provider, or upgrading an existing common-ai 0.1.x setup to 0.4.x (multimodal prompts, toolsets, embedding operators); also when common-ai provider, AIP-99, a pydanticai connection or migrating away from airflow-ai-sdk come up.","included_files":[],"skill_md_contents":"---\nname: migrating-ai-sdk-to-common-ai\ndescription: Migrates Airflow projects from airflow-ai-sdk to apache-airflow-providers-common-ai 0.4.0+. Use when replacing airflow-ai-sdk with the official Airflow AI provider - migrating LLM decorators (@task.llm, @task.agent, @task.llm_branch, @task.embed), switching from model strings/objects to connection-based LLM configuration, updating imports from airflow_ai_sdk to the new provider, or upgrading an existing common-ai 0.1.x setup to 0.4.x (multimodal prompts, toolsets, embedding operators); also when common-ai provider, AIP-99, a pydanticai connection or migrating away from airflow-ai-sdk come up.\n---\n\n# Migrate airflow-ai-sdk to apache-airflow-providers-common-ai\n\nThis skill migrates Airflow projects from `airflow-ai-sdk` to `apache-airflow-providers-common-ai` (target **0.4.0+**), the official Airflow AI provider built on PydanticAI. It also covers upgrading projects already on common-ai 0.1.x, since several capabilities (multimodal prompts, `toolsets`, embedding operators, structured-output XCom behavior) changed between 0.1.0 and 0.4.0.\n\n> **CRITICAL**: The new provider requires **Airflow 3.0+** and (for 0.4.0) **pydantic-ai-slim >= 1.71.0**. The API surface has changed: LLM configuration moves from code (model strings/objects) to Airflow connections (`pydanticai` type). There is no `@task.embed` in the new provider; embeddings move to the LlamaIndex integration or a plain `@task` (see Step 3).\n\n## Before starting\n\nUse the Grep tool with the pattern below to inventory everything that needs to migrate:\n\n```\nairflow_ai_sdk|airflow-ai-sdk|ai_sdk|@task\\.llm|@task\\.agent|@task\\.llm_branch|@task\\.embed\n```\n\nFrom the results, capture:\n\n1. All files importing `airflow-ai-sdk` / `airflow_ai_sdk`\n2. Which decorators are in use: `@task.llm`, `@task.agent`, `@task.llm_branch`, `@task.embed`\n3. The model configuration pattern (string names like `\"gpt-5\"`, or `OpenAIModel(...)` objects)\n4. Any `airflow_ai_sdk.BaseModel` subclasses used as `output_type`\n\nUse this inventory to drive the steps below.\n\n---\n\n## Step 1: Update requirements.txt\n\n**Remove:**\n```\nairflow-ai-sdk[openai]\n# or any variant: airflow-ai-sdk[openai]==0.1.7, airflow-ai-sdk[anthropic], etc.\n```\n\n**Add:**\n```\napache-airflow-providers-common-ai[openai]>=0.4.0\n```\n\nUse the latest available 0.x version unless the user has pinned a specific one. Available extras (0.4.0): `[openai]`, `[anthropic]`, `[google]`, `[bedrock]`, `[llamaindex]`, `[langchain]`, `[mcp]`, plus file-format extras (`[pdf]`, `[docx]`, `[parquet]`, `[avro]`) for `DocumentLoaderOperator` and `[sql]`/`[common-sql]` for the SQL operators. There are no `[groq]`/`[mistral]` extras; for those providers install the matching `pydantic-ai-slim` extra yourself.\n\nAdd `[llamaindex]` if the project migrates `@task.embed` to the `LlamaIndexEmbeddingOperator` (recommended, see Step 3). In that case `sentence-transformers` and `torch` can usually be **removed**, which shrinks the image considerably. Keep them only if the project stays on local sentence-transformers embeddings via plain `@task`.\n\n---\n\n## Step 2: Create PydanticAI connection\n\nThe new provider uses an Airflow connection instead of model strings or objects in code.\n\n**Connection type:** `pydanticai`\n**Default connection ID:** `pydanticai_default`\n\n### Via environment variable (.env)\n\n```bash\nAIRFLOW_CONN_PYDANTICAI_DEFAULT='{\n    \"conn_type\": \"pydanticai\",\n    \"password\": \"<api-key>\",\n    \"extra\": {\n        \"model\": \"<provider>:<model-name>\"\n    }\n}'\n```\n\n### Model format\n\nThe model field uses `provider:model` format:\n\n| Provider | Example model value |\n|----------|-------------------|\n| OpenAI | `openai:gpt-5` |\n| Anthropic | `anthropic:claude-sonnet-4-20250514` |\n| Google | `google:gemini-2.5-pro` |\n| Groq | `groq:llama-3.3-70b-versatile` |\n| Mistral | `mistral:mistral-large-latest` |\n| Bedrock | `bedrock:us.anthropic.claude-sonnet-4-20250514-v1:0` |\n\n### Custom endpoints (Ollama, vLLM, Snowflake Cortex, etc.)\n\nSet `host` to the base URL:\n```bash\nAIRFLOW_CONN_PYDANTICAI_CORTEX='{\n    \"conn_type\": \"pydanticai\",\n    \"password\": \"<api-key>\",\n    \"host\": \"https://my-endpoint.com/v1\",\n    \"extra\": {\n        \"model\": \"openai:<model-name>\"\n    }\n}'\n```\n\nUse the `openai:` prefix for any OpenAI-compatible API, regardless of the actual provider.\n\n### Connection ID convention\n\nThe env var name determines the connection ID:\n- `AIRFLOW_CONN_PYDANTICAI_DEFAULT` creates `pydanticai_default`\n- `AIRFLOW_CONN_PYDANTICAI_CORTEX` creates `pydanticai_cortex`\n\n### Model resolution priority\n\n1. `model_id` parameter on the decorator/operator (highest)\n2. `model` in connection's extra JSON (fallback)\n\n### Other connection types (0.4.0)\n\nBesides `pydanticai`, the provider registers vendor-specific connection types: `pydanticai-azure` (Azure OpenAI: host = endpoint, extra `api_version`), `pydanticai-bedrock` (AWS credentials/region in extra), and `pydanticai-vertex` (GCP project/location in extra). The LlamaIndex and LangChain hooks read API key/host/extra from whatever connection ID they are given, so a single `pydanticai_default` connection can serve LLM calls **and** embeddings: one API key entry for the whole project.\n\n---\n\n## Step 3: Migrate decorators\n\n### @task.llm\n\n```python\n# BEFORE (airflow-ai-sdk)\nimport airflow_ai_sdk as ai_sdk\n\nclass MyOutput(ai_sdk.BaseModel):\n    field: str\n\n@task.llm(\n    model=\"gpt-5\",                    # or model=OpenAIModel(...)\n    system_prompt=\"You are helpful.\",\n    output_type=MyOutput,\n)\ndef my_task(text: str) -> str:\n    return text\n\n# AFTER (apache-airflow-providers-common-ai)\nfrom pydantic import BaseModel\n\nclass MyOutput(BaseModel):\n    field: str\n\n@task.llm(\n    llm_conn_id=\"pydanticai_default\",  # Airflow connection ID\n    system_prompt=\"You are helpful.\",\n    output_type=MyOutput,\n)\ndef my_task(text: str) -> str:\n    return text\n```\n\n**Parameter mapping:**\n\n| airflow-ai-sdk | common-ai provider | Notes |\n|----------------|-------------------|-------|\n| `model=\"gpt-5\"` | `llm_conn_id=\"pydanticai_default\"` | Model specified in connection |\n| `model=OpenAIModel(...)` | `llm_conn_id=\"pydanticai_default\"` | Model + endpoint in connection |\n| `system_prompt=\"...\"` | `system_prompt=\"...\"` | Unchanged |\n| `output_type=MyModel` | `output_type=MyModel` | Unchanged |\n| `result_type=MyModel` | `output_type=MyModel` | `result_type` was already deprecated |\n| (not available) | `model_id=\"openai:gpt-5\"` | Override connection's model |\n| (not available) | `require_approval=True` | Built-in HITL review |\n| (not available) | `agent_params={...}` | Extra kwargs for pydantic-ai Agent |\n| (not available) | `serialize_output=True` | Force dict shape for BaseModel output |\n\n**Multimodal prompts (0.4.0+):** the translation function may return a `Sequence[UserContent]` instead of a string, e.g. for vision:\n\n```python\n@task.llm(llm_conn_id=\"pydanticai_default\", system_prompt=\"...\", output_type=ReviewAnalysis)\ndef analyze(text: str, image_path: str | None = None):\n    if image_path:\n        with open(image_path, \"rb\") as f:\n            return [text, BinaryContent(data=f.read(), media_type=\"image/jpeg\")]\n    return text\n```\n\nThis matches the old airflow-ai-sdk vision pattern, so vision code migrates unchanged. Note: common-ai **0.1.x only accepted strings** — if a project disabled vision to migrate to 0.1.0, re-enable it when bumping to 0.4.0. Non-string prompts are incompatible with `require_approval=True` / `enable_hitl_review=True` (both render the prompt as text).\n\n**Structured output via XCom (0.4.0 behavior change):** with `output_type=<BaseModel subclass>`, the model **instance** flows through XCom on Airflow cores whose task SDK has `SUPPORTS_OPERATOR_DESERIALIZATION_WALKER` (attribute access downstream); on older cores (including Astro Runtime 3.2 task SDK 1.2.x) the provider automatically dumps to a **dict** (subscript access). Check which shape arrives at runtime before choosing attribute vs dict access downstream, or set `serialize_output=True` to force the dict shape everywhere. The `output_type` class must be defined at **module scope** (nested classes cannot be deserialized from XCom).\n\n### @task.llm_branch\n\n```python\n# BEFORE\n@task.llm_branch(\n    model=\"gpt-5\",\n    system_prompt=\"Choose a team...\",\n    allow_multiple_branches=False,\n)\ndef route(text: str) -> str:\n    return text\n\n# AFTER\n@task.llm_branch(\n    llm_conn_id=\"pydanticai_default\",\n    system_prompt=\"Choose a team...\",\n    allow_multiple_branches=False,    # same parameter, unchanged\n)\ndef route(text: str) -> str:\n    return text\n```\n\nOnly change: `model=` becomes `llm_conn_id=`.\n\n### @task.agent\n\nThis has the biggest API change. The Agent is no longer pre-built in user code.\n\n```python\n# BEFORE (airflow-ai-sdk) - Agent built at module level\nfrom pydantic_ai import Agent\n\nmy_agent = Agent(\n    \"gpt-5\",\n    system_prompt=\"You are a research assistant.\",\n    tools=[search_tool, lookup_tool],\n)\n\n@task.agent(agent=my_agent)\ndef research(question: str) -> str:\n    return question\n\n# AFTER (common-ai provider) - No Agent object, config via parameters\nfrom pydantic_ai.toolsets import FunctionToolset\n\n@task.agent(\n    llm_conn_id=\"pydanticai_default\",\n    system_prompt=\"You are a research assistant.\",\n    toolsets=[FunctionToolset(tools=[search_tool, lookup_tool])],\n)\ndef research(question: str) -> str:\n    return question\n```\n\n**Parameter mapping:**\n\n| airflow-ai-sdk | common-ai provider | Notes |\n|----------------|-------------------|-------|\n| `agent=Agent(model, ...)` | `llm_conn_id=\"...\"` | Model from connection |\n| Agent's `system_prompt` | `system_prompt=\"...\"` | Now a decorator param |\n| Agent's `tools=[...]` | `toolsets=[FunctionToolset(tools=[...])]` | Preferred: gets automatic tool-call logging |\n| Agent's `tools=[...]` | `agent_params={\"tools\": [...]}` | Also works, but no tool-call logging |\n| Agent's `output_type` | `output_type=MyModel` | Now a decorator param |\n| (not available) | `durable=True` | Step-level caching (needs `[common.ai] durable_cache_path`) |\n| (not available) | `enable_hitl_review=True` | Iterative human review loop (see below) |\n\n**Key insight:** Everything that was configured on the `Agent()` constructor now goes into either a top-level decorator parameter or `agent_params`. The `agent_params` dict is passed directly to pydantic-ai's `Agent` constructor. Prefer `toolsets` over `agent_params[\"tools\"]`: the operator wraps each toolset in a `LoggingToolset`, so every tool call appears in the task log with timing.\n\n**enable_hitl_review behavior:** the task generates a first draft, then **blocks** until a human acts. The reviewer uses the **HITL Review** tab/extra link on the task instance (chat UI from the provider's auto-registered `hitl_review` plugin) to request changes (agent regenerates with the feedback in its message history) or approve. Constraints: requires a string prompt, incompatible with `durable=True`, and the final (possibly regenerated) output is what flows to XCom. Warn users that the Dag run waits indefinitely at this task unless `hitl_timeout` is set. For headless testing, the plugin exposes REST endpoints under `/hitl-review`: `GET /sessions/find`, `POST /sessions/feedback`, `POST /sessions/approve`, `POST /sessions/reject` (query params `dag_id`, `task_id`, `run_id`, `map_index`).\n\n### @task.embed (NO EQUIVALENT — three replacement options)\n\nThe new provider does NOT include an embed decorator. Pick the replacement based on what the project needs:\n\n**Option A (recommended): `LlamaIndexEmbeddingOperator`** (0.4.0, `[llamaindex]` extra). Connection-based, one task embeds the whole document list, and with `persist_dir` the resulting vector index is persisted for retrieval (pairs with `LlamaIndexRetrievalOperator`):\n\n```python\nfrom airflow.providers.common.ai.operators.llamaindex_embedding import LlamaIndexEmbeddingOperator\n\n_embeddings = LlamaIndexEmbeddingOperator(\n    task_id=\"create_embeddings\",\n    documents=[{\"text\": \"...\", \"metadata\": {\"id\": 1}}, ...],  # templated, accepts XComArg\n    llm_conn_id=\"pydanticai_default\",   # reuses the same connection (API key only)\n    embed_model=\"text-embedding-3-small\",\n    persist_dir=f\"{AIRFLOW_HOME}/include/my_index\",  # optional; local path or s3://, gs://, ...\n)\n```\n\nThe operator returns `{\"chunks\": [{\"text\", \"metadata\", \"vector\"}], ...}`. Put a stable key into each document's `metadata` — it round-trips through chunking, so vectors can be mapped back to source records.\n\n**Option B: `LlamaIndexHook` for raw vectors** (no operator, no persisted index). Shortest path when vectors go straight to a database:\n\n```python\n@task\ndef create_embeddings(rows):\n    from airflow.providers.common.ai.hooks.llamaindex import LlamaIndexHook\n    embed_model = LlamaIndexHook(\n        llm_conn_id=\"pydanticai_default\",\n        embed_model=\"text-embedding-3-small\",\n    ).get_embedding_model()\n    vectors = embed_model.get_text_embedding_batch([r[\"text\"] for r in rows])\n    return list(zip([r[\"id\"] for r in rows], vectors))\n```\n\n**Option C: plain `@task` with sentence-transformers** (keeps the old local/offline behavior, no API cost; requires keeping `sentence-transformers` + `torch` in requirements):\n\n```python\n@task\ndef embed_texts(texts: list[str]) -> list[list[float]]:\n    from sentence_transformers import SentenceTransformer\n    model = SentenceTransformer(\"all-MiniLM-L6-v2\")\n    return model.encode(texts, normalize_embeddings=True).tolist()\n```\n\nNote on dimensions: switching from `all-MiniLM-L6-v2` (384) to `text-embedding-3-small` (1536) changes vector size — existing stored embeddings must be regenerated, and fixed-size vector columns (e.g. pgvector `vector(384)`) need a schema change. Embed all texts in one task/batch call rather than `.expand()` per text: batching is one API round-trip and avoids per-task model loading.\n\n---\n\n## Step 4: Update imports\n\n| Old import | New import |\n|-----------|-----------|\n| `import airflow_ai_sdk as ai_sdk` | Remove entirely |\n| `from airflow_ai_sdk import BaseModel` | `from pydantic import BaseModel` |\n| `from airflow_ai_sdk.models.base import BaseModel` | `from pydantic import BaseModel` |\n| `class Foo(ai_sdk.BaseModel):` | `class Foo(BaseModel):` |\n| `from pydantic_ai import Agent` | Remove if Agent was only used for `@task.agent` |\n| `from pydantic_ai.models.openai import OpenAIModel` | Remove (model config in connection now) |\n| (new) | `from pydantic_ai.toolsets import FunctionToolset` for `@task.agent` toolsets |\n\nThe `@task.llm`, `@task.agent`, `@task.llm_branch` decorators are auto-registered by the provider. No explicit import needed beyond `from airflow.sdk import task`.\n\n`pydantic_ai` imports for non-decorator usage (e.g., `BinaryContent` for multimodal) are still valid since the new provider depends on `pydantic-ai-slim` (>= 1.71.0 for provider 0.4.0).\n\n---\n\n## Step 5: Update connections.yaml (if used for local testing)\n\n```yaml\npydanticai_default:\n  conn_type: pydanticai\n  password: <api-key>\n  extra:\n    model: \"openai:gpt-5\"\n```\n\nFor custom endpoints:\n```yaml\npydanticai_cortex:\n  conn_type: pydanticai\n  password: <api-key>\n  host: https://my-endpoint.com/v1\n  extra:\n    model: \"openai:llama3.1-8b\"\n```\n\n---\n\n## Step 6: Clean up env vars\n\nThe new provider reads model config from the `pydanticai` connection, so env vars that previously fed the model in code are usually redundant. Before removing any of them, grep the project (and any sibling scripts/services) to confirm nothing else still references them:\n\n```\nOPENAI_API_KEY|OPENAI_BASE_URL|ANTHROPIC_API_KEY|GOOGLE_API_KEY\n```\n\nCandidates for removal **only if no other code references them**:\n- `OPENAI_API_KEY` (now in the pydanticai connection's password field)\n- `OPENAI_BASE_URL` (now in the connection's host field)\n- Custom model name vars (now in the connection's extra.model)\n\nIf anything outside the migrated DAGs still uses them (other DAGs not yet migrated, helper scripts, non-Airflow services sharing the `.env`), leave them in place.\n\n**Keep** `AIRFLOW_CONN_*` env vars for all connections.\n\n---\n\n## Step 7: Verify\n\nAfter migration, grep the codebase to confirm no stale references remain:\n\n```\nairflow_ai_sdk|airflow-ai-sdk|ai_sdk\\.BaseModel|from pydantic_ai import Agent|from pydantic_ai.models\n```\n\nVerify:\n- [ ] No imports from `airflow_ai_sdk`\n- [ ] No `Agent()` objects created for `@task.agent` (unless used outside decorators)\n- [ ] No `model=` parameter on LLM decorators (should be `llm_conn_id=`)\n- [ ] All `@task.embed` replaced (LlamaIndex operator/hook or plain `@task`); stored embeddings regenerated if the model/dimensions changed\n- [ ] Vision translation functions return `[text, BinaryContent(...)]` again if they were string-only-restricted under common-ai 0.1.x\n- [ ] Downstream consumers of `output_type=BaseModel` results use the XCom shape that actually arrives (dict on older cores, instance on newer; `serialize_output=True` pins it)\n- [ ] `pydanticai` connection configured in `.env` or connections.yaml\n- [ ] `requirements.txt` has `apache-airflow-providers-common-ai[...]` instead of `airflow-ai-sdk[...]`; `torch`/`sentence-transformers` removed if no longer used\n- [ ] Run the Dags end-to-end: tasks with `enable_hitl_review=True` or `require_approval=True` wait for human input, so the test plan must include acting on them (UI tab or `/hitl-review` REST)\n\n---\n\n## Quick reference: New features in common-ai provider\n\nThese features are available after migration but have no airflow-ai-sdk equivalent:\n\n| Feature | Parameter / API | Since | Description |\n|---------|-----------------|-------|-------------|\n| HITL approval | `require_approval=True` on `@task.llm` | 0.1.0 | Pause for human review before returning |\n| HITL review loop | `enable_hitl_review=True` on `@task.agent` | 0.1.0 | Iterative review with regeneration (chat UI via `hitl_review` plugin) |\n| Durable execution | `durable=True` on `@task.agent` | 0.1.0 | Step-level caching for resilience |\n| Tool logging | `enable_tool_logging=True` on `@task.agent` | 0.1.0 | INFO-level tool call logs (default: on; requires `toolsets`) |\n| Model override | `model_id=\"openai:gpt-5\"` | 0.1.0 | Override connection's model per-task |\n| File analysis | `@task.llm_file_analysis` | 0.1.0 | Analyze files/images via ObjectStoragePath |\n| NL-to-SQL | `@task.llm_sql` | 0.1.0 | Generate SQL from natural language |\n| Multimodal prompts | Translation function returns `Sequence[UserContent]` | 0.4.0 | Vision and other binary content in `@task.llm` / `@task.agent` / `@task.llm_branch` |\n| Pydantic instance via XCom | `output_type=BaseModel` (with `serialize_output` opt-out) | 0.4.0 | Instance flows through XCom on capable cores; dict fallback otherwise |\n| Embeddings | `LlamaIndexEmbeddingOperator` (+ `persist_dir`) | 0.4.0 | Connection-based embeddings + persisted vector index |\n| Retrieval | `LlamaIndexRetrievalOperator` | 0.4.0 | Top-k similarity search over a persisted index |\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}