{"id":11232,"plugin_id":"plugin_asdk_app_6a79985b0ef881918e389c82400ea95c","kind":"skill","collection_source":null,"comparison_source":null,"observed_at":"2026-09-30T22:59:03.048Z","digest":"b696332fae5fd18bb90ae77377a88bf94f67208de2078d8cafc63764f981f57f","against":null,"payload":{"description":"Use when deploying a Flyte v2 (flyte-binary / flyte2) cluster on AWS from scratch — provisions EKS + S3 + RDS PostgreSQL + AWS Load Balancer Controller, then helm-installs the flyte-binary chart behind an ALB, with optional TLS and Okta SSO. Trigger words: \"deploy flyte\", \"flyte v2 on AWS\", \"flyte EKS\".","included_files":[],"name":"flyte-deploy-aws","skill_md_contents":"---\nname: flyte-deploy-aws\ndescription: 'Use when deploying a Flyte v2 (flyte-binary / flyte2) cluster on AWS from scratch — provisions EKS + S3 + RDS PostgreSQL + AWS Load Balancer Controller, then helm-installs the flyte-binary chart behind an ALB, with optional TLS and Okta SSO. Trigger words: \"deploy flyte\", \"flyte v2 on AWS\", \"flyte EKS\".'\n---\n\n# Deploying Flyte v2 on AWS (EKS + RDS + S3 + ALB)\n\nFlyte v2 ships as a single unified binary (`flyte-binary-v2`) plus a separate console\nimage. One HTTP ingress serves the console (`/v2`), the `flyteidl2.*` Connect API, and\nauth-discovery — there is no separate gRPC port. You scale it vertically.\n\nThe chart does NOT provision infrastructure. Stand up four things first:\n**EKS cluster, S3 bucket, PostgreSQL (RDS), and (for external access) an ingress\ncontroller.** This skill does all four with `eksctl` + `aws` + `helm`, then installs the\n`flyte-binary` chart (the v2 chart; defaults to `flyte-binary-v2` + `flyteconsole-v2`).\n\n**Get the chart first.** Install from the published Helm repo — this is what the official\ndocs do, and the released chart pins the Flyte image tag to the chart version (see Image\nselection in Step 5). Also `helm pull --untar` a local copy so the TaskAction CRD file is on\ndisk for the Step 5 idempotency check:\n```bash\nhelm repo add flyteorg https://flyteorg.github.io/flyte && helm repo update\nhelm pull flyteorg/flyte-binary --untar   # ./flyte-binary/templates/crds/flyte.org_taskactions.yaml now resolves\n```\n(Alternatively clone the repo — `git clone https://github.com/flyteorg/flyte` — and install\nfrom the local `./charts/flyte-binary` path for the bleeding-edge chart; its default image tag\nis a floating `:latest`, so pair it with `pullPolicy: Always` — see Image selection.) Official\ndocs: https://www.union.ai/docs/v2/flyte/oss-deployment/aws-deployment/. Validated end-to-end\non EKS.\n\n> Replace every placeholder in angle brackets and the example hostnames/IDs with your own.\n\n## Prerequisites & decisions\n\n- CLIs: `aws` v2, **`eksctl` ≥ 0.227** (older caps out at k8s 1.29 — see gotcha), `kubectl`, `helm`, `jq`.\n- Admin (or EKS+RDS+IAM+S3+EC2) creds. STS/SSO works — export the 3 env vars + region.\n- eksctl writes the kubeconfig context (e.g. `<user>@flyte-v2.<region>.eksctl.io`). Pass\n  `kubectl --context <ctx>` (and `helm --kube-context <ctx>`) per command rather than\n  `kubectl config use-context` — that way you don't mutate the operator's current context.\n- Decide: region, name prefix, and **exposure**: ALB+TLS needs a Route53 zone + ACM cert;\n  **ALB HTTP-only needs neither** (reached at the auto `*.elb.amazonaws.com` name) — the\n  simplest default when you own no domain. (This skill provisions **RDS PostgreSQL** for the DB;\n  an in-cluster Postgres is out of scope here — RDS is assumed by Steps 3–5.)\n\n**Persist your variables.** This deploy spans many commands and derives values you can't\nrecover later — most critically the **random `DBPW`** (Step 3), plus `ACCT`, `BUCKET`,\n`RDS_HOST`, `IRSA_ARN`, etc. If your shell state resets between steps (or your AWS session\ntoken expires and you re-auth in a fresh shell), these are gone. Keep them in a file you\nre-`source` at the start of every step, and append each derived value as you compute it:\n```bash\nexport AWS_ACCESS_KEY_ID=... AWS_SECRET_ACCESS_KEY=... AWS_SESSION_TOKEN=...\nexport AWS_DEFAULT_REGION=us-west-2\nENVF=~/flyte-deploy.env                                   # source this at every step\n{ echo \"export PREFIX=flyte-v2 REGION=us-west-2 CLUSTER=flyte-v2\"\n  echo \"export ACCT=$(aws sts get-caller-identity --query Account --output text)\"  # confirm the RIGHT account\n} >> $ENVF && source $ENVF\n# As you create infra, append its outputs, e.g.:  echo \"export DBPW='$DBPW' RDS_HOST=$RDS_HOST\" >> $ENVF\n# Check for an existing domain/cert (empty => go ALB HTTP-only):\naws route53 list-hosted-zones --query 'HostedZones[].Name' --output text\naws acm list-certificates --region $REGION --query 'CertificateSummaryList[].DomainName' --output text\n```\n\n## Step 0 — Reuse an existing cluster?\n\nBefore creating anything, list the EKS clusters already in the account/region and **ask the\nuser whether to deploy onto one of them or stand up a fresh cluster**. Reusing skips Step 1\n(~15-20 min + the EKS control-plane + node cost).\n\n```bash\naws eks list-clusters --region $REGION --query 'clusters' --output text\n```\n\nPresent the list and let the user pick one (or choose \"create new\"). If they reuse one:\n\n```bash\nCLUSTER=<chosen-cluster>\naws eks update-kubeconfig --region $REGION --name $CLUSTER --alias $CLUSTER   # writes + selects context\nkubectl --context $CLUSTER get nodes                                          # confirm reachable + Ready\n# Confirm IRSA is possible (the chart needs an OIDC provider on the cluster):\naws eks describe-cluster --region $REGION --name $CLUSTER \\\n  --query 'cluster.identity.oidc.issuer' --output text                        # empty => run: eksctl utils associate-iam-oidc-provider --cluster $CLUSTER --approve\n```\n\nThen **skip Step 1** and continue from Step 2. S3 (Step 2), RDS (Step 3), and the ALB\ncontroller (Step 4) may already exist on a reused cluster — check before recreating\n(`aws s3 ls`, `aws rds describe-db-instances`, `kubectl --context $CLUSTER -n kube-system get deploy aws-load-balancer-controller`)\nand reuse what's there. Otherwise proceed normally. Pass `--context $CLUSTER` /\n`--kube-context $CLUSTER` on the later kubectl/helm commands.\n\n## Step 0.5 — Confirm deployment parameters (ASK up front, never assume)\n\n**Before provisioning or installing anything, gather the deploy parameters by ASKING the\nuser — do NOT silently reuse values you happen to find.** A previous deploy leaves identifiers\nlying around (an old `values-eks.yaml` with `HOST=`/`certificate-arn`/`password`, a live\n`flyte*-console-oidc` k8s Secret, `authMetadata.flyteClient.clientId`, a memory of the last\nrun). These are **suggestions to confirm, not defaults.** Silently reusing the prior\nhostname, OIDC client ID/secret, or cert is the #1 way this skill does the wrong thing.\n\nFor each parameter below, **discover any prior value, then present it as a choice** — e.g.\n\"reuse previous (`test.uniondemo.run`, loaded from the old values file / the in-cluster\nSecret), enter a new one, or pick a different existing one\" — and let the user decide. Restate\nthe final set back to them before `helm install`.\n\n| Parameter | Where a prior value hides | Notes |\n|---|---|---|\n| Region / name prefix / cluster | Step 0, current kube-context | |\n| Exposure (HTTP-only / TLS / TLS+SSO) | — | drives which params below apply |\n| Hostname | `HOST=` in old `values-eks.yaml`; existing Route53 record | drives cert, OIDC redirect URI, DNS |\n| ACM cert ARN | old values `certificate-arn`; `aws acm list-certificates` | must match the chosen hostname |\n| OIDC issuer / client ID / **client secret** | `authMetadata` in old values; `flyte*-console-oidc` Secret; the IdP app | **never echo/ask for the secret in chat** — have the user create the Secret themselves (see ALB edge SSO) |\n| OIDC CLI/PKCE client ID | `authMetadata.flyteClient.clientId` | |\n| S3 bucket / RDS host+password | Step 2/3 outputs; old values | reuse the live infra's real values |\n\nOnly after the user confirms each value do you write `values-eks.yaml` (Step 5). If reusing a\nsecret/credential, confirm the user still wants *that* IdP app — switching IdP is just a new\nSecret + issuer refs (no ALB/DNS churn).\n\n## Step 1 — EKS cluster (eksctl)\n\n`cluster.yaml` — `iam.withOIDC: true` is what makes IRSA possible:\n\n```yaml\napiVersion: eksctl.io/v1alpha5\nkind: ClusterConfig\nmetadata: { name: flyte-v2, region: us-west-2, version: \"1.33\" }\niam: { withOIDC: true }\nmanagedNodeGroups:\n  - name: ng-default\n    instanceType: m5.large\n    desiredCapacity: 2\n    minSize: 2\n    maxSize: 3\n    volumeSize: 50\n    iam: { withAddonPolicies: { ebs: true } }\naddons: [{name: vpc-cni},{name: coredns},{name: kube-proxy},{name: aws-ebs-csi-driver}]\n```\n\n```bash\neksctl create cluster -f cluster.yaml     # ~15-20 min; writes kubeconfig + sets context\nkubectl get nodes                          # expect Ready\n```\n\nThe VPC + private subnets exist within ~2 min (before the control plane finishes), so you\ncan start RDS (step 3) in parallel.\n\n## Step 2 — S3 bucket + IRSA role\n\n```bash\nBUCKET=$PREFIX-data-$ACCT       # account-id suffix => globally unique\naws s3api create-bucket --bucket $BUCKET --region $REGION \\\n  --create-bucket-configuration LocationConstraint=$REGION\naws s3api put-public-access-block --bucket $BUCKET --public-access-block-configuration \\\n  BlockPublicAcls=true,IgnorePublicAcls=true,BlockPublicPolicy=true,RestrictPublicBuckets=true\naws s3api put-bucket-encryption --bucket $BUCKET --server-side-encryption-configuration \\\n  '{\"Rules\":[{\"ApplyServerSideEncryptionByDefault\":{\"SSEAlgorithm\":\"AES256\"}}]}'\n\n# Scoped S3 policy: ListBucket on the bucket, Get/Put/Delete on its objects.\ncat > s3-policy.json <<EOF\n{\"Version\":\"2012-10-17\",\"Statement\":[\n {\"Effect\":\"Allow\",\"Action\":[\"s3:ListBucket\"],\"Resource\":\"arn:aws:s3:::$BUCKET\"},\n {\"Effect\":\"Allow\",\"Action\":[\"s3:GetObject\",\"s3:PutObject\",\"s3:DeleteObject\"],\"Resource\":\"arn:aws:s3:::$BUCKET/*\"}]}\nEOF\nPOLICY_ARN=$(aws iam create-policy --policy-name $PREFIX-s3-access \\\n  --policy-document file://s3-policy.json --query Policy.Arn --output text)\n\n# --role-only: create the IAM role with OIDC trust for system:serviceaccount:flyte:flyte,\n# but NOT the k8s SA (the chart creates+annotates it). Works before the ns exists.\neksctl create iamserviceaccount --cluster $CLUSTER --region $REGION \\\n  --namespace flyte --name flyte --role-name $PREFIX-irsa \\\n  --attach-policy-arn \"$POLICY_ARN\" --role-only --approve\nIRSA_ARN=$(aws iam get-role --role-name $PREFIX-irsa --query Role.Arn --output text)\n```\n\n## Step 3 — RDS PostgreSQL  (can run in parallel with step 1)\n\n```bash\nVPC=$(aws ec2 describe-vpcs --region $REGION \\\n  --filters \"Name=tag:alpha.eksctl.io/cluster-name,Values=$CLUSTER\" --query 'Vpcs[0].VpcId' --output text)\n# Private subnets (internal-elb role tag):\nSUBNETS=$(aws ec2 describe-subnets --region $REGION --filters \"Name=vpc-id,Values=$VPC\" \\\n  \"Name=tag:kubernetes.io/role/internal-elb,Values=1\" --query 'Subnets[].SubnetId' --output text)\n# CRITICAL: source SG must be the EKS-managed cluster SG actually on the NODES,\n# NOT ClusterSharedNodeSecurityGroup. Pod egress uses the node primary-ENI SG.\naws rds create-db-subnet-group --region $REGION --db-subnet-group-name $PREFIX-db-subnets \\\n  --db-subnet-group-description \"Flyte v2 private DB subnets\" --subnet-ids $SUBNETS\nRDSSG=$(aws ec2 create-security-group --region $REGION --group-name $PREFIX-rds-sg \\\n  --description \"Flyte v2 RDS 5432 from cluster nodes\" --vpc-id $VPC --query GroupId --output text)\nDBPW=$(LC_ALL=C tr -dc 'A-Za-z0-9' </dev/urandom | head -c 28)\necho \"export DBPW='$DBPW' RDSSG=$RDSSG VPC=$VPC\" >> $ENVF   # persist (DBPW is unrecoverable)\naws rds create-db-instance --region $REGION --db-instance-identifier $PREFIX-db \\\n  --engine postgres --db-instance-class db.t3.micro --allocated-storage 20 --storage-type gp3 \\\n  --master-username flyte --master-user-password \"$DBPW\" --db-name flyte \\\n  --vpc-security-group-ids $RDSSG --db-subnet-group-name $PREFIX-db-subnets \\\n  --no-publicly-accessible --backup-retention-period 1\n# Endpoint (when status=available):\nRDS_HOST=$(aws rds describe-db-instances --region $REGION --db-instance-identifier $PREFIX-db \\\n  --query 'DBInstances[0].Endpoint.Address' --output text)\necho \"export RDS_HOST=$RDS_HOST\" >> $ENVF\n```\n\n**Open 5432 from the nodes — do this once the nodegroup is up** (`kubectl get nodes` Ready),\nnot before: pod egress uses the **EKS-managed cluster SG on the nodes** (`eks-cluster-sg-*`),\nNOT `ClusterSharedNodeSecurityGroup` (gotcha 2). If you started RDS in parallel with Step 1,\nthe nodes may not exist yet — that's why this is its own step. The DB just needs this one rule:\n```bash\nNODESG=$(aws ec2 describe-instances --region $REGION \\\n  --filters \"Name=tag:eks:cluster-name,Values=$CLUSTER\" \"Name=instance-state-name,Values=running\" \\\n  --query 'Reservations[0].Instances[0].SecurityGroups[?contains(GroupName,`eks-cluster-sg`)].GroupId' --output text)\n[ -n \"$NODESG\" ] || { echo \"no running nodes yet — wait for the nodegroup, then re-run\"; }\naws ec2 authorize-security-group-ingress --region $REGION --group-id $RDSSG \\\n  --protocol tcp --port 5432 --source-group $NODESG   # init container retries until this lands\n```\n\n## Step 4 — AWS Load Balancer Controller (for ALB ingress)\n\n```bash\n# Use the policy matching the controller version the chart installs — currently v3.x.\ncurl -sL https://raw.githubusercontent.com/kubernetes-sigs/aws-load-balancer-controller/v3.4.0/docs/install/iam_policy.json -o alb-iam-policy.json\nALB_POLICY_ARN=$(aws iam create-policy --policy-name AWSLoadBalancerControllerIAMPolicy \\\n  --policy-document file://alb-iam-policy.json --query Policy.Arn --output text)\neksctl create iamserviceaccount --cluster $CLUSTER --region $REGION \\\n  --namespace kube-system --name aws-load-balancer-controller \\\n  --role-name $PREFIX-alb-controller --attach-policy-arn \"$ALB_POLICY_ARN\" --approve\nhelm repo add eks https://aws.github.io/eks-charts && helm repo update eks\nhelm upgrade --install aws-load-balancer-controller eks/aws-load-balancer-controller -n kube-system \\\n  --set clusterName=$CLUSTER --set serviceAccount.create=false \\\n  --set serviceAccount.name=aws-load-balancer-controller --set region=$REGION --set vpcId=$VPC\nkubectl -n kube-system rollout status deploy/aws-load-balancer-controller\n```\n\nIf the controller image is newer than the policy you fetched, you'll see `AccessDenied`\non actions like `elasticloadbalancing:DescribeListenerAttributes`. Fix WITHOUT reinstalling:\n```bash\ncurl -sL .../aws-load-balancer-controller/v<INSTALLED>/docs/install/iam_policy.json -o p.json\naws iam create-policy-version --policy-arn $ALB_POLICY_ARN --policy-document file://p.json --set-as-default\n```\n(Check version: `kubectl -n kube-system get deploy aws-load-balancer-controller -o jsonpath='{..image}'`.)\n\n## Step 5 — helm install flyte-binary\n\n`values-eks.yaml` (ALB HTTP-only variant). The UPPERCASE tokens (`BUCKET`, `RDS_HOST`, `DBPW`,\n`IRSA_ARN`) and `region:` are placeholders — substitute your real values before installing, e.g.\n`sed -i \"s/BUCKET/$BUCKET/g; s/RDS_HOST/$RDS_HOST/; s/DBPW/$DBPW/; s#IRSA_ARN#$IRSA_ARN#; s/us-west-2/$REGION/g\" values-eks.yaml`\n(or hand-edit). Note this chart uses `metadataContainer` (no `userDataContainer`) and its run\noutput prefix defaults to a nonexistent `s3://flyte-data` — override `storagePrefix` to your bucket:\n\n```yaml\nfullnameOverride: flyte\nflyte-core-components:\n  runs: { storagePrefix: \"s3://BUCKET\" }   # under `runs`, NOT `runs.server` (else ignored)\n# no image override needed: the repo chart pins its image tag to the chart version\n# (only the git-main chart floats `:latest` — see Image selection below)\nconfiguration:\n  database:\n    postgres:\n      host: RDS_HOST\n      port: 5432\n      dbname: flyte\n      username: flyte\n      password: \"DBPW\"\n      options: \"sslmode=require\"\n  storage:\n    metadataContainer: BUCKET\n    provider: s3\n    providerConfig: { s3: { region: us-west-2, authType: iam } }   # set to your $REGION\n  inline: { executor: { defaultK8sServiceAccount: flyte } }   # task pods inherit S3 via IRSA\nserviceAccount:\n  create: true\n  name: flyte\n  annotations: { eks.amazonaws.com/role-arn: IRSA_ARN }\ningress:\n  create: true\n  host: \"\"                       # empty => rule matches any host => reach by ALB DNS name\n  ingressClassName: alb\n  httpAnnotations:\n    alb.ingress.kubernetes.io/scheme: internet-facing\n    alb.ingress.kubernetes.io/target-type: ip\n    alb.ingress.kubernetes.io/listen-ports: '[{\"HTTP\": 80}]'\n    alb.ingress.kubernetes.io/healthcheck-path: /healthz      # binary serves /healthz on :8090\n    alb.ingress.kubernetes.io/healthcheck-port: \"8090\"\n```\n\nFor **TLS**: add `certificate-arn`, `listen-ports: '[{\"HTTP\":80},{\"HTTPS\":443}]'`,\n`ssl-redirect: \"443\"`, and set `ingress.host` to the cert hostname + a Route53 record.\nSee the TLS section below — works even when DNS lives in a different AWS account.\n\n```bash\nhelm install flyte flyteorg/flyte-binary -n flyte --create-namespace -f values-eks.yaml --dry-run  # check\nhelm install flyte flyteorg/flyte-binary -n flyte --create-namespace -f values-eks.yaml\nkubectl -n flyte get pods   # flyte stuck Init:0/1 => wait-for-db can't reach RDS (see gotchas)\n```\n\n**ALWAYS confirm the TaskAction CRD is present after install** — the chart ships it as a\nplain template, so in a shared cluster it's easily deleted out-of-band, and the binary then\nloops `Failed to watch ... taskactions.flyte.org` and every run sticks at \"queued\" (gotcha 8).\nMake it idempotent at the end of every deploy:\n```bash\nkubectl --context <ctx> get crd taskactions.flyte.org >/dev/null 2>&1 \\\n  || kubectl --context <ctx> apply -f ./flyte-binary/templates/crds/flyte.org_taskactions.yaml   # from `helm pull --untar`\nkubectl --context <ctx> get crd taskactions.flyte.org -o jsonpath='{.status.conditions[?(@.type==\"Established\")].status}'  # => True\n# Pre-existing CRD blocks helm adopt? patch ownership, then (re)install:\n#   kubectl annotate crd taskactions.flyte.org meta.helm.sh/release-name=flyte meta.helm.sh/release-namespace=flyte --overwrite\n#   kubectl label    crd taskactions.flyte.org app.kubernetes.io/managed-by=Helm --overwrite\n```\nOnly `rollout restart` if you applied the CRD onto an **already-running** binary that was\nmissing it (the watch won't retry a resource that 404'd at boot). On a normal install the chart\ncreates the CRD before the pod is Ready, so the watch establishes on first boot — don't restart\nreflexively, it's a wasted second rollout (+ image re-pull on a floating tag).\n\n**Image selection.** The published repo chart (what the official docs install) **pins the\nFlyte image tag to the chart version** — e.g. chart `v2.0.27` runs\n`cr.flyte.org/flyteorg/flyte-binary-v2:v2.0.27` — plus console\n`ghcr.io/unionai-oss/flyteconsole-v2:latest`, all with the default `pullPolicy: IfNotPresent`.\nNo image override is needed: the binary and the DB migrations it runs ship as a matched pair,\nand upgrading is `helm repo update && helm upgrade` (a new chart brings its new pinned image).\nOnly the **git-main chart** (local `./charts/flyte-binary` from a clone) still defaults to a\nfloating `:latest`, which CI pushes on every merge to `main`. If you deploy that one, set:\n```yaml\ndeployment:\n  image:\n    pullPolicy: Always    # :latest is floating + kept current by CI; re-pull on (re)start\nconsole:\n  image:\n    pullPolicy: Always\n```\nWith a floating tag, `kubectl rollout restart deploy/flyte -n flyte` forces a fresh pull (the\nwait-for-db init container uses a fixed `postgres` tag — leave it `IfNotPresent`). **For a timed\ndemo, a digest pin** (`repository@sha256:…`) + `pullPolicy: IfNotPresent` is stricter still —\nthe layer already on the node is reused and the build can't shift under you. The image and the\nschema are a matched pair (the binary owns its migrations), so running an *older* image against\na DB a *newer* one migrated (rolled back, or switched from `:latest` to the pinned repo chart)\ncan hit the schema-mismatch error in gotcha 9 — a non-issue on a fresh DB.\n\n## Step 6 — Verify\n\nThe HTTP ingress is named `<fullname>-http` — `flyte-http` with `fullnameOverride: flyte` above.\n```bash\nALB=$(kubectl -n flyte get ingress flyte-http -o jsonpath='{.status.loadBalancer.ingress[0].hostname}')\ncurl -s -X POST \"http://$ALB/flyteidl2.project.ProjectService/ListProjects\" \\\n  -H 'Content-Type: application/json' -d '{}'   # => JSON listing flytesnacks (seeded on first boot)\ncurl -s -o /dev/null -w \"%{http_code}\\n\" \"http://$ALB/v2\"   # => 200 (console)\n```\nALB takes ~2-3 min after the address appears to pass health checks and serve 200.\n\n## Step 7 — Cost dashboard (SHOW THIS after a successful deploy)\n\nOnce the deploy verifies, present this breakdown so the user knows the steady-state\nspend. Figures are **us-west-2 on-demand list prices at ~730 hrs/month** for the default\nsizing in this skill (2× `m5.large`, `db.t3.micro`, 1 ALB); scale for your instance\ntypes / region. Excludes data-transfer/egress, which depends on usage.\n\nAll amounts are **USD** (avoid bare `$` so GitHub doesn't render it as math).\n\n| AWS service | Sizing | Rate | ~ USD/mo |\n|---|---|---|---:|\n| EKS control plane | 1 cluster | 0.10/hr | 73 |\n| EC2 worker nodes | 2 × m5.large | 0.096/hr each | 140 |\n| EBS (node disks) | 2 × 50 GB gp3 | 0.08/GB-mo | 8 |\n| NAT gateway ⚠️ | 1 (eksctl default) | 0.045/hr + 0.045/GB | 33 + data |\n| RDS PostgreSQL | db.t3.micro single-AZ | 0.018/hr | 13 |\n| RDS storage | 20 GB gp3 | ~0.115/GB-mo | 2 |\n| Application Load Balancer | 1 ALB | 0.0225/hr + ~1 LCU | 16 + ~6 |\n| S3 | metadata + task I/O | 0.023/GB-mo + requests | <1 |\n| ACM / IAM / OIDC / IRSA / Secrets | — | free | 0 |\n| Route53 hosted zone (if added) | per zone | 0.50/zone-mo | 0.50 |\n| **Total (default sizing)** | | | **≈ 290–310** |\n\nCall-outs to make when showing it:\n- **Nodes + control plane ≈ 70%** of the bill.\n- **The NAT gateway is easy to miss** — eksctl's default VPC creates one (~33 USD/mo + data)\n  just so private-subnet nodes reach the internet.\n- **Actual spend so far** is only ~0.40–0.60 USD/hr × hours-up, not the monthly figure.\n- Levers: spot nodes (~−65% of node cost), 1 node / smaller instances, drop the NAT gateway\n  (public-subnet nodes, dev only), or **tear down when idle** (see Teardown) to stop the meter.\n- For real numbers, offer to pull month-to-date from Cost Explorer (`aws ce get-cost-and-usage`,\n  needs `ce:GetCostAndUsage`).\n\n## TLS (ACM + ALB), including cross-account DNS\n\n**ASK the user which hostname to use** before requesting the cert or setting `ingress.host` —\ndo NOT default to `flyte.example.com`, and do NOT silently reuse a value left over from a\nprior deploy (e.g. a `HOST=` in an old `values-eks.yaml`). The hostname drives the ACM cert,\nthe OIDC redirect URI, and the DNS record, so it must be the user's choice. If you find a\nleftover value, surface it as a *suggestion* to confirm, not a default. Then set `HOST` below\nto their answer.\n\nThe **cert and ALB must be in the Flyte account + the ALB's region**; the **DNS zone can\nlive in another account** — you just add two records there. No domain transfer/delegation.\nTLS is a prerequisite for browser SSO (ALB `authenticate-oidc` only runs on HTTPS listeners;\nIdPs reject non-`https` redirect URIs for non-localhost hosts).\n\n```bash\nHOST=flyte.example.com; ZONE=<YOUR_ROUTE53_ZONE_ID>   # zone is in whichever account owns DNS\n# 1. Flyte account: request a DNS-validated cert (same region as the ALB)\nCERT=$(aws acm request-certificate --region $REGION --domain-name $HOST \\\n  --validation-method DNS --query CertificateArn --output text)\naws acm describe-certificate --region $REGION --certificate-arn $CERT \\\n  --query 'Certificate.DomainValidationOptions[0].ResourceRecord'    # -> {Name,Type,Value}\n# 2. DNS account: UPSERT that validation CNAME into the zone (change-resource-record-sets)\n# 3. Flyte account: wait until status == ISSUED (aws acm describe-certificate ...)\n# 4. Flyte account: helm upgrade with ingress.host=$HOST and these httpAnnotations:\n#      listen-ports: '[{\"HTTP\": 80}, {\"HTTPS\": 443}]'\n#      ssl-redirect: \"443\"\n#      certificate-arn: <CERT>\n# 5. DNS account: UPSERT a CNAME  $HOST -> <ALB DNS name>  (subdomain => CNAME is fine; no\n#    cross-account alias-target dance needed)\n# 6. Verify: curl https://$HOST/v2 -> 200; http -> 301; openssl s_client shows CN=$HOST\n```\n\nTwo AWS credential sets may be in play (Flyte acct + DNS acct) — keep them in separate env\nfiles and `source` the right one per command. STS/SSO session tokens expire mid-deploy;\nwhen a call returns `ExpiredTokenException`, refresh that account's creds (kubectl/helm to\nthe cluster also need the Flyte account's creds, via `aws eks get-token`).\n\n## ALB edge SSO (Okta / OIDC)\n\nGates the console at the load balancer via ALB `authenticate-oidc` — the binary is\nunchanged. **Requires HTTPS** (see TLS above). Tradeoff: the action applies to ALL paths on\nthe ingress, so the browser console works (same-origin API calls carry the ALB session\ncookie) but **CLI/SDK clients get 302'd** — add a higher-precedence `ingress.apiJwtIngress`\nthat matches `Authorization: Bearer*` and JWT-validates it at the ALB (see CLI Bearer-bypass\nbelow).\n\n> **The Flyte binary does NOT validate tokens.** Its server wires no auth interceptor — it\n> trusts whatever reaches it (`TrustForwardedIdentityHeaders`). So auth has to be enforced at\n> the edge. A Bearer-match ingress with *no* validation action just forwards the token blindly,\n> leaving the API **wide open** (any `Authorization: Bearer anything` returns 200). To actually\n> lock it down you need ALB-native JWT validation on that ingress (the `jwt-validation`\n> annotation below) — not just a Bearer-match condition.\n\n1. Add the redirect URI **`https://<host>/oauth2/idpresponse`** to the OIDC app (fixed ALB\n   callback path) — login fails without it.\n2. Create the OIDC Secret in the **ingress namespace** (keys exactly `clientID`/`clientSecret`).\n   **Have the USER run this command themselves** — do NOT ask them to paste the client secret\n   into the chat. Give them the command and ask them to run it (inline with a leading `! ` or\n   in their own terminal) so the secret goes straight into kubectl and never reaches the assistant:\n   ```bash\n   kubectl -n flyte create secret generic flyte-console-oidc \\\n     --from-literal=clientID=<oidc-client-id> --from-literal=clientSecret=<oidc-secret>\n   ```\n3. Add these `ingress.httpAnnotations` and `helm upgrade`:\n   ```yaml\n   alb.ingress.kubernetes.io/auth-type: oidc\n   alb.ingress.kubernetes.io/auth-scope: openid email profile   # profile => given/family name in x-amzn-oidc-data\n   alb.ingress.kubernetes.io/auth-on-unauthenticated-request: authenticate\n   alb.ingress.kubernetes.io/auth-session-timeout: \"604800\"\n   alb.ingress.kubernetes.io/auth-idp-oidc: '{\"issuer\":\"https://<idp>/oauth2/default\",\"authorizationEndpoint\":\".../v1/authorize\",\"tokenEndpoint\":\".../v1/token\",\"userInfoEndpoint\":\".../v1/userinfo\",\"secretName\":\"flyte-console-oidc\"}'\n   ```\n4. **Grant the controller RBAC to read that Secret** (REQUIRED — otherwise the rule silently\n   stays plain `forward` and you keep getting HTTP 200 instead of 302). The controller SA\n   `kube-system:aws-load-balancer-controller` has no secret access in app namespaces by default:\n   ```bash\n   kubectl -n flyte create role alb-oidc-secret-reader --verb=get,list,watch --resource=secrets\n   kubectl -n flyte create rolebinding alb-oidc-secret-reader --role=alb-oidc-secret-reader \\\n     --serviceaccount=kube-system:aws-load-balancer-controller\n   ```\n   Symptom if missing: controller logs `secrets \"flyte-console-oidc\" is forbidden`.\n   - **Okta issuer host:** use the **non-admin** org domain (`https://<org>.okta.com/oauth2/default`),\n     NOT the `-admin` console host — tokens carry the non-admin host as `iss`, so jwt-validation\n     fails if you use `-admin`. Confirm via `…/oauth2/default/.well-known/openid-configuration`.\n     Switching IdPs is pure config (Secret + the issuer refs + `flyteClient.clientId`); no ALB/DNS churn.\n5. Verify: `curl -s -o /dev/null -w '%{http_code} %{redirect_url}' https://<host>/v2`\n   → `302 https://<idp>/oauth2/default/v1/authorize?client_id=...&redirect_uri=https://<host>/oauth2/idpresponse`.\n6. **IdP-side (can't be fixed from the cluster):** the user must be **assigned to the app**,\n   and (on Okta) the `default` auth server's **Access Policy** must permit the app + the\n   requested scopes (`openid email profile`). Okta error *\"Bad Request — Policy evaluation\n   failed\"* after the redirect = the access-policy rule is missing the app or restricts scopes\n   → Security → API → Authorization Servers → default → Access Policies, allow the app with\n   \"Any scopes\" (or add `email`/`profile`).\n\n### CLI Bearer-bypass (dual-auth: keep CLI/SDK working alongside console SSO)\n\nEdge SSO alone 302s CLI clients. To let token clients through, add two more ingresses in\nthe SAME ALB group plus `authMetadata` so the CLI knows to fetch a token. Three ingresses,\nordered by `group.order` (lower = higher precedence), all sharing one `group.name`:\n\n| Ingress | order | matches | auth |\n|---|---|---|---|\n| `wellknownIngress` | -150 | `/.well-known/*`, `AuthMetadataService` | none (discovery before token) |\n| `apiJwtIngress` | -140 | `Authorization: Bearer*` (via `conditions.<fullname>-http`) | ALB `jwt-validation` vs IdP JWKS |\n| http (main) | -100 | everything else | `authenticate-oidc` (cookie) |\n\n1. Tell the binary to advertise the IdP + the PKCE CLI client — under **`runs`, NOT `runs.server`**:\n   ```yaml\n   flyte-core-components:\n     runs:\n       storagePrefix: \"s3://<bucket>\"        # also belongs under runs, not runs.server\n       authMetadata:\n         externalAuthServerBaseUrl: \"https://<idp>/oauth2/default\"\n         flyteClient: { clientId: <native-PKCE-app>, redirectUri: http://localhost:53593/callback, scopes: [openid, profile, offline_access] }\n   ```\n2. Add `group.name: <group>` + `group.order: \"-100\"` to the main `httpAnnotations`. Then add\n   `ingress.apiJwtIngress` (enabled, order -140) and `ingress.wellknownIngress` (enabled, order\n   -150, no auth). All three need cert-arn + listen-ports + ssl-redirect. The apiJwtIngress\n   carries **two** annotations that together do the lock-down — the Bearer-match condition AND\n   ALB-native JWT validation (the `conditions.*` key targets the rendered backend service name,\n   `<fullname>-http`, e.g. `flyte-http`):\n   ```yaml\n   ingress:\n     apiJwtIngress:\n       enabled: true\n       annotations:\n         alb.ingress.kubernetes.io/group.name: <group>\n         alb.ingress.kubernetes.io/group.order: \"-140\"\n         alb.ingress.kubernetes.io/conditions.<fullname>-http: '[{\"field\":\"http-header\",\"httpHeaderConfig\":{\"httpHeaderName\":\"Authorization\",\"values\":[\"Bearer*\"]}}]'\n         # ALB checks signature + iss/exp against the IdP JWKS; bad/expired token -> 401 at the edge\n         alb.ingress.kubernetes.io/jwt-validation: '{\"jwksEndpoint\":\"https://<idp>/oauth2/default/v1/keys\",\"issuer\":\"https://<idp>/oauth2/default\"}'\n         # plus the shared cert-arn / listen-ports / ssl-redirect / target-type / healthcheck-* annotations\n   ```\n   **`jwt-validation` needs a recent controller** — ALB native JWT verification shipped Nov 2025,\n   exposed by aws-load-balancer-controller as `alb.ingress.kubernetes.io/jwt-validation`. On an\n   older controller the annotation is silently ignored and the API stays open — confirm the\n   controller is new enough (check its release date, not just a high-looking version number) and\n   verify the 401 below. The `conditions.*` match alone (no `jwt-validation`) does NOT validate\n   anything; it only routes Bearer requests past the cookie flow.\n3. `helm upgrade`. **Adding group.name recreates the ALB under a new name** (`k8s-<group>-*`)\n   and deletes the old standalone one — **re-point the DNS CNAME to the new ALB DNS name.**\n4. Verify (use `curl --connect-to host:443:<alb>:443` before DNS propagates):\n   - `POST .../AuthMetadataService/GetOAuth2Metadata` (with `Content-Type: application/json`) → 200\n   - API `+ Authorization: Bearer fake` → **401, not 302** (matched the JWT ingress, validated)\n   - `/v2` and API without Bearer → 302 (cookie path)\n\nGotchas: (a) `authMetadata`/`storagePrefix` go under `runs`, not `runs.server` — misplaced,\nthey're silently ignored and `GetOAuth2Metadata` returns `unimplemented`. (b) config-only helm\nchanges may not roll the pod — `kubectl rollout restart deploy/flyte` to be sure. (c) the\n`conditions.*` key must match the rendered backend service name (`<fullname>-http`).\n\n## App serving (optional — Knative + Kourier)\n\nFlyte v2 can host long-running **apps** (deployed via the SDK), each published at\n`{name}-{project}-{domain}.<base-domain>`. It's **off by default**: the binary always exposes\n`AppService`, but with no controller behind it the console's Apps tab and any\n`flyteidl2.app.AppService/List` call return **`{\"code\":\"unimplemented\",\"message\":\"404 Not Found\"}`**\nuntil you enable it. Apps run as **Knative Services**, so this needs Knative Serving + a Knative\nnetworking layer (Kourier) installed first — your cloud's ALB controller can't be Knative's\nnetworking layer. Skip this section unless the user wants apps. Official doc:\nhttps://www.union.ai/docs/v2/flyte/oss-deployment/app-serving/.\n\n**1. Install Knative Serving + Kourier.** Pick a Knative release that supports your cluster's\nk8s version — Knative only supports the most recent k8s minors, so the upstream doc's pinned\nversion is often too old (e.g. on k8s 1.34, v1.17 is rejected; v1.22 works). Check the latest\nthat matches, and use the **same version** for serving and net-kourier:\n```bash\nKV=knative-v1.22.1   # must support your k8s version; serving + net-kourier must match\nkubectl --context <ctx> apply -f https://github.com/knative/serving/releases/download/$KV/serving-crds.yaml\nkubectl --context <ctx> apply -f https://github.com/knative/serving/releases/download/$KV/serving-core.yaml\nkubectl --context <ctx> apply -f https://github.com/knative-extensions/net-kourier/releases/download/$KV/kourier.yaml\nkubectl --context <ctx> patch configmap/config-network -n knative-serving --type merge \\\n  -p '{\"data\":{\"ingress-class\":\"kourier.ingress.networking.knative.dev\"}}'\nkubectl --context <ctx> wait --for=condition=Available deploy --all -n knative-serving --timeout=180s\nkubectl --context <ctx> wait --for=condition=Available deploy --all -n kourier-system --timeout=180s\n```\n\n**2. Configure the apps domain — single-label so one wildcard cert covers every app.** Set\n`config-domain` to your base domain and drop the namespace from the hostname template (Knative's\ndefault `{{.Name}}.{{.Namespace}}.{{.Domain}}` is two labels, which `*.<base-domain>` can't match):\n```bash\nkubectl --context <ctx> patch configmap/config-domain -n knative-serving --type merge \\\n  -p '{\"data\":{\"<base-domain>\":\"\"}}'                       # e.g. flyte-v2.example.com\nkubectl --context <ctx> patch configmap/config-network -n knative-serving --type merge \\\n  -p '{\"data\":{\"domain-template\":\"{{.Name}}.{{.Domain}}\"}}'\n```\n\n**3. Expose Kourier behind the existing ALB** (reuse the same `group.name` so apps share the\nFlyte load balancer — no second LB, no extra cost). Switch the Kourier Service to `ClusterIP`\nand add an Ingress in `kourier-system` joined to that group, with a **wildcard cert** and the\n`*.<base-domain>` host. Match the group-level annotations (scheme / target-type / listen-ports /\nssl-redirect) to the Flyte ingresses or the controller errors on conflicting group config:\n```bash\nkubectl --context <ctx> patch svc kourier -n kourier-system --type merge -p '{\"spec\":{\"type\":\"ClusterIP\"}}'\n```\n```yaml\n# kourier-alb-ingress.yaml\napiVersion: networking.k8s.io/v1\nkind: Ingress\nmetadata:\n  name: kourier-alb\n  namespace: kourier-system\n  annotations:\n    alb.ingress.kubernetes.io/group.name: flyte          # SAME group as the Flyte ingresses\n    alb.ingress.kubernetes.io/group.order: \"-50\"\n    alb.ingress.kubernetes.io/scheme: internet-facing\n    alb.ingress.kubernetes.io/target-type: ip\n    alb.ingress.kubernetes.io/listen-ports: '[{\"HTTP\": 80}, {\"HTTPS\": 443}]'\n    alb.ingress.kubernetes.io/ssl-redirect: \"443\"\n    alb.ingress.kubernetes.io/certificate-arn: <WILDCARD_CERT_ARN>\n    alb.ingress.kubernetes.io/healthcheck-path: /\n    alb.ingress.kubernetes.io/success-codes: \"200,404\"   # Kourier 404s an unmatched host while healthy\nspec:\n  ingressClassName: alb\n  rules:\n    - host: \"*.<base-domain>\"\n      http:\n        paths:\n          - { path: /, pathType: Prefix, backend: { service: { name: kourier, port: { number: 80 } } } }\n```\n- **Wildcard cert:** request `*.<base-domain>` in ACM. ACM derives the DNS-validation CNAME from\n  the base name, so if you already validated the console's `<base-domain>` cert, the wildcard\n  reuses the **same** validation record and auto-issues — often no new DNS needed.\n- **Wildcard DNS:** add `*.<base-domain>` → the ALB DNS name (cross-account zones: a CNAME in the\n  DNS account, same as the console record).\n- (Alternative: leave the Kourier Service as `LoadBalancer` and point `*.<base-domain>` at that NLB\n  — simpler, but a second load balancer + you manage its TLS separately.)\n\n**3b. Require authentication for apps (optional but recommended).** By default any app URL is\n**public** — anyone who can reach the ALB opens it. To gate every app behind the same OIDC login\nas the console, add `authenticate-oidc` to the **`kourier-alb`** Ingress (same mechanism as the\nconsole's edge SSO). Three parts, all required:\n- **Annotations** on the `kourier-alb` Ingress (alongside the step-3 annotations):\n  ```yaml\n  alb.ingress.kubernetes.io/auth-type: oidc\n  alb.ingress.kubernetes.io/auth-on-unauthenticated-request: authenticate\n  alb.ingress.kubernetes.io/auth-scope: openid email profile\n  alb.ingress.kubernetes.io/auth-session-timeout: \"604800\"\n  alb.ingress.kubernetes.io/auth-idp-oidc: '{\"issuer\":\"https://<idp>/oauth2/default\",\"authorizationEndpoint\":\"https://<idp>/oauth2/default/v1/authorize\",\"tokenEndpoint\":\"https://<idp>/oauth2/default/v1/token\",\"userInfoEndpoint\":\"https://<idp>/oauth2/default/v1/userinfo\",\"secretName\":\"flyte-console-oidc\"}'\n  ```\n- **The OIDC Secret must exist in `kourier-system`** (the controller reads it from the Ingress's\n  own namespace). Reuse the console's client by copying the existing Secret over:\n  ```bash\n  kubectl --context <ctx> get secret flyte-console-oidc -n flyte -o json \\\n    | jq '.metadata={name:\"flyte-console-oidc\",namespace:\"kourier-system\"}' \\\n    | kubectl --context <ctx> apply -f -\n  ```\n- **RBAC for the controller to read it in `kourier-system`** — its Secret access is per-namespace,\n  so without a Role here the Ingress fails with `secrets \"…\" is forbidden`, which **stalls\n  reconciliation of the whole ALB group** (not just this Ingress — it can take the console down):\n  ```bash\n  kubectl --context <ctx> -n kourier-system create role alb-oidc-secret-reader \\\n    --verb=get,list,watch --resource=secrets\n  kubectl --context <ctx> -n kourier-system create rolebinding alb-oidc-secret-reader \\\n    --role=alb-oidc-secret-reader --serviceaccount=kube-system:aws-load-balancer-controller\n  ```\n- **IdP redirect URI:** ALB's callback is `https://<app-host>/oauth2/idpresponse` and every app\n  has a different hostname, so register the **wildcard** `https://*.<base-domain>/oauth2/idpresponse`\n  as a sign-in redirect URI on the OIDC app (your IdP must allow wildcard redirect URIs). Without\n  it, login dead-ends after the redirect with a `redirect_uri` error.\n\nVerify: `curl -s -o /dev/null -w '%{http_code}'` an app host → **302** to the IdP (was 200/404).\nNote this gates apps with the **cookie** flow (browser) — app-to-app/API calls would need the same\nBearer-bypass treatment as the main API if you want programmatic access.\n\n**4. Enable the controller in Flyte values** (under `configuration.inline`), then upgrade.\n`baseDomain` MUST equal the `config-domain` from step 2 so advertised URLs match what Knative serves:\n```yaml\nconfiguration:\n  inline:\n    internalApps:\n      enabled: true\n      baseDomain: <base-domain>\n      scheme: https\n      ingressAppsPort: 0                 # apps sit behind the ALB on 443; omit the port\n      defaultServiceAccount: flyte       # app pods run under the IRSA'd SA (S3 access)\n```\n```bash\nhelm upgrade flyte flyteorg/flyte-binary -n flyte -f values-eks.yaml --kube-context <ctx>\nkubectl --context <ctx> -n flyte rollout restart deploy/flyte   # config-only change may not roll the pod\n```\nThe chart auto-grants the `serving.knative.dev` RBAC when `internalApps.enabled` (flyte#7557).\n\n**5. Verify.**\n```bash\nkubectl --context <ctx> auth can-i create services.serving.knative.dev \\\n  --as=system:serviceaccount:flyte:flyte -n flyte                      # => yes\n# AppService in-cluster (bypasses the ALB auth gate) — 200 + {} (NOT 404/unimplemented):\nkubectl --context <ctx> -n flyte run c --rm -i --image=curlimages/curl:8.10.1 --restart=Never -- \\\n  curl -s -o /dev/null -w '%{http_code}\\n' -X POST \\\n  http://flyte-http.flyte:8090/flyteidl2.app.AppService/List -H 'Content-Type: application/json' -d '{}'\n# App path via ALB (before DNS, use --connect-to): wildcard TLS served + 404 from Kourier = healthy, no app yet:\ncurl -s -o /dev/null -w '%{http_code}\\n' --connect-to \"noapp.<base-domain>:443:<alb>:443\" https://noapp.<base-domain>/\n```\nThe console's Apps tab now loads; deploy an app with the SDK and open\n`https://<name>-<project>-<domain>.<base-domain>`.\n\n**Gotchas:** (a) Knative version too new for your k8s → `kubectl apply` rejects the manifests;\ninstall an older Knative (serving + net-kourier matched). (b) Two-label app hostname → wildcard\nTLS error; confirm the single-label `domain-template`. (c) `baseDomain` ≠ `config-domain` → URLs\nFlyte advertises don't match what Knative serves. (d) **Apps are unauthenticated at the edge** —\nthe Kourier ingress carries no OIDC/JWT, so app URLs are public once DNS resolves (the console\nApps *tab* is still SSO-gated) — gate them with step 3b. (e) `List` still 404s after enabling →\nthe binary didn't roll; `rollout restart`. (f) Enabling app auth without the `kourier-system`\nSecret RBAC (step 3b) → `secrets \"…\" is forbidden` stalls the **whole** ALB group, which can take\nthe console down too — add the Role/RoleBinding before (or with) the auth annotations.\n\n## Optional `configuration.inline` tuning\n\nAnything under `configuration.inline` is merged into the rendered Flyte config — it's how you\nset options the top-level values don't expose. All of the below go in `values-eks.yaml`; apply\nwith `helm upgrade flyte ... -f values-eks.yaml` (config-only changes may not roll the pod —\n`kubectl rollout restart deploy/flyte -n flyte` if it doesn't pick them up).\n\n**Default task resources.** CPU/memory requests for task pods that don't set their own:\n```yaml\nconfiguration:\n  inline:\n    plugins:\n      k8s:\n        default-cpus: 500m\n        default-memory: 1Gi\n```\n\n**Default task scheduling.** Tolerations / affinity / node selectors / injected env on every\ntask pod (same `plugins.k8s` block — `default-env-vars` is also where gotcha 7's callback vars\ngo on older charts):\n```yaml\nconfiguration:\n  inline:\n    plugins:\n      k8s:\n        default-tolerations:\n          - { key: flyte.org/node-role, operator: Equal, value: worker, effect: NoSchedule }\n        default-affinity: {}             # a standard core/v1 Affinity\n        default-env-vars:\n          - MY_ENV_VAR: value            # injected into every task pod\n```\n\n**OpenTelemetry.** Off by default (`otel.type: noop`). Point it at an OTLP collector — prefer\n`otlpgrpc` (the `otlphttp` metric exporter reuses the trace endpoint path):\n```yaml\nconfiguration:\n  inline:\n    otel:\n      type: otlpgrpc                     # noop | file | jaeger | otlpgrpc | otlphttp\n      otlpgrpc: { endpoint: http://otel-collector.flyte.svc.cluster.local:4317 }\n      sampler: { parentSampler: traceid, traceIdRatio: 0.01 }   # keep 1% of traces in prod\n```\n\n**DB password (and S3 keys) from a Secret.** Setting `configuration.database.postgres.password`\nalready writes it into a mounted k8s Secret (not the plaintext ConfigMap); same for S3 access\nkeys when `authType: accesskey`. To keep the password out of the values file entirely, leave\n`password` empty and either reference an existing Secret with\n`configuration.extraInlineSecretRefs`, or mount it as a file and point\n`configuration.database.postgres.passwordPath` at it (`password` and `passwordPath` are mutually\nexclusive). This is the better choice than the plaintext `password:` shown in Step 5 when the\nvalues file is committed or shared.\n\n## Stable ALB across redeploys (anchor ingress)\n\nBy default the ALB is owned by Flyte's ingresses, so `helm uninstall` deletes it and the next\ninstall mints a **new ALB with a new DNS name** — forcing a DNS re-point every cycle. To keep\none stable endpoint, exploit the controller's rule that it keeps exactly **one ALB per\n`group.name` as long as ≥1 ingress in that group exists**: add a permanent \"anchor\" ingress in\nthe group, applied **out-of-band (NOT in the helm release)**, so the ALB survives uninstall.\n\n```yaml\napiVersion: networking.k8s.io/v1\nkind: Ingress\nmetadata:\n  name: flyte-alb-anchor\n  namespace: flyte\n  annotations:\n    alb.ingress.kubernetes.io/group.name: flyte            # SAME group as the flyte ingresses\n    alb.ingress.kubernetes.io/group.order: \"100\"           # lowest precedence; never shadows flyte rules\n    alb.ingress.kubernetes.io/scheme: internet-facing      # group-level annotations MUST match the\n    alb.ingress.kubernetes.io/target-type: ip              # flyte ingress group, else the controller\n    alb.ingress.kubernetes.io/listen-ports: '[{\"HTTP\": 80}, {\"HTTPS\": 443}]'   # errors on conflicting config\n    alb.ingress.kubernetes.io/certificate-arn: <CERT_ARN>\n    # fixed-response backend => the anchor needs NO real service, so it stands alone when flyte is uninstalled:\n    alb.ingress.kubernetes.io/actions.anchor-ok: '{\"type\":\"fixed-response\",\"fixedResponseConfig\":{\"contentType\":\"text/plain\",\"statusCode\":\"200\",\"messageBody\":\"flyte-alb-anchor\"}}'\nspec:\n  ingressClassName: alb\n  rules:\n    - http:\n        paths:\n          - { path: /__alb_anchor, pathType: ImplementationSpecific, backend: { service: { name: anchor-ok, port: { name: use-annotation } } } }\n```\n\n```bash\nkubectl --context <ctx> apply -f alb-anchor.yaml          # provisions the ALB once\n# point DNS at THIS ALB's name one time; it never changes again:\nkubectl --context <ctx> -n flyte get ingress flyte-alb-anchor -o jsonpath='{.status.loadBalancer.ingress[0].hostname}'\n```\n\nNow `helm uninstall flyte` removes Flyte's rules but the anchor keeps the ALB (same DNS name)\nalive; `helm install` re-adds Flyte's listener rules (including the `authenticate-oidc` SSO\nrule) onto the surviving ALB. Verify the anchor still answers between deploys:\n`curl http://<alb>/__alb_anchor` → `200 flyte-alb-anchor`. (To intentionally delete the ALB,\nremove the anchor too.) Keeps the annotation-driven SSO config — the alternative, a fully\npre-provisioned BYO ALB via the controller's `TargetGroupBinding` CRD + `ingress.create:\nfalse`, is more stable still but makes you hand-manage every listener/SSO rule yourself.\n\n## Pruning run data from the DB\n\nTo wipe run history without reinstalling, prune the DB directly. The v2 run data lives in\njust two tables (Postgres `flyte` DB): **`actions`** (one row per run/action) and\n**`action_events`** (per-attempt events). They're linked by `(project, domain, run_name,\nname)` — there's no FK, so delete events first, then actions. **Keep** `projects` (seeded\n`flytesnacks`), `schema_migrations` (migration state), and `task_specs` (registered tasks).\n\nRDS is private (no public access), so run psql from an **ephemeral in-cluster pod** rather\nthan your laptop. Prune only **finished** runs by keying on `ended_at IS NOT NULL` — that\nskips anything still in-flight (an open run has `ended_at` null):\n\n```bash\nDBHOST=<rds-endpoint>; DBPW=<db-password>   # from your values-eks.yaml\nkubectl --context <ctx> run pgcli --rm -i --restart=Never -n flyte \\\n  --image=postgres:16 --env PGPASSWORD=\"$DBPW\" --command -- \\\n  psql \"host=$DBHOST user=flyte dbname=flyte sslmode=require\" -P pager=off -v ON_ERROR_STOP=1 \\\n  -c \"begin;\n      delete from action_events ae using actions a\n        where ae.project=a.project and ae.domain=a.domain\n          and ae.run_name=a.run_name and ae.name=a.name and a.ended_at is not null;\n      delete from actions where ended_at is not null;\n      commit;\"\n```\n\nInspect first with `select relname,n_live_tup from pg_stat_user_tables order by 2 desc;`.\nNote a run **stuck \"queued\"** (e.g. from the missing-CRD gotcha 8) has `ended_at` null, so\nthis leaves it untouched — delete those explicitly by `run_name` once you've confirmed no\ntask pod / TaskAction CR backs them. To wipe **everything** instead, `truncate actions,\naction_events;` (projects/migrations survive). To reset the whole DB, see Teardown +\nreinstall, or `drop database flyte; create database flyte;` and rollout-restart the binary.\n\n## Teardown\n\n```bash\nhelm uninstall flyte -n flyte          # deletes the ingress => controller removes the ALB\n# helm uninstall leaves the run/task pods behind (the controller created them, not Helm) —\n# delete them explicitly so the namespace is clean for a redeploy:\nkubectl --context <ctx> -n flyte delete pods --all\nhelm uninstall aws-load-balancer-controller -n kube-system\naws rds delete-db-instance --region $REGION --db-instance-identifier $PREFIX-db --skip-final-snapshot --delete-automated-backups\naws rds delete-db-subnet-group --region $REGION --db-subnet-group-name $PREFIX-db-subnets\naws s3 rb s3://$PREFIX-data-$ACCT --force\neksctl delete cluster -f cluster.yaml   # tears down VPC, nodegroup, OIDC, IRSA stacks\n# Delete the standalone IAM policies (detach first if needed):\naws iam delete-policy --policy-arn arn:aws:iam::$ACCT:policy/$PREFIX-s3-access\naws iam delete-policy --policy-arn arn:aws:iam::$ACCT:policy/AWSLoadBalancerControllerIAMPolicy\n```\n\n## Gotchas (each one bit during a real run)\n\n1. **eksctl too old → \"unsupported Kubernetes version\".** eksctl 0.175 only offers up to\n   1.29, but EKS has dropped 1.29 from standard support → CFN `ControlPlane` fails ~30s in\n   and rolls back. Use eksctl ≥ 0.227 (defaults to a current version); pin a supported one\n   (1.33 worked). After a failed create, delete the `ROLLBACK_COMPLETE` stack before retrying.\n2. **RDS unreachable: wrong source SG.** The pod stays `Init:0/1` (`wait-for-db ... no\n   response`). EKS managed-nodegroup nodes run with the **EKS-managed cluster SG**\n   (`eks-cluster-sg-<cluster>-*`), NOT `ClusterSharedNodeSecurityGroup`. Pod egress (VPC CNI\n   secondary IPs on the primary ENI) uses the node-ENI SG. Authorize 5432 on the RDS SG from\n   the actual node SG (`describe-instances ... SecurityGroups`), not the shared one. Init\n   container retries on its own once the rule lands.\n3. **ALB controller IAM lag.** The eks chart installs the latest controller (v3.x), which\n   needs newer IAM actions (e.g. `DescribeListenerAttributes`) than older policy JSON.\n   Match `iam_policy.json` to the installed controller version (create-policy-version\n   --set-as-default; no reinstall needed).\n4. **Default storagePrefix is fake.** `flyte-core-components.runs.storagePrefix` (under `runs`,\n   NOT `runs.server`) defaults to `s3://flyte-data` — override to your real bucket or run I/O\n   fails. Misplaced under `runs.server` it's silently ignored.\n5. **ALB by DNS name:** leave `ingress.host: \"\"` so the rule matches any host; the binary\n   serves `/healthz` on `:8090` for the ALB health check. Add ACM cert + Route53 for TLS.\n6. Postgres default major from RDS is fine (chart needs ≥12).\n7. **Task pods loop/recreate every ~75s — missing control-plane callback env vars.** A run's\n   task pod (image `ghcr.io/flyteorg/flyte:py3.x-vX`) calls back to the backend to enqueue\n   child actions / watch state. Without config it uses the **devbox default\n   `host.docker.internal:8090`** → `dns error: Name or service not known` → retries exhaust →\n   controller recreates the pod, forever. Recent `flyte-binary` chart versions inject these by\n   default; on older charts add them via `configuration.inline.plugins.k8s.default-env-vars`:\n   ```yaml\n   configuration:\n     inline:\n       plugins:\n         k8s:\n           default-env-vars:\n             - _U_EP_OVERRIDE: flyte-http.flyte:8090   # in-cluster HTTP svc = <fullname>-http.<ns>:8090\n             - _U_INSECURE: \"true\"                     # svc is plain HTTP on :8090; without this the\n                                                       # SDK uses https:// → \"received corrupt message\n                                                       # of type InvalidContentType\"\n             - _U_USE_ACTIONS: \"1\"                     # enable the QueueService/actions path\n   ```\n   Verify a task pod: `kubectl -n flyte get pod <run>-a0-0 -o jsonpath='{..env[*].name}'` shows\n   `_U_EP_OVERRIDE`, and its logs no longer mention `host.docker.internal` or `InvalidContentType`.\n8. **Runs stuck \"queued\" — missing TaskAction CRD.** The chart ships `taskactions.flyte.org`\n   under `templates/crds/` (NOT Helm's delete-protected `crds/` dir), so it's an ordinary,\n   release-owned template. Two consequences bite:\n   - `helm uninstall` DELETES it (and all TaskAction CRs); a later `helm install` doesn't\n     reliably re-establish it, and Helm won't recreate it while the release sits at\n     `deployed` even though `helm get manifest` still lists it.\n   - It's **cluster-scoped but owned by a namespaced release** — so in a **shared cluster**,\n     uninstalling *any* Flyte release (or a stray `kubectl delete crd`) wipes it for\n     everyone, and it can vanish *after* a successful install with the binary still running.\n   Symptom, two variants by binary version: older builds keep *running* and log `Failed to\n   watch ... could not find the requested resource (get taskactions.flyte.org)` every few\n   seconds (runs sit at \"queued\"); the **current `flyte-binary-v2` hard-fails at startup** —\n   `Error: setup failed: actions: failed to start TaskAction watcher: ... no matches for kind\n   \"TaskAction\" in version \"flyte.org/v1\"`, exit 1 → **CrashLoopBackOff** (the console still\n   serves and the ALB still 302s, so check the *binary* pod, not just the URL). Same root cause,\n   same fix:\n   ```bash\n   kubectl --context <ctx> apply -f ./flyte-binary/templates/crds/flyte.org_taskactions.yaml   # from `helm pull --untar`\n   kubectl --context <ctx> -n flyte rollout restart deploy/flyte   # re-establish the watch\n   ```\n   A normal `helm uninstall` → `helm install` cycle mostly self-heals: uninstall deletes the CRD,\n   install recreates it as a template. **But CRD registration and the pod start race.** If the\n   CRD wins, the binary boots clean on the first try (verified once: rollout Ready ~41s, single\n   rollout). If the pod wins, the current binary **crashloops** until the CRD is discoverable,\n   then comes up after a restart or two (~30–60s; `--wait` rides it out) — also verified in the\n   same cluster minutes later, so treat the crashloop as expected, not a failure. Either way\n   **verify the binary pod is `1/1 Running` and the CRD is `Established` after every deploy**\n   (see Step 5) — a green `helm install` and a `302` from the URL do NOT prove the binary is up\n   (the ALB 302s and the console serves even while the binary crashloops).\n   **Truly protecting it from `helm uninstall` requires removing it from the release manifest**\n   — stripping the live CRD's Helm ownership *labels* does NOT work, because uninstall deletes\n   by manifest membership, not by label (observed: labels stripped, uninstall still deleted it).\n   The ownership strip only avoids the *install-time* adopt conflict (Step 5). To survive\n   uninstall, manage the CRD entirely out-of-band: delete it from the chart's `templates/crds/`\n   (so no release ever lists it) and `kubectl apply` it yourself once. Otherwise just rely on the\n   self-healing install above — simpler, and fine for demos/redeploys.\n9. **Listing runs/tasks fails: `missing destination name <col> in *[]*models.Action`.** The\n   install is green, the console loads, but opening a project's runs (or any\n   `RunService/ListActions` / `ListRuns` call) returns\n   `{\"code\":\"internal\",\"message\":\"failed to list actions: missing destination name created_by in *[]*models.Action\"}`\n   (the column varies — `created_by`, `executed_by`, …). **Root cause: the binary image and the DB\n   schema disagree** — the `actions` table has a column the running `models.Action` struct can't\n   scan, i.e. the schema was migrated by a *different* image than the one running. A deploy on a\n   fresh DB won't hit this (any image — the pinned repo-chart tag or `:latest` — migrates its own\n   schema). It shows up when you **reuse a DB that a newer/feature-branch image migrated** and then\n   run an older image against it (e.g. rolled back, pinned a digest, or switched from a `:latest`\n   deploy to the older pinned repo-chart image). Diagnose, then make the schema match the image —\n   on a fresh/disposable DB just let the running image re-migrate from scratch:\n   ```bash\n   # see the extra columns the DB has, and check there's no run data worth keeping:\n   kubectl --context <ctx> -n flyte run pg --rm -i --restart=Never --image=postgres:16 \\\n     --env PGPASSWORD=<pw> --command -- psql \"host=<rds> user=flyte dbname=flyte sslmode=require\" -A -t \\\n     -c \"select column_name from information_schema.columns where table_name='actions' and column_name like '%_by%';\" \\\n     -c \"select count(*) from actions;\"\n   # if count is 0 (or disposable), reset the schema and let the image re-migrate on boot:\n   kubectl --context <ctx> -n flyte scale deploy/flyte --replicas=0\n   kubectl --context <ctx> -n flyte run pg --rm -i --restart=Never --image=postgres:16 \\\n     --env PGPASSWORD=<pw> --command -- psql \"host=<rds> user=flyte dbname=flyte sslmode=require\" \\\n     -c \"DROP SCHEMA public CASCADE; CREATE SCHEMA public; GRANT ALL ON SCHEMA public TO flyte; GRANT ALL ON SCHEMA public TO public;\"\n   kubectl --context <ctx> -n flyte scale deploy/flyte --replicas=1   # re-migrates clean\n   ```\n   Confirm in-cluster (bypasses the ALB JWT gate):\n   `kubectl -n flyte run c --rm -i --image=curlimages/curl --restart=Never -- curl -s -XPOST\n   http://flyte-http.flyte:8090/flyteidl2.workflow.RunService/ListActions -H 'Content-Type: application/json'\n   -d '{\"project_id\":{\"domain\":\"development\",\"name\":\"flytesnacks\"}}'` → `{}` (not the error).\n   A green `helm install` and a `302` from `/v2` do NOT prove run-listing works (the binary serves\n   the console and auth-redirects even when this query is broken), so exercise `ListActions` after a\n   deploy that pins/reuses an image or DB.\n"},"changes":[],"summary":"First saved snapshot. No earlier version is available for comparison.","summary_kind":"deterministic","summary_metadata":{}}