← AWS Data AnalyticsCONTENT HISTORY

Update to AWS Data Analytics

Snapshot Sep 30, 2026 · 22:53 UTC · version 1.0.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "connecting-to-data-source",
  "description": "Create and troubleshoot AWS Glue connections to JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS), Redshift, Snowflake, and BigQuery. Gathers connection hints from user, discovers existing connections and RDS/Redshift candidates, registers credentials in Secrets Manager or IAM DB auth, configures VPC, and tests. Triggers on: connect to database, set up Glue connection, register data source, connect to Snowflake/BigQuery/RDS, connection timeout, test connection, troubleshoot connection. Do NOT use for moving data (use ingesting-into-data-lake), creating tables (use creating-data-lake-table), queries (use querying-data-lake), catalog exploration (use exploring-data-catalog), or SaaS (Salesforce, ServiceNow, SAP, MongoDB, Kafka).",
  "included_files": [
    {
      "relative_path": "references/bigquery-setup.md",
      "size_in_bytes": 2645
    },
    {
      "relative_path": "references/credential-security.md",
      "size_in_bytes": 4400
    },
    {
      "relative_path": "references/discovery.md",
      "size_in_bytes": 3668
    },
    {
      "relative_path": "references/jdbc-setup.md",
      "size_in_bytes": 3569
    },
    {
      "relative_path": "references/network-setup.md",
      "size_in_bytes": 4010
    },
    {
      "relative_path": "references/snowflake-setup.md",
      "size_in_bytes": 2687
    },
    {
      "relative_path": "references/troubleshooting.md",
      "size_in_bytes": 7810
    }
  ],
  "skill_md_contents": "---\nname: connecting-to-data-source\ndescription: >-\n  Create and troubleshoot AWS Glue connections to JDBC databases (Oracle, SQL Server,\n  PostgreSQL, MySQL, RDS), Redshift, Snowflake, and BigQuery. Gathers connection hints\n  from user, discovers existing connections and RDS/Redshift candidates, registers\n  credentials in Secrets Manager or IAM DB auth, configures VPC, and tests. Triggers\n  on: connect to database, set up Glue connection, register data source, connect to\n  Snowflake/BigQuery/RDS, connection timeout, test connection, troubleshoot connection.\n  Do NOT use for moving data (use ingesting-into-data-lake), creating tables (use\n  creating-data-lake-table), queries (use querying-data-lake), catalog exploration\n  (use exploring-data-catalog), or SaaS (Salesforce, ServiceNow, SAP, MongoDB, Kafka).\nmetadata:\n  version: \"1\"\n  argument-hint: \"'[source-type|connection-name|hostname]'\"\n---\n\n# Connect to Data Source\n\nRegister an external data source with AWS Glue so downstream skills (ingesting-into-data-lake) can move data from it. A Glue connection stores the network config, driver, and credential reference for one source. Create once per source, reuse across jobs.\n\n## Philosophy\n\n**A connection is a named pipe, not a pipeline.** This skill produces a tested, reusable Glue connection. It does not move data.\n\n## Common Tasks\n\nYou MUST execute commands using AWS MCP server tools when connected -- they provide validation, sandboxed execution, and audit logging. Fall back to AWS CLI only if MCP is unavailable. You MUST explain each step before executing.\n\n## Workflow\n\n### 1. Verify Dependencies and Context\n\n- You MUST check whether AWS MCP tools or AWS CLI are available and inform the user if missing\n- You MUST confirm target AWS region and verify credentials with `aws sts get-caller-identity`\n\n### 2. Classify the Source\n\nAsk the user which source type they want to connect to, or infer from hints:\n\n| User says... | Source type | Connection type | Reference |\n|---|---|---|---|\n| \"Oracle\", \"SQL Server\", \"Postgres\", \"MySQL\", \"RDS \\<engine\\>\" | JDBC database | `JDBC` | [jdbc-setup.md](references/jdbc-setup.md) |\n| \"Redshift\", \"my cluster\", \"my data warehouse on AWS\" | Redshift | `JDBC` | [jdbc-setup.md](references/jdbc-setup.md) (Redshift section) |\n| \"Snowflake\" | Snowflake | `SNOWFLAKE` | [snowflake-setup.md](references/snowflake-setup.md) |\n| \"BigQuery\", \"Google analytics warehouse\" | BigQuery | `BIGQUERY` | [bigquery-setup.md](references/bigquery-setup.md) |\n\nIf the user names DynamoDB or a local file, stop and tell them: DynamoDB is read directly by Glue without a connection, and local files belong in the ingesting-into-data-lake skill's local-upload workflow.\n\n### 3. Gather Connection Hints from the User\n\nYou MUST ask for hints the user can provide -- do not guess.\n\n**For all sources:**\n\n- Desired connection name (lowercase, hyphens: `oracle-prod-sales`, `snowflake-analytics`)\n- Existing Secrets Manager secret, or create one\n- Is source reachable from a Glue VPC (same, peered, VPN, Direct Connect)\n\n**JDBC:** hostname/endpoint, port, database, whether RDS/Aurora/self-managed, IAM DB auth enabled (Aurora/RDS MySQL/Postgres), SSL required.\n\n**Snowflake:** account identifier, warehouse, role, default database, auth (password, key-pair, OAuth).\n\n**BigQuery:** GCP project ID, location, whether service account JSON is provisioned.\n\n### 4. Discover Existing Connections and Candidate Sources\n\nCheck what exists before creating.\n\n**Existing Glue connections:**\n\n```bash\naws glue get-connections --filter ConnectionType=<TYPE> --region <REGION>\n```\n\nIf a suitable one exists, confirm and skip to Step 7.\n\n**Candidate sources in account** (JDBC/Redshift only):\n\n- RDS: `aws rds describe-db-instances`\n- Aurora: `aws rds describe-db-clusters`\n- Redshift: `aws redshift describe-clusters`\n\nPresent candidates to user; let them pick. See [discovery.md](references/discovery.md).\n\n### 5. Register Credentials\n\nYou MUST encourage AWS Secrets Manager over plaintext passwords. You SHOULD prefer IAM database authentication where supported (Aurora/RDS MySQL and PostgreSQL, Redshift). See [credential-security.md](references/credential-security.md).\n\n- You MUST confirm with user before creating a new Secrets Manager secret\n- You MUST NOT write plaintext credentials into chat or logs\n- For IAM DB auth, no secret is needed\n\n### 6. Create the Glue Connection\n\nFollow the source-specific reference for connection properties:\n\n```bash\naws glue create-connection --connection-input '<JSON>' --region <REGION>\n```\n\nPrivate sources require `PhysicalConnectionRequirements` (SubnetId, SecurityGroupIdList, AvailabilityZone). See [network-setup.md](references/network-setup.md).\n\n### 7. Test the Connection\n\nYou MUST test before handing off. Testing is two-phase: a quick API check, then an engine-level verification.\n\n#### Phase A: Glue TestConnection (network and credential sanity check)\n\n```bash\naws glue test-connection --connection-name <NAME> --region <REGION>\n```\n\nThis validates that Glue can reach the source and authenticate. It does NOT prove the connection works end-to-end with the query engine the user plans to use.\n\n#### Phase B: Engine-level verification\n\nAfter TestConnection passes, verify the connection works with the user's intended engine by running a minimal query through it:\n\n- **Glue ETL (default):** Run a smoke-test Glue job that reads one row via the connection. See [troubleshooting.md](references/troubleshooting.md).\n- **Athena:** If the user plans to query via Athena with a federated connector, run a `SELECT 1` through the Athena connection to confirm the Lambda-based connector can reach the source.\n- **Glue Crawler:** If the user plans to crawl the source, run a test crawl on a single table.\n\nPhase B catches issues that TestConnection misses: driver compatibility at job runtime, catalog configuration, Spark-level serialization, and engine-specific auth flows (e.g., Snowflake SNOWFLAKE type works in ETL but not via JDBC crawlers).\n\nOn success in both phases, tell user the connection name is ready for `ingesting-into-data-lake`. On failure in either phase, Step 8.\n\n### 8. Troubleshoot (only if test failed)\n\nDiagnose in order: network, credentials, driver. See [troubleshooting.md](references/troubleshooting.md).\n\n**Constraints:**\n\n- You MUST check VPC routing, security groups, and S3 VPC endpoint before blaming credentials\n- You MUST verify Glue role can read the Secrets Manager secret\n- You MUST NOT rotate credentials without user confirmation\n\n## Argument Routing\n\n- No args: Walk through Steps 1-7 interactively\n- Source type keyword (e.g., `snowflake`, `oracle`): Skip to Step 2 with the type prefilled\n- Existing connection name: Skip to Step 7 (test) then Step 8 if failing\n- Hostname or RDS endpoint: Skip to Step 4 with the candidate prefilled\n\n## Gotchas\n\n- Glue's `SNOWFLAKE` connection type is distinct from `JDBC` configured for Snowflake. You MUST use `SNOWFLAKE` for Spark ETL jobs; do not use JDBC.\n- Connection names are immutable. Choose carefully.\n- `PhysicalConnectionRequirements.AvailabilityZone` MUST match the subnet's AZ or the connection fails at job runtime, not creation time.\n- IAM database authentication tokens expire in 15 minutes. The Glue job generates a fresh token on each connection; do not cache.\n- An S3 VPC gateway endpoint MUST exist in the VPC used by private-source connections. Without it, Glue jobs cannot read their scripts or write results to S3.\n\n## Troubleshooting\n\n| Error | Likely cause | Fix |\n|---|---|---|\n| `Connect timed out` | VPC routing, SG rule, or NAT gateway missing | See [troubleshooting.md](references/troubleshooting.md) |\n| `Access denied for user` / `ORA-01017` | Credentials wrong, Secrets Manager access missing, or IAM DB auth misconfigured | See [troubleshooting.md](references/troubleshooting.md) |\n| `No suitable driver found` | Custom driver JAR not set or wrong class name | See [troubleshooting.md](references/troubleshooting.md) |\n| `SSL handshake failed` | `JDBC_ENFORCE_SSL` mismatch between Glue and source | See [troubleshooting.md](references/troubleshooting.md) |\n| `UnableToFindVpcEndpoint` | S3 VPC endpoint missing | Create S3 gateway endpoint in the connection's VPC |\n\n## References\n\n- [jdbc-setup.md](references/jdbc-setup.md) -- Oracle, SQL Server, PostgreSQL, MySQL, RDS, Redshift\n- [snowflake-setup.md](references/snowflake-setup.md) -- Glue `SNOWFLAKE` type, auth modes\n- [bigquery-setup.md](references/bigquery-setup.md) -- Glue `BIGQUERY` type, GCP service accounts\n- [discovery.md](references/discovery.md) -- Finding existing connections and candidate sources\n- [credential-security.md](references/credential-security.md) -- Secrets Manager and IAM DB auth\n- [network-setup.md](references/network-setup.md) -- VPC, subnets, security groups, endpoints\n- [troubleshooting.md](references/troubleshooting.md) -- Connection errors and diagnostic flow\n"
}

SHA-256: 9d3e67a32a4cf5114d124fe89e1bbf51dbdeb446f57bdafac13f9b8b5368e504