← Files Catalyst by ZohoARCHIVED FILE

skills/catalyst-by-zoho/references/troubleshooting.md

23.8 KB · Oct 2, 2026 · 00:06 UTC

↓ Download file

# Catalyst Troubleshooting Guide

## When to use this file
Load this file when the user reports something is broken, asks "why is X failing", or needs
diagnostic help with: deployments, function errors, ZCQL/Data Store issues, MCP tool errors,
AppSail crashes, Circuits failures, cron problems, or API Gateway errors.

---

## Deployment Failures

### Function deployment fails
- **Missing `.class` files (Java)** — If uploaded via console, the `.class` files must be included.
  When deploying via CLI, `catalyst deploy` auto-compiles and creates missing dependency files.
  Fix: use `catalyst deploy` instead of manual console upload.
- **`catalyst.json` missing or malformed** — Must exist at project root. Never create manually;
  it is auto-generated by `catalyst init`. If missing, re-run `catalyst init`.
- **Wrong directory structure** — Functions must be under `functions/`, web client under `client/`.
  Deviating from this structure causes deployment to fail silently or partially.
- **Deploy specific components only:**
  - Functions: `catalyst deploy --only functions`
  - AppSail: `catalyst deploy appsail`

### `is_deployed: false` in API responses does not indicate a problem
`List_All_Functions` and related API/MCP responses return `"is_deployed": false` for all
functions — including functions that are actively deployed and handling requests. This field
does not reflect actual deployment status and should be ignored. Verify function status from
the Catalyst console (Functions list) or by invoking the function endpoint directly.

### AppSail deployment fails or app not responding after deploy
- **Port mismatch** — The `port` field in `catalyst-config.json` must exactly match the port
  the app listens on. Using a different port causes requests to fail.
- **Cold start timeout** — When the first request hits an inactive AppSail app, a new instance
  is spawned. The app **must start listening on the port within 10 seconds** or the instance is
  killed and the next request triggers a new cold start cycle.
  Fix: minimize initialization logic; start listening as early as possible.
- **Missing env variable** — Always use `process.env.X_ZOHO_CATALYST_LISTEN_PORT || 9000`
  (Node.js) for the port. Hard-coding a port can work locally but breaks in Catalyst's environment.
- **Docker image build failure (custom runtime)** — Check the Dockerfile and ensure the image
  builds successfully locally before pushing to Catalyst's container registry.

### Slate deployment fails
- **Repository not in standard Catalyst project structure** — The `catalyst.json` file MUST be
  present in the GitHub repository's default branch. Without it, deployment fails.
- **Default branch has wrong structure** — Functions and Web Client Hosting won't be updated if
  the structure is incorrect; no changes are reflected.
- **Build command fails** — Check Slate build logs in the console for the specific error.
- **Recovery — Sync Now** — If a deployment fails, use the "Sync Now" feature in the Catalyst
  console to merge the latest Git commit to the current deployment. You can also rollback to a
  previous successful deployment from the console.
  CLI: `catalyst deploy slate` (all apps) or `catalyst deploy --only slate:appname` (specific app)

### GitHub-based deployment fails
- Repository must contain resources in standard Catalyst project directory format.
- `catalyst.json` MUST be in the repository for deployment to succeed.
- If deployment is unsuccessful, no changes are reflected in Functions or Web Client Hosting.
- Fix: verify directory structure, add/fix `catalyst.json`, and redeploy.

---

## Function Execution Errors

### Timeout errors
Execution time limits by function type:

| Function Type | Timeout |
|---------------|---------|
| Basic I/O | 30 seconds |
| Advanced I/O | 30 seconds |
| Event | 15 minutes |
| Cron | 15 minutes |
| Integration | 30 seconds |
| Job | 15 minutes |
| Browser Logic | 30 seconds |

- Use `context.getRemainingExecutionTimeMs()` inside the function to check remaining time.
- Use `context.getMaxExecutionTimeMs()` to get the configured maximum (constant value).
- For operations exceeding 30 seconds, switch to Event, Job, or Cron function types (15-min limit).
- For operations exceeding 15 minutes, migrate to AppSail (no function-level timeout).
- For large batch processing, use Circuits batch state — it runs multiple function instances
  in parallel, avoiding timeout on a single long-running execution.

### Concurrency limit reached (429 error)
Error message: `"CONCURRENCY_LIMIT_REACHED FOR THE FEATURE FUNCTIONS"`
- **Production limit**: 1500 concurrent executions (for a function that runs in 10ms)
- **Development limit**: 1000 concurrent executions (for a function that runs in 10ms)
- Fix: implement retry with exponential backoff in the caller, or route heavy processing
  through Job Scheduling for async execution instead of direct function invocation.

### Memory issues
- Default function memory: 128 MB
- Default AppSail memory: 512 MB
- Configurable range: 128–1024 MB (functions), 256–2048 MB (AppSail)
- Use DevOps → APM to identify memory-intensive operations and compare response times
  across memory settings. Start at the lowest setting and increase as needed.

### Basic I/O limitations causing unexpected errors
- Basic I/O returns **STRING only** — if you need JSON, use Advanced I/O.
- `basicIO.write()` can only be called **once per execution**. Calling it multiple times
  causes unexpected behavior.
- Basic I/O does **not support HTTP request/response headers**. Use Advanced I/O for
  custom status codes, headers, or content types.

### Advanced I/O — Node.js `res` object issues
In `node20`, the response object is a raw `http.ServerResponse`, **not** an Express response.
- `res.status()`, `res.json()`, `res.send()` do NOT exist.
- Use `res.writeHead(statusCode, headers)` and `res.end(JSON.stringify(data))`.
- Helper pattern to use in all Advanced I/O functions:
  ```javascript
  function sendJson(res, statusCode, data) {
    res.writeHead(statusCode, { 'Content-Type': 'application/json' });
    res.end(JSON.stringify(data));
  }
  ```

---

## Data Store / ZCQL Errors

### ZCQL query errors
- **"Empty query" error** — Wrong body key. Use `{ "zcql": "SELECT ..." }` not `{ "query": "..." }`.
- **Max rows** — ZCQL returns a maximum of 300 rows per query. Paginate with
  `LIMIT offset, count` (e.g., `LIMIT 0, 300`, then `LIMIT 300, 300`).
- **Max columns per SELECT** — ZCQL limits `SELECT` to 20 columns per query. Use explicit column
  names instead of `SELECT *` on tables with more than 20 columns.
- **Case sensitivity** — Table names and column names are case-sensitive; must match the
  console exactly.
- **String quoting** — String values in ZCQL must be in **single quotes**: `WHERE name = 'Alice'`.
- **Row ID field** — Use `ROWID` (uppercase), not `id`. The system primary key is always `ROWID`.
- **Result unwrapping** — `executeZCQLQuery` returns `[{ TableName: { ROWID: ..., col: ... } }]`.
  Always unwrap: `result.map(r => r.TableName)`. The key matches the table name as defined in
  the console (case-sensitive).

### Column creation errors
- **`INVALID_INPUT: max_length cannot be null`** — `varchar` columns MUST specify `max_length`
  (e.g., `"max_length": 255`). Omitting this field always causes this error.
- **System columns already exist** — Do NOT create `ROWID`, `CREATORID`, `CREATEDTIME`, or
  `MODIFIEDTIME` columns. Catalyst adds these automatically to every table.
- **Boolean fields as strings** — All boolean column properties (`is_mandatory`, `is_unique`,
  `search_index_enabled`, `audit_consent`) must be passed as **string values**: `"true"` or
  `"false"`, not JSON booleans (`true` / `false`).
- **Batch column creation** — If one column definition in a batch `Create_Column` call is invalid,
  the entire batch fails. Validate all column types before sending.

### Data Store permissions error
Error: `"No privileges to perform this action"`
- DataStore table permissions default to **Read-only for App Users**.
- Insert, Update, and Delete must be explicitly enabled per table in the console:
  Data Store → [table] → Permissions → App User role.
- Alternative: use admin-scoped SDK for DataStore operations:
  `catalyst.initialize(req, { scope: 'admin' })`

### CREATEDTIME timezone issues
- Catalyst stores `CREATEDTIME` in the project's configured timezone (e.g. IST) WITHOUT an
  offset marker. Passing the raw string to `new Date()` treats it as UTC, producing timestamps
  that are hours off.
- Fix: always append the project timezone offset before parsing the date string.

### DataStore SDK methods hang silently in Job functions (Python SDK)
All `zcatalyst_sdk` DataStore `Table` methods (`get_paged_rows`, `delete_rows`, `insert_rows`,
etc.) hardcode `CredentialUser.USER` internally. Job functions run under admin credentials only —
`CredentialUser.USER` has no token, so every DataStore SDK call makes a request with no auth,
waits indefinitely, and fails silently. No exception is raised; the function eventually hits the
15-minute hard timeout.
- Fix: call the underlying requester directly with `CredentialUser.ADMIN`:
  ```python
  from zcatalyst_sdk.credential import CredentialUser

  rows = table._requester.request(
      method="GET",
      path=f"/datastore/table/{table_identifier}/tablerow",
      user=CredentialUser.ADMIN,
      timeout=10,
  )
  ```
- **Warning:** `_requester` is a private API and may change between SDK versions. Check
  compatibility after SDK upgrades.
- Cache SDK methods are NOT affected — they already use `CredentialUser.ADMIN` by default.

### Emoji / 4-byte UTF-8 silently corrupted
- Data Store does NOT support emoji or 4-byte UTF-8 characters (many CJK extensions).
- These are silently stored as `?`.
- Workaround: store a string key (e.g., `"happy"`) and map to emoji in application code.

---

## Cache Errors

### `segment.delete()` / `Delete_Cache_Item` does not remove the key
`segment.delete(key)` (SDK) and the `Delete_Cache_Item` MCP tool both set `cache_value = null`
and clear the TTL, but the key continues to exist indefinitely. A subsequent `segment.getValue(key)`
does NOT raise a "key not found" error — it returns HTTP 200 with `cache_value: null`. Code that
checks for key absence by catching an exception will incorrectly treat the null-value key as
present.
- Fix: check the value itself, not the exception:
  ```python
  lock_val = segment.get_value(key)  # does not raise even if deleted
  if lock_val:  # truthy check — null/empty means absent
      # key is live
  ```

### `segment.update()` without expiry preserves the original TTL (not the documented 48-hour default)
`segment.update(key, value)` called without an expiry argument does **not** reset TTL to 48 hours
as the documentation states. Instead, it **preserves the original TTL** from the initial `put()`.
The expiry clock keeps counting down from its original wall-clock time. This means:
- If you `put(key, value, 1)` (1 hour) then `update(key, newValue)` 30 minutes later, the key
  still expires 30 minutes after the update — not 48 hours later.
- Fix: always pass the same TTL used at creation time to be explicit:
  ```python
  segment.update(key, value, 1)  # 1-hour TTL, matching segment.put(key, value, 1)
  ```
- This applies to lock tokens, session state, and any key originally created with a TTL.

---

## Zoho MCP Tool Errors

### `PERMISSION_NEEDED`
Causes (check in this order):
1. **Wrong `projectId`** — The `id` returned by `List_All_Projects` is not always the correct
   project ID for tool calls. Get the correct ID from the Catalyst console URL:
   `https://console.catalyst.zoho.com/baas/<org_id>/project/<project_id>/...`
2. **"On Demand" auth not enabled** — Go to mcp.zoho.com → Connections → Edit →
   select "On Demand" → Update. Verify `Catalyst by Zoho` shows Status: Connected.
3. **Accessing Production** — Production requires separate authorization. Default to
   `"Development"` environment; only access production when explicitly requested.

### `INVALID_ORG`
- Wrong `Catalyst-org` header value.
- Fix: re-run `List_All_Organizations` (no parameters needed) to get the correct org ID.

### Tool not available / not found
- The tool has not been added to the MCP server.
- Fix: go to mcp.zoho.com → Config Tools → search for the tool → Add Now.

### MCP connection not authorized
- Fix: in the MCP console → Connections → enable "On Demand" authorization.

---

## AppSail Specific Errors

### Cold start behavior
- The first request to an inactive app triggers a cold start (new server instance spawned).
- The app must start listening on the configured port **within 10 seconds** of instance creation.
- If no process is found listening on the port within 10 seconds, the user instance is killed
  and the next request triggers a new cold start.
- Mitigation: keep startup/initialization code minimal; move heavy setup to after `listen()`.

### AppSail cross-origin issues with Slate frontend
- Slate-hosted frontends calling AppSail APIs may get blocked by Catalyst's auth layer
  (manifests as "Unable to Fetch" or "Failed to fetch").
- Fix: serve the frontend from AppSail itself using `express.static()` so all calls are
  same-origin. This eliminates CORS and auth-layer issues entirely.
- **Note:** This issue is specific to AppSail. For Serverless Functions, cross-domain
  from Slate DOES work — see the "Duplicate Access-Control-Allow-Origin" entry under Auth Errors.

---

## Auth Errors

### Server-side `userManagement().getCurrentUser()` throws 401 even when user is logged in
- Web client fetch calls to Catalyst functions must include `credentials: 'include'`.
  Without it, auth cookies are not forwarded.
- Fix: always add `credentials: 'include'` to fetch options:
  ```javascript
  fetch('/server/my_function/execute', {
    credentials: 'include',
    // ...
  });
  ```

### `catalyst.auth.getCurrentUser is not a function` in the Web SDK
- This method does not exist in the Web SDK.
- Use `catalyst.auth.isUserAuthenticated()` instead — it resolves with the full user object
  (`result.content.email_id`, `result.content.first_name`, etc.) on success, and rejects
  with 401 on failure (platform auto-redirects to login).

### `signOut()` crashes with `Cannot read properties of undefined`
- `catalyst.auth.signOut()` requires a redirect URL argument.
- Calling it without an argument crashes because the SDK calls `.startsWith("/")` on `undefined`.
- Fix: `catalyst.auth.signOut(window.location.origin);` (Slate apps) or `catalyst.auth.signOut(window.location.origin + '/app/index.html');` (legacy Web Client)
- Note: `constructSignOutUrl()` does not exist — do not use a two-step pattern.

### `Authorization: Bearer` header intercepted before the function handler runs
Catalyst validates any `Authorization: Bearer <token>` header as a Zoho OAuth token **before**
passing the request to the function handler — even when the endpoint's Security Rule has
`authentication: optional`. If your function uses its own Bearer token for server-to-server
auth (e.g. a shared secret), Catalyst returns `INVALID_TOKEN` before your code runs.
- Fix: use a non-standard header for application-level secrets:
  ```
  # Not intercepted — use this for custom auth
  X-My-App-Token: <secret>

  # Intercepted by Catalyst's auth layer — avoid for custom secrets
  Authorization: Bearer <secret>
  ```

### ZAID mismatch between environments
- ZAID (Zoho Application ID) differs between Development and Production environments.
- This is the #1 source of auth issues when promoting to production.
- Fix: always retrieve ZAID from the correct environment's console and update accordingly.

### Duplicate `Access-Control-Allow-Origin` header (Slate + Serverless Functions)

Error in browser console:
```
The 'Access-Control-Allow-Origin' header contains multiple values
'https://myapp.onslate.com, https://myapp.onslate.com', but only one is allowed.
```

**Cause:** The Catalyst gateway injects `Access-Control-Allow-Origin` when the Slate domain is
in Authorized Domains (Console → Authentication → Whitelisting). If your Express code ALSO sets
this header (via `cors()` middleware or manual `res.setHeader`), the browser receives two values
in one header and rejects the response.

**Fix:** Remove ALL Express/function-level CORS headers for production origins. Only set CORS
headers for `localhost` (local dev, where no gateway is present):
```javascript
app.use((req, res, next) => {
  const origin = req.headers.origin || '';
  if (/^http:\/\/localhost(:\d+)?$/.test(origin)) {
    res.setHeader('Access-Control-Allow-Origin', origin);
    res.setHeader('Access-Control-Allow-Credentials', 'true');
    res.setHeader('Access-Control-Allow-Methods', 'GET, POST, PUT, DELETE, OPTIONS');
    res.setHeader('Access-Control-Allow-Headers', 'Content-Type, Authorization');
    if (req.method === 'OPTIONS') return res.status(204).end();
  }
  next();
});
```

**Key rule:** The gateway owns CORS headers for deployed origins. Express must not touch them.

### `getCurrentUser()` returns `null` — collaborator vs app user

`userManagement().getCurrentUser()` calls the internal `/project-user/current` endpoint using the
user token. This only returns data for **registered app users** — users who signed up through
Catalyst's auth flow (signUp/signIn).

**Collaborators and project admins** (people added via the Catalyst console) are NOT registered
app users. `getCurrentUser()` returns `null` for them, causing `Cannot read properties of null`
errors if your code doesn't check for it.

**Fix — add a null check with admin-scope fallback:**
```javascript
const userApp = catalyst.initialize(req);            // user-scope
const userData = await userApp.userManagement().getCurrentUser();

if (!userData || !userData.user_id) {
  // Collaborator/admin — not in the project-user table.
  // Fall back to admin-scope user list, or use a default identity.
  const adminApp = catalyst.initialize(req, { scope: 'admin' });
  const allUsers = await adminApp.userManagement().getAllUsers();
  // Match by email or use a system identity
}
```

### User-scope vs admin-scope: when to use each

The SDK supports two initialization scopes. Using the wrong one causes auth failures:

| Scope | Init call | Use for | What it CAN'T do |
|-------|-----------|---------|-------------------|
| **User** (default) | `catalyst.initialize(req)` | `getCurrentUser()`, user-identity operations | DataStore writes (if App User perms not enabled) |
| **Admin** | `catalyst.initialize(req, { scope: 'admin' })` | DataStore CRUD, Stratus, ZCQL, Cache, all data operations | `getCurrentUser()` — throws "no user credentials present" |

**Pattern for apps that need both auth and data ops:**
```javascript
// User-scope for identity
const userApp = catalyst.initialize(req);
const currentUser = await userApp.userManagement().getCurrentUser();

// Admin-scope for data operations
const adminApp = catalyst.initialize(req, { scope: 'admin' });
const dataStore = adminApp.datastore();
const table = dataStore.table('MyTable');
```

### `Authorization` header is `undefined` inside the function

The Catalyst gateway **strips** the `Authorization` header after validating the token. It then
injects internal `x-zc-*` headers that the SDK reads directly. Do not try to read
`req.headers['authorization']` — it will be `undefined`. The SDK handles this internally via
`catalyst.initialize(req)`.

---

## Slate Deployment Errors

### `slate-config.toml` wiped by build commands

The `.catalyst/slate-config.toml` file lives inside the build output directory (e.g., `dist/`).
Build commands that clean the output directory (Vite `--clean`, Expo `--clear`, `rm -rf dist/`)
delete this file. Without it, `catalyst deploy slate` fails.

**Fix — recreate after every clean build:**
```bash
# Example for a React/Vite app
npm run build && mkdir -p dist/.catalyst && \
  echo -e 'framework = "static"\ndeployment_name = "default"' > dist/.catalyst/slate-config.toml && \
  catalyst deploy slate
```

### Assets returning 404 on Slate (framework `baseUrl` issue)

If your build tool has a `baseUrl` or `basePath` configured for a non-root path (e.g.,
`/server/my_function`), all JS/CSS asset URLs will be prefixed with that path on Slate — but
Slate serves from root `/`. This causes all assets to 404.

**Fix:** Remove `baseUrl`/`basePath` from your build config before building for Slate deployment.
Only set it when serving the frontend from inside a function or AppSail sub-path.

---

## Circuits Errors

### State execution failures
- Function state supports three default error handlers: **On TimeOut**, **On Authorization Failure**,
  **On Execution Failure**.
- For each, configure **Retry** (max attempts + delay) and **Fallback** (alternative state to
  go to when all retries fail).
- Custom error handlers match by Error Code or Error Message from the function response.
- HTTP status codes outside 200–299 trigger the error handler check for Basic I/O functions.

### Batch state issues
- Nested parallel states are **not supported** — success and failure states are end states and
  Catalyst does not support nested parallel states within a parallel state.
- Batch and circuit states only support **Retry** action on error — there is no Fallback for these
  (unlike function states, which support both Retry and Fallback).

---

## Cron / Job Scheduling Errors

### Cron auto-disabled after repeated failures
- Third-party URL crons are automatically disabled after **50 consecutive failures**.
- Cron functions (not URL-based) are **NOT auto-disabled** regardless of repeated failures.
- Fix: configure Application Alerts on the cron job to be notified immediately on failure,
  check execution history from the console, fix the underlying bug, and re-enable the cron.

---

## API Gateway Errors

### 429 Too Many Requests
- Throttling rate limit exceeded.
- Catalyst uses a **sliding window rate limiting algorithm**: checks request count in the window
  preceding the current second (not a fixed window).
- Two throttle types:
  - **General throttling**: max hits per time unit for all users combined.
  - **IP-based throttling**: max hits per IP per time unit.
- Fix: review and adjust rate limit settings in API Gateway console; implement client-side
  retry with exponential backoff.

---

## Event Listener Errors

### "Configuration Needed" error on rules
- **Integration account deleted** — All rules associated with the deleted account show
  "Configuration Needed". Fix: re-associate rules with a different account, or re-integrate
  the deleted account.
- **Service deleted** — All accounts in the service are removed, and all associated rules show
  the error. Fix: create a new service and re-integrate accounts.

---

## DevOps — Where to find logs

| What you need | Where to look in the console |
|---------------|------------------------------|
| Function execution logs | DevOps → Logs (filters auto-applied when accessed from function page) |
| AppSail instance logs | AppSail → Instances → click Logs icon → redirects to Catalyst Logs |
| Circuit execution details | Circuits → Execution History → click execution → View Logs |
| Cron execution history | Cron → Details → Execution History |
| Performance reports | DevOps → APM (Application Performance Monitoring) |
| Audit trail (console + app) | Settings → Audit Logs |
| Automated failure alerts | DevOps → Application Alerts |

---

## Quick diagnostic checklist

When an agent reports "something is broken", work through this in order:

1. **What component?** Function / AppSail / Slate / Circuits / Cron / Data Store / MCP tool
2. **What error message?** Look at the exact error text first.
3. **Is it a deployment issue or a runtime issue?**
   - Deployment: check `catalyst.json`, directory structure, Java `.class` files, port config
   - Runtime: check DevOps → Logs, look at execution history
4. **Is it an auth/permission issue?** Check ZAID environment, DataStore permissions, API Gateway rules
5. **Is it a data issue?** Check ZCQL syntax, 300-row limit, case sensitivity, column types

SHA-256: fa920b08bf57d48d15acc7e38a8c17595635237c947c20c1fa84bf0db32bbfab