← CerebriumCONTENT HISTORY

Update to Cerebrium

Snapshot Sep 30, 2026 · 23:15 UTC · version 0.1.0

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "cerebrium",
  "description": "Use for any Cerebrium task: deploying Python code to serverless GPU or CPU, writing or fixing cerebrium.toml, choosing hardware and regions, calling deployed endpoints (REST, streaming, WebSocket, async), autoscaling and concurrency, cold starts, secrets, CI/CD, and debugging a build or a running app from the terminal. Covers the cerebrium CLI, configuration defaults the API actually applies, accepted GPU identifiers with per-plan limits, and troubleshooting.",
  "included_files": [
    {
      "relative_path": "references/cli.md",
      "size_in_bytes": 5378
    },
    {
      "relative_path": "references/config.md",
      "size_in_bytes": 9705
    },
    {
      "relative_path": "references/hardware.md",
      "size_in_bytes": 6430
    },
    {
      "relative_path": "references/troubleshooting.md",
      "size_in_bytes": 4416
    }
  ],
  "skill_md_contents": "---\nname: cerebrium\ndescription: >-\n  Use for any Cerebrium task: deploying Python code to serverless GPU or CPU, writing or fixing\n  cerebrium.toml, choosing hardware and regions, calling deployed endpoints (REST, streaming,\n  WebSocket, async), autoscaling and concurrency, cold starts, secrets, CI/CD, and debugging a\n  build or a running app from the terminal. Covers the cerebrium CLI, configuration defaults the\n  API actually applies, accepted GPU identifiers with per-plan limits, and troubleshooting.\nlicense: MIT\nmetadata:\n  author: cerebrium\n  version: \"0.1.0\"\n---\n\n# Cerebrium\n\nCerebrium runs Python workloads on serverless GPU and CPU: REST endpoints, SSE streaming,\nWebSockets, and async jobs, with scale-to-zero and per-second billing. One `cerebrium.toml`\ndescribes hardware, scaling, dependencies and runtime; one CLI (`cerebrium`) drives everything.\n\nReach for it when the workload is an inference API for an LLM, an embedding model or a vision\nmodel; a real-time voice, video or streaming app; bursty traffic that should scale from zero\nwithout holding idle GPUs; a deployment that has to run in several regions for latency or data\nresidency; or a migration off Replicate, Hugging Face or Mystic, each of which has a guide under\n`https://cerebrium.ai/docs/migrations`.\n\nThis file carries the workflow and the rules. Load the reference that matches the task:\n\n| Read | When |\n| --- | --- |\n| [`references/cli.md`](https://github.com/CerebriumAI/cerebrium-skills/blob/master/skills/cerebrium/references/cli.md) | Running any `cerebrium` command: the full surface, flags, non-interactive auth, CI/CD, which commands cost money. |\n| [`references/config.md`](https://github.com/CerebriumAI/cerebrium-skills/blob/master/skills/cerebrium/references/config.md) | Writing or fixing `cerebrium.toml`: every key, the default the API applies when it is omitted, accepted ranges, rebuild triggers. |\n| [`references/hardware.md`](https://github.com/CerebriumAI/cerebrium-skills/blob/master/skills/cerebrium/references/hardware.md) | Choosing `compute`, `gpu_count`, `region`, `provider`, `compute_tier`: accepted GPU identifiers, per-GPU and per-plan limits, regional availability, storage. |\n| [`references/troubleshooting.md`](https://github.com/CerebriumAI/cerebrium-skills/blob/master/skills/cerebrium/references/troubleshooting.md) | A build that failed, an app that 5xxs or queues, slow cold starts, settings that reverted. |\n\n## Rules for agents\n\n1. **Deploys cost money.** `cerebrium deploy`, `cerebrium run` and `cerebrium apps scale` start\n   billable compute, and `cerebrium apps delete` is destructive. State what will run on what\n   hardware and get the user's confirmation before the first one in a session.\n2. **`cerebrium run` is not local.** It packages the working directory, uploads it, and executes\n   in the cloud on the hardware in `cerebrium.toml`. There is no local emulator.\n3. **A `cerebrium.toml` key you leave out is reset to its default on deploy**, not left alone,\n   and a misspelled key does nothing in the CLI while still reaching the backend. Keep every\n   value that matters in the file, spelled as in [references/config.md](references/config.md).\n4. **Never invent config keys or GPU identifiers.** Both are validated server-side and a wrong\n   value fails the deploy. The accepted sets are in the references.\n5. **Adapt an example before writing from scratch.** `https://github.com/CerebriumAI/examples`\n   holds runnable references (vLLM, SDXL, Pipecat voice agents, ASGI apps), each with a working\n   `cerebrium.toml`.\n6. **Check the live docs when this skill does not cover it**, rather than guessing: the\n   `cerebrium-docs` MCP server (search plus docs filesystem), any docs page as markdown by\n   appending `.md` to its URL, or the index at `https://cerebrium.ai/docs/llms.txt`.\n\n## First run: check state before acting\n\n```bash\ncerebrium version                   # installed? if not: pip install cerebrium\ncerebrium projects current          # authenticated, and pointed at the intended project?\n```\n\n`cerebrium login` opens a browser and fails without a TTY. In CI or headless environments set\n`CEREBRIUM_SERVICE_ACCOUNT_TOKEN` (or pass `--service-account-token`) instead: see\n[`references/cli.md`](https://github.com/CerebriumAI/cerebrium-skills/blob/master/skills/cerebrium/references/cli.md).\n\n## Zero to a deployed endpoint\n\nStarting with no account: create one at `https://dashboard.cerebrium.ai`. The dashboard is also\nwhere API keys and authentication tokens are created. Compute is billed per second; current rates\nand any starting credit are at `https://www.cerebrium.ai/pricing`.\n\n```bash\npip install cerebrium              # thin wrapper that fetches the Go binary on first use\ncerebrium login                    # interactive only, needs an account\ncerebrium init my-app && cd my-app\ncerebrium deploy\n```\n\nThe full loop:\n\n1. Create an account at `https://dashboard.cerebrium.ai`, then `cerebrium login`.\n2. `cerebrium init my-app` writes `main.py` and `cerebrium.toml`.\n3. Write a function in `main.py` that takes and returns JSON-serialisable values. Everything at\n   module scope runs once per replica at startup: load models there, not inside the function.\n4. Set the config\n   ([`references/config.md`](https://github.com/CerebriumAI/cerebrium-skills/blob/master/skills/cerebrium/references/config.md)).\n   Do not skip `disable_auth` (the scaffold ships `true`, which makes the endpoint public) or\n   `max_replicas` (default 1, the most common cause of queueing).\n5. `cerebrium run main.py::run --prompt \"test\"` executes remotely on the configured hardware.\n6. `cerebrium deploy` builds, uploads, starts the app, and prints the endpoint. Build output\n   streams from this command and nowhere else.\n7. Once running, `cerebrium logs APP_NAME` shows runtime logs.\n\n## Choosing the runtime\n\nCortex is the default and needs no configuration. A custom runtime means the container starts\nyour own web server, and you own routing, middleware and auth. Opt in by adding\n`[cerebrium.runtime.custom]` with an `entrypoint` and a matching `port`.\n\n| The app | Runtime | Why |\n| --- | --- | --- |\n| A Python function you want reachable as an endpoint | Cortex | Cerebrium builds the route, parses the request, applies auth. |\n| Streaming output token by token | Cortex | `yield` from the function and the response is SSE. A custom runtime buys nothing here. |\n| FastAPI, ASGI, Gradio, anything already serving its own HTTP | Custom | Two servers cannot both own the port. |\n| WebSockets, or anything bidirectional | Custom | Cortex serves HTTP only. Clients connect over `wss://`. |\n| A self-contained server such as vLLM or Triton, custom batching, custom auth | Custom | The process is already the server. |\n\n## Choosing a scaling metric\n\n`scaling_metric` picks what the autoscaler watches, `scaling_target` is the level it holds.\n\n| `scaling_metric` | Reach for it when | `scaling_target` means |\n| --- | --- | --- |\n| `concurrency_utilization` (default) | GPU inference, and anything whose request times vary | Percent of `replica_concurrency` held per replica. At `replica_concurrency = 200`, target 80 holds 160 in flight. |\n| `requests_per_second` | You have measured a rate one replica sustains | Requests per second per replica. Target 5 holds 5 req/s. |\n| `cpu_utilization` | CPU-bound work | Percent of `cpu`. At `cpu = 2`, target 80 holds 1.6 cores. |\n| `memory_utilization` | Memory-bound work | Percent of `memory`. At `memory = 10`, target 80 holds 8 GB. |\n\n`cpu_utilization` and `memory_utilization` need a live replica to measure, so the API rejects\nboth with `min_replicas = 0`, and rejects `scaling_buffer` alongside either. Ranges, the rest of\n`[cerebrium.scaling]`, and how to pick `load_balancing_algorithm` are in\n[`references/config.md`](https://github.com/CerebriumAI/cerebrium-skills/blob/master/skills/cerebrium/references/config.md).\n\n## Calling the endpoint\n\n```\nPOST https://api.cerebrium.ai/v4/{PROJECT_ID}/{APP_NAME}/{FUNCTION_NAME}\n```\n\n`PROJECT_ID` already includes its `p-` prefix (`p-abcd1234`), so the path reads\n`/v4/p-abcd1234/my-app/run`. Do not add a second `p-`.\n\n```bash\ncurl -X POST 'https://api.cerebrium.ai/v4/p-abcd1234/my-app/run' \\\n  -H 'Authorization: Bearer <JWT>' \\\n  -H 'Content-Type: application/json' \\\n  -d '{\"prompt\": \"hello\"}'\n```\n\nResponse: `{ \"run_id\": \"...\", \"run_time_ms\": 326.34, \"result\": { ... } }`\n\n- The token comes from the API Keys page of the dashboard, or from a service account.\n- With `disable_auth = true` the endpoint takes unauthenticated requests from anyone.\n- A function whose name starts with `_` is not exposed. Use that for helpers.\n- **Streaming**: `yield` from the function; the response is `text/event-stream` (SSE).\n- **Async**: append `?async=true` for fire-and-forget, bounded by `response_grace_period`\n  (default 900 seconds, ceiling 12 hours).\n- **WebSockets**: require a custom runtime (`[cerebrium.runtime.custom]`) and a `wss://` client.\n- Regional hostnames such as `api.aws.us-east-1.cerebrium.ai` still resolve but proxy through\n  the global router and add latency. Prefer `api.cerebrium.ai`.\n\n## Secrets and automatic environment variables\n\n```bash\ncerebrium secrets add KEY=VALUE OTHER=VALUE   # project-wide; --app APP_ID scopes to one app\n```\n\nSecrets arrive as environment variables, read at container start, so an existing replica needs a\nrestart or redeploy to see a new one. Set automatically for every app: `APP_NAME`, `PROJECT_ID`\n(`p-` prefixed), `BUILD_ID`, and `HF_HOME` (`/persistent-storage/.cache/huggingface`).\n"
}

SHA-256: d5b5f11443e434c1bc78f9131f22a2092fb16d9ef0d711de26e901167c2a06f9