← Mixpanel HeadlessCONTENT HISTORY

Update to Mixpanel Headless

Snapshot Sep 30, 2026 · 22:51 UTC · version 0.1.2

Collection source: not recorded for this historical snapshot.

WHAT CHANGED · RULE-BASED ANALYSIS

First saved snapshot

No earlier snapshot is available to establish a change.

Compare saved observations

Download comparison JSON
Full technical diff · 0 changed fields
Full snapshot data
{
  "name": "mixpanelyst",
  "description": "This skill should be used when the user asks about Mixpanel product analytics, event data, funnel analysis, retention curves, cohort analysis, segmentation queries, user behavior, conversion rates, churn, DAU/MAU, ARPU, revenue metrics, feature adoption, A/B test results, user paths, flow analysis, or any request to query, explore, visualize, or analyze Mixpanel data using Python. Also use when the user asks to read, write, or manage Mixpanel \"business context\" — the markdown documentation that grounds AI assistants in an organization's structure and goals.",
  "included_files": [
    {
      "relative_path": "scripts/auth_manager.py",
      "size_in_bytes": 14621
    },
    {
      "relative_path": "scripts/help.py",
      "size_in_bytes": 32074
    }
  ],
  "skill_md_contents": "---\nname: mixpanelyst\ndescription: This skill should be used when the user asks about Mixpanel product analytics, event data, funnel analysis, retention curves, cohort analysis, segmentation queries, user behavior, conversion rates, churn, DAU/MAU, ARPU, revenue metrics, feature adoption, A/B test results, user paths, flow analysis, or any request to query, explore, visualize, or analyze Mixpanel data using Python. Also use when the user asks to read, write, or manage Mixpanel \"business context\" — the markdown documentation that grounds AI assistants in an organization's structure and goals.\nallowed-tools: Bash Read Write WebFetch\n---\n\n# mixpanel_headless API Reference\n\nAnalyze Mixpanel data by writing and executing Python code using the `mixpanel_headless` library and `pandas`.\nBefore running bundled helper scripts, set `SKILL_DIR` to the absolute path of this\n`skills/mixpanelyst` directory.\n\n```python\nimport mixpanel_headless as mp\nws = mp.Workspace()\nresult = ws.query(\"Login\", last=30)\nprint(result.df.head())\n```\n\n## Query Engines\n\n| Question | Method | Returns |\n|----------|--------|---------|\n| How much? How many? Trends? | `ws.query()` | `QueryResult` |\n| Do users convert through a sequence? | `ws.query_funnel()` | `FunnelQueryResult` |\n| Do users come back? | `ws.query_retention()` | `RetentionQueryResult` |\n| What paths do users take? | `ws.query_flow()` | `FlowQueryResult` |\n| Who are they? What do they look like? | `ws.query_user()` | `UserQueryResult` |\n\nAll result types have a `.df` property returning a pandas DataFrame and a `.params` dict containing the bookmark JSON.\n`FlowQueryResult` also has `.graph` (NetworkX DiGraph) and `.anytree` (list of tree roots).\n\n**Quick lookups** use `python3 -c \"...\"` one-liners. **Multi-step analysis** writes `.py` files.\n\n## Discovery — ALWAYS Do Both Steps Before Querying\n\nGuessing event names causes silent empty results. Guessing API parameters causes TypeErrors and invalid queries. **Discover both the data schema AND the API surface before writing any query.**\n\n### Step 1: Discover the data schema\n\n```python\nimport mixpanel_headless as mp\nfrom mixpanel_headless import Filter, GroupBy, Metric\nws = mp.Workspace()\n\n# 1. Find real event names\nevents = ws.events()\ntop = ws.top_events(limit=10)\nprint(\"Events:\", events[:20])\nprint(\"Top:\", [(e.event, e.count) for e in top])\n\n# 2. Find real property names for the event you'll query\nprops = ws.properties(\"Login\")  # use an actual event name from step 1\nprint(\"Properties:\", props)\n\n# 3. (Optional) Check property values to validate filter inputs\nvals = ws.property_values(\"platform\", event=\"Login\")\nprint(\"Platforms:\", vals)\n```\n\n### Step 2: Discover the API surface with `help.py`\n\n**`help.py` is the source of truth for method signatures, parameter names, type constructors, and enum values.** The method signatures later in this document are summaries — always verify with `help.py` before using a method or type you haven't looked up.\n\n**Never guess parameter names.** If you're unsure whether a parameter is called `property` or `math_property`, or what arguments `GroupBy()` accepts, run `help.py` first. Wrong parameter names cause TypeErrors that waste tool calls.\n\n```bash\n# Look up a query method BEFORE writing the query\npython3 $SKILL_DIR/scripts/help.py Workspace.query\npython3 $SKILL_DIR/scripts/help.py Workspace.query_funnel\n\n# Look up types BEFORE constructing them\npython3 $SKILL_DIR/scripts/help.py Filter          # → classmethods: .equals(), .less_than(), etc.\npython3 $SKILL_DIR/scripts/help.py Metric          # → property=, NOT math_property=\npython3 $SKILL_DIR/scripts/help.py GroupBy          # → property, property_type only\npython3 $SKILL_DIR/scripts/help.py MathType         # → enum values\n\n# Look up result types to know what columns .df returns\npython3 $SKILL_DIR/scripts/help.py QueryResult\npython3 $SKILL_DIR/scripts/help.py FlowQueryResult\n\n# Search when you're not sure of the exact name\npython3 $SKILL_DIR/scripts/help.py search cohort   # → CohortBreakdown, CohortMetric, CohortDefinition, ...\npython3 $SKILL_DIR/scripts/help.py search retention # → query_retention, RetentionEvent, RetentionMathType, ...\n\n# List everything\npython3 $SKILL_DIR/scripts/help.py types            # all public types\npython3 $SKILL_DIR/scripts/help.py exceptions        # all exceptions\n```\n\nFor tutorials and guides: `WebFetch(url=\"https://mixpanel.github.io/mixpanel-headless/llms.txt\")`\n\n### Discovery method signatures\n\n```python\ndef events(self) -> list[str]: ...\n    # List all event names (cached).\n\ndef properties(self, event: str) -> list[str]: ...\n    # List all property names for an event (cached).\n\ndef property_values(self, property_name: str, *, event: str | None = None, limit: int = 100) -> list[str]: ...\n    # Get sample values for a property.\n\ndef top_events(self, *, type: Literal['general', 'average', 'unique'] = 'general', limit: int | None = None) -> list[TopEvent]: ...\n    # Get today's most active events. TopEvent has .event (str), .count (int), .percent_change (float).\n\ndef funnels(self) -> list[FunnelInfo]: ...\n    # List saved funnels.\n\ndef cohorts(self) -> list[SavedCohort]: ...\n    # List saved cohorts.\n\ndef list_bookmarks(self, bookmark_type: BookmarkType | None = None) -> list[BookmarkInfo]: ...\n    # List all saved reports (bookmarks).\n\ndef lexicon_schemas(self, *, entity_type: EntityType | None = None) -> list[LexiconSchema]: ...\n    # List Lexicon schemas (event/property definitions).\n\ndef clear_discovery_cache(self) -> None: ...\n    # Clear cached discovery results.\n# User Guide: WebFetch(url=\"https://mixpanel.github.io/mixpanel-headless/guide/discovery/index.md\")\n```\n\n## Exploratory Analysis Workflow\n\nWhen exploring an unfamiliar dataset or asked to \"find insights,\" follow this systematic approach. Do NOT skip to querying — explore first.\n\n### Step 1: Orient — Map the Event Schema\n\n```python\nimport mixpanel_headless as mp\nws = mp.Workspace()\n\nevents = ws.events()\ntop = ws.top_events(limit=15)\nprint(\"Events:\", events)\nprint(\"Top events:\", [(e.event, e.count) for e in top])\n\n# Profile the top 3-5 events by volume\nfor event in [e.event for e in top[:5]]:\n    props = ws.properties(event)\n    print(f\"\\n{event} ({len(props)} properties):\")\n    for p in props[:15]:\n        vals = ws.property_values(p, event=event, limit=10)\n        print(f\"  {p}: {vals}\")\n```\n\n### Step 2: Classify Properties\n\nInfer property types from sampled values to decide how to use each:\n\n- **Boolean**: values are `['true', 'false']` — segment with `group_by`, often pre-computed behavioral flags\n- **Low-cardinality categorical** (<10 values): `platform`, `tier`, `category` — use for `group_by`\n- **Numeric**: values parse as int/float: `price`, `total`, `count` — use for `math='average'` or `math='sum'` with `math_property=`\n- **High-cardinality** (>100 values): IDs, names — skip for `group_by`, may need custom property cleanup\n- **Temporal**: ISO dates or epoch values — use for time-based analysis\n\nProperty naming patterns that signal analytical value:\n- `is_*`, `has_*`, `was_*`, `post_*` → boolean flags, often pre-computed behavioral segments worth investigating\n- `*_total`, `*_count`, `*_value`, `*_amount` → numeric, aggregate with avg/sum/median\n- `*_name`, `*_type`, `*_category`, `*_tier` → categorical, use for breakdowns\n- `*_id`, `*_uuid` → identifiers, skip for breakdowns\n\n### Step 3: Scan for Significant Segments\n\nFor each boolean and low-cardinality categorical property on key events, run a quick breakdown against a numeric metric:\n\n```python\n# Example: scan all interesting properties on a purchase event\nnumeric_prop = 'order_total'  # or whatever the key metric is\ninteresting_props = [p for p in props if not p.endswith('_id')]\n\nfor prop in interesting_props:\n    vals = ws.property_values(prop, event=event, limit=10)\n    if len(set(vals)) <= 10:  # low cardinality — worth a breakdown\n        result = ws.query(event, math='average', math_property=numeric_prop,\n                           group_by=prop, last=90, mode='total')\n        print(f\"\\n{numeric_prop} by {prop}:\")\n        print(result.df.to_string(index=False))\n        # Flag segments where metric differs >15% from overall\n```\n\n### Step 4: Deep Dive on Significant Findings\n\nWhen a breakdown reveals a notable difference (>15% between segments):\n1. **Quantify**: calculate the exact ratio between segments\n2. **Cross-reference**: does this segment differ on other metrics too?\n3. **Investigate causally**: run funnels or retention filtered by the segment\n4. **Control for confounds**: add a second `group_by` dimension to check if the effect holds\n\n### Step 5: Analyze Messy String Properties\n\nWhen string properties have complex/unreadable values (e.g., campaign names from tools like Braze):\n1. Sample 15-20 values to identify the naming convention\n2. Look for structural patterns: date codes, targeting prefixes, channel suffixes, audience tags\n3. Design regex cleanup rules, one layer per structural element\n4. Create a custom property with `ws.create_custom_property(CreateCustomPropertyParams(...))`\n5. Verify by querying with the custom property as `group_by`\n\n### Custom Property Formula Reference\n\nFormulas use a SQL-like expression language. Variables (A, B, _A, etc.) map to properties via `composedProperties`.\n\n**Variable binding:** `LET(name, expression, body)` — define intermediate results:\n```\nLET(raw, A, REGEX_REPLACE(raw, \"pattern\", \"replacement\"))\nLET(x, A * B, IFS(x < 50, \"low\", x < 200, \"mid\", TRUE, \"high\"))\n```\n\n**Conditionals:** `IF(cond, then, else)`, `IFS(cond1, val1, cond2, val2, ..., TRUE, default)`\n\n**String functions:** `UPPER(s)`, `LOWER(s)`, `LEN(s)`, `LEFT(s, n)`, `RIGHT(s, n)`, `MID(s, start, count)`, `SPLIT(s, delim, n)`, `HAS_PREFIX(s, p)`, `HAS_SUFFIX(s, p)`, `PARSE_URL(s, \"domain\")`\n\n**Regex functions (PCRE2 engine):**\n- `REGEX_MATCH(haystack, pattern)` — returns true/false\n- `REGEX_EXTRACT(haystack, pattern, capture_group)` — returns match or capture group\n- `REGEX_REPLACE(haystack, pattern, replacement)` — replaces all matches\n\n**Regex engine quirks (Mixpanel-specific):**\n- **Case-insensitive by default** — use `(?-i)` to switch to case-sensitive matching within a pattern\n- **Backreferences work** — `$1`, `$2` capture groups and `$0` whole-match all work in `REGEX_REPLACE` replacements\n- **`{n,m}` quantifiers conflict with formula syntax** — curly braces are parsed as formula constructs. Use repeated character classes instead (e.g., `[0-9][0-9][0-9][0-9]` instead of `[0-9]{4}`)\n- **`\\d`, `\\w` shorthand classes don't work** — use `[0-9]`, `[A-Za-z0-9_]` explicitly\n- **Escape backslashes carefully** — in formula strings, `\\\\\\\\` may be needed for a literal `\\` depending on how the formula is constructed (Python string → JSON → regex engine)\n\n**CamelCase splitting** — insert space between lowercase→uppercase boundaries:\n```\nREGEX_REPLACE(text, \"(?-i)([a-z])([A-Z])\", \"$1 $2\")\n// ChickenSundaysApril → Chicken Sundays April\n```\n\n**Practical multi-step cleanup example** (campaign names from Braze):\n```\nLET(s1, REGEX_REPLACE(A, \"^[0-9][0-9][0-9][0-9][0-9]*_\", \"\"),\nLET(s2, REGEX_REPLACE(s1, \"^(NW|TARGETED|REGIONAL|NTL)_\", \"\"),\nLET(s3, REGEX_REPLACE(s2, \"_(Push|Email|NotificationCenter|ModalInAppMessage)_.*$\", \"\"),\nLET(s4, REGEX_REPLACE(s3, \"_\", \" \"),\n  REGEX_REPLACE(s4, \" +\", \" \")\n))))\n```\n\n**Type functions:** `STRING(x)`, `NUMBER(x)`, `BOOLEAN(x)`, `DEFINED(x)`\n\n**Math:** `+`, `-`, `*`, `/`, `%`, `MIN(a,b)`, `MAX(a,b)`, `FLOOR(n)`, `CEIL(n)`, `ROUND(n)`\n\n**Date:** `DATEDIF(start, end, unit)` — units: D, M, Y, MD, YM, YD. `TODAY()` for current date.\n\n**List:** `SUM(list)`, `ANY(x, list, expr)`, `ALL(x, list, expr)`, `FILTER(x, list, expr)`, `MAP(x, list, expr)`\n\n**Comparison:** `==`, `!=`, `<`, `>`, `<=`, `>=` (case-insensitive for strings), `IN` for list membership\n\n**Logical:** `AND`, `OR`, `NOT(x)`\n\n**Constants:** `TRUE`, `FALSE`, `UNDEFINED`\n\n**Creating a custom property via the API:**\n```python\nfrom mixpanel_headless import CreateCustomPropertyParams, ComposedPropertyValue\n\nparams = CreateCustomPropertyParams(\n    name=\"Clean Campaign Name\",\n    resource_type=\"events\",\n    display_formula='LET(raw, A, REGEX_REPLACE(REGEX_REPLACE(raw, \"^[0-9]+_\", \"\"), \"_\", \" \"))',\n    composed_properties={\n        \"A\": ComposedPropertyValue(\n            resource_type=\"event\", type=\"string\", value=\"campaign_name\",\n            label=\"Campaign Name\", property_default_type=\"string\",\n        )\n    }\n)\nprop = ws.create_custom_property(params)\nref = CustomPropertyRef(prop.custom_property_id)\nresult = ws.query(event, group_by=GroupBy(ref), last=30, mode='total')\n```\n\n## Workspace\n\n```python\nclass Workspace:\n    \"\"\"Unified entry point for Mixpanel data operations (042 redesign).\"\"\"\n\n    def __init__(\n        self,\n        *,\n        account: str | None = None,\n        project: str | None = None,\n        workspace: int | None = None,\n        target: str | None = None,\n        session: Session | None = None,\n    ) -> None:\n        \"\"\"Create a new Workspace. Resolution per axis is independent\n        (env > param > target > bridge > config); see\n        ``mixpanel_headless.auth_types`` and the resolver.\n\n        With ``session=`` supplied, all other axis kwargs are ignored\n        (full bypass).\n        \"\"\"\n        ...\n\n    # --- Properties (read-only) ---\n    account: Account            # Resolved Account (discriminated union)\n    project: Project            # Resolved Project\n    workspace: WorkspaceRef | None  # Resolved workspace, or None for lazy-resolve\n    session: Session            # The (account, project, workspace) tuple\n    api: MixpanelAPIClient      # Direct API client access (escape hatch)\n\n    def use(\n        self,\n        *,\n        account: str | None = None,\n        project: str | None = None,\n        workspace: int | None = None,\n        target: str | None = None,\n        persist: bool = False,\n    ) -> Self:\n        \"\"\"Switch any axis in-session. Returns self for chaining.\n        Preserves the underlying httpx.Client. With persist=True, also\n        writes to ~/.mp/config.toml [active]. ``target=`` is mutex\n        with the per-axis kwargs.\n        \"\"\"\n        ...\n```\n\nSupports context manager: `with mp.Workspace() as ws: ...`\n\n### Project & Workspace Management\n\n```python\ndef me(self, *, force_refresh: bool = False) -> Any: ...\n    # Get /me response for current credentials (cached 24h).\n\ndef projects(self) -> list[Project]: ...\n    # List accessible projects (v3; returns Project records — id, name,\n    # organization_id, timezone). Replaces deprecated discover_projects().\n\ndef workspaces(self, *, project_id: str | None = None) -> list[WorkspaceRef]: ...\n    # List workspaces in a project (v3; returns WorkspaceRef records —\n    # id, name, is_default). Replaces deprecated discover_workspaces().\n\ndef list_workspaces(self) -> list[PublicWorkspace]: ...\n    # List all public workspaces for the current project (App API).\n\ndef resolve_workspace_id(self) -> int: ...\n    # Auto-discover and resolve workspace ID (lazy-resolve helper).\n\ndef close(self) -> None: ...\n    # Close all resources (HTTP client). Idempotent.\n```\n\n> **Removed (042 redesign — FR-038):** `Workspace.workspace_id` property,\n> `set_workspace_id()`, `switch_project()`, `switch_workspace()`,\n> `discover_projects()`, `discover_workspaces()`, `current_project`,\n> `current_credential`, `test_credentials()`. Use `ws.session.workspace_id`,\n> `ws.use(workspace=N)`, `ws.use(project=P)`, `ws.projects()`,\n> `ws.workspaces()`, `ws.project`, `ws.account`, and `mp.accounts.test(NAME)`\n> respectively.\n\n### Insights Query\n\nRun `python3 $SKILL_DIR/scripts/help.py Workspace.query` for the full signature.\n\n```python\ndef query(\n    self,\n    events: str | Metric | CohortMetric | Formula | Sequence[...],\n    *,\n    from_date: str | None = None,        # YYYY-MM-DD, overrides last\n    to_date: str | None = None,          # YYYY-MM-DD, requires from_date\n    last: int = 30,                      # relative days (ignored if from_date set)\n    unit: QueryTimeUnit = 'day',\n    math: MathType = 'total',            # aggregation: total, unique, dau, average, sum, ...\n    math_property: str | None = None,    # top-level shorthand; Metric() uses property= instead\n    per_user: PerUserAggregation | None = None,\n    percentile_value: int | float | None = None,\n    group_by: str | GroupBy | CohortBreakdown | FrequencyBreakdown | list[...] | None = None,\n    where: Filter | FrequencyFilter | list[...] | None = None,\n    formula: str | None = None,          # e.g. \"(B / A) * 100\", requires 2+ events\n    formula_label: str | None = None,\n    rolling: int | None = None,\n    cumulative: bool = False,\n    mode: Literal['timeseries', 'total', 'table'] = 'timeseries',\n    time_comparison: TimeComparison | None = None,\n    data_group_id: int | None = None,\n) -> QueryResult:\n    # .df columns: timeseries → [date, event, count]\n    #              total → [event, count]\n    #              with group_by → adds segment column\n    ...\n```\n\n### Funnel Query\n\nRun `python3 $SKILL_DIR/scripts/help.py Workspace.query_funnel` for the full signature.\n\n```python\ndef query_funnel(\n    self,\n    steps: list[str | FunnelStep],      # at least 2 steps required\n    *,\n    conversion_window: int = 14,\n    conversion_window_unit: Literal['second', 'minute', 'hour', 'day', 'week', 'month', 'session'] = 'day',\n    order: Literal['loose', 'any'] = 'loose',\n    from_date: str | None = None, to_date: str | None = None, last: int = 30,\n    unit: QueryTimeUnit = 'day',\n    math: FunnelMathType = 'conversion_rate_unique',\n    math_property: str | None = None,\n    group_by: str | GroupBy | CohortBreakdown | list[...] | None = None,\n    where: Filter | list[Filter] | None = None,\n    exclusions: list[str | Exclusion] | None = None,\n    holding_constant: str | HoldingConstant | list[...] | None = None,\n    mode: Literal['steps', 'trends', 'table'] = 'steps',\n    reentry_mode: FunnelReentryMode | None = None,\n    time_comparison: TimeComparison | None = None,\n    data_group_id: int | None = None,\n) -> FunnelQueryResult:\n    # .df columns: [step, event, count, step_conv_ratio, avg_time]\n    # .overall_conversion_rate: float\n    ...\n```\n\n### Retention Query\n\nRun `python3 $SKILL_DIR/scripts/help.py Workspace.query_retention` for the full signature.\n\n```python\ndef query_retention(\n    self,\n    born_event: str | RetentionEvent,\n    return_event: str | RetentionEvent,\n    *,\n    retention_unit: TimeUnit = 'week',\n    alignment: RetentionAlignment = 'birth',\n    bucket_sizes: list[int] | None = None,\n    from_date: str | None = None, to_date: str | None = None, last: int = 30,\n    unit: QueryTimeUnit = 'day',\n    math: RetentionMathType = 'retention_rate',\n    group_by: str | GroupBy | CohortBreakdown | list[...] | None = None,\n    where: Filter | list[Filter] | None = None,\n    mode: RetentionMode = 'curve',\n    unbounded_mode: RetentionUnboundedMode | None = None,\n    retention_cumulative: bool = False,\n    time_comparison: TimeComparison | None = None,\n    data_group_id: int | None = None,\n) -> RetentionQueryResult:\n    # .df columns: [cohort_date, bucket, count, rate]  (+ segment with group_by)\n    # .average: synthetic average across cohorts\n    ...\n```\n\n### Flow Query\n\nRun `python3 $SKILL_DIR/scripts/help.py Workspace.query_flow` for the full signature.\n\n```python\ndef query_flow(\n    self,\n    event: str | FlowStep | Sequence[str | FlowStep],\n    *,\n    forward: int = 3, reverse: int = 0,\n    from_date: str | None = None, to_date: str | None = None, last: int = 30,\n    conversion_window: int = 7,\n    conversion_window_unit: Literal['day', 'week', 'month', 'session'] = 'day',\n    count_type: Literal['unique', 'total', 'session'] = 'unique',\n    cardinality: int = 3,\n    collapse_repeated: bool = False,\n    hidden_events: list[str] | None = None,\n    mode: Literal['sankey', 'paths', 'tree'] = 'sankey',\n    where: Filter | list[Filter] | None = None,\n    segments: str | GroupBy | CohortBreakdown | FrequencyBreakdown | list[...] | None = None,\n    exclusions: list[str] | None = None,\n    data_group_id: int | None = None,\n) -> FlowQueryResult:\n    # .df, .graph (NetworkX DiGraph), .anytree (tree mode)\n    # .top_transitions(n), .drop_off_summary()\n    ...\n```\n\n### User Profile Query\n\nRun `python3 $SKILL_DIR/scripts/help.py Workspace.query_user` for the full signature.\n\n```python\ndef query_user(\n    self,\n    *,\n    where: Filter | list[Filter] | str | None = None,\n    cohort: int | CohortDefinition | None = None,\n    properties: list[str] | None = None,\n    sort_by: str | None = None,\n    sort_order: Literal['ascending', 'descending'] = 'descending',\n    limit: int | None = 1,              # None = fetch all matching\n    search: str | None = None,\n    distinct_id: str | None = None,     # single user lookup\n    distinct_ids: list[str] | None = None,  # batch lookup\n    group_id: str | None = None,        # query group profiles\n    as_of: str | int | None = None,     # point-in-time\n    mode: Literal['profiles', 'aggregate'] = 'aggregate',\n    aggregate: Literal['count', 'extremes', 'percentile', 'numeric_summary'] = 'count',\n    aggregate_property: str | None = None,\n    percentile: float | None = None,\n    segment_by: list[int] | None = None,\n    parallel: bool = False, workers: int = 5,\n    include_all_users: bool = False,\n) -> UserQueryResult:\n    # .df, .total, .profiles\n    ...\n```\n\n### Build Params (without executing)\n\nSame parameters as the corresponding query methods, but return `dict[str, Any]` bookmark params without making an API call. Useful for creating saved reports (bookmarks).\n\n```python\ndef build_params(self, events, **kwargs) -> dict[str, Any]: ...\ndef build_funnel_params(self, steps, **kwargs) -> dict[str, Any]: ...\ndef build_retention_params(self, born_event, return_event, **kwargs) -> dict[str, Any]: ...\ndef build_flow_params(self, event, **kwargs) -> dict[str, Any]: ...\ndef build_user_params(self, **kwargs) -> dict[str, Any]: ...\n```\n\n### Multi-Step Analysis Patterns\n\nEvery query engine has parameters that look like simple settings but are actually analytical choices with outsized influence on results. Before running any query, apply these principles:\n\n- [ ] **Find the master dial.** Each engine has one parameter (or small set) that reshapes all downstream metrics. Changing it changes the story. Know which parameter it is and choose deliberately — don't accept defaults blindly.\n- [ ] **Match parameters to the domain.** There are no universal \"correct\" values. A social app needs daily retention; a B2B SaaS needs monthly. A food-ordering funnel needs a 1-hour window; an onboarding funnel needs 14 days. The product's natural usage cadence dictates the setting.\n- [ ] **Distrust averages.** Averages include outliers — one extreme value distorts the whole metric. Use medians (`math='median'`, `percentile=50`) to see what's typical. If the mean and median diverge, the distribution is skewed and the mean is misleading.\n- [ ] **Counting methodology is a modeling choice.** \"Unique users,\" \"total events,\" and \"sessions\" aren't just modes — they answer fundamentally different questions. \"How many people?\" vs \"How much activity?\" vs \"How many engagement moments?\" Choose the counting method that matches the business question.\n- [ ] **Know the silent defaults.** Parameters are sometimes silently ignored (e.g., `math_property` with `math='unique'`), silently constraining (e.g., no funnel re-entry by default), or silently inflating (e.g., `unbounded_mode='carry_forward'` in retention). If results look surprising, check whether a default is shaping them.\n- [ ] **Sweep, don't guess.** When unsure which parameter value to use, try several and observe how metrics shift. Where the metric stabilizes or diverges reveals the true signal. The code examples below demonstrate this for each engine.\n\n#### Comparing Segments Across Multiple Dimensions\n\nWhen a single breakdown shows a difference, verify it holds across dimensions:\n\n```python\n# Step 1: Find the interesting segment\nresult = ws.query(event, math='average', math_property='order_total',\n                   group_by='deal_sweet_spot', last=90, mode='total')\n# Found: deal_sweet_spot=true has 37% higher AOV\n\n# Step 2: Check if it holds across another dimension\nresult = ws.query(event, math='average', math_property='order_total',\n                   group_by=['deal_sweet_spot', 'platform'], last=90, mode='total')\n# Does the sweet spot hold for both iOS and Android?\n\n# Step 3: Check segment rates across a third dimension\nresult = ws.query(event, math='unique',\n                   group_by=['loyalty_tier', 'deal_sweet_spot'], last=90, mode='total')\n# Which tier is most likely to achieve the sweet spot?\n```\n\n#### Insights Analysis: MathType, Per-User, and the Unit of Analysis\n\n**MathType is the most critical Insights parameter.** It defines what you're counting — `total` (event volume), `unique` (user reach), `dau/wau/mau` (time-bounded engagement), `average/median/percentile` (property distributions), `sum` (revenue totals). Choosing the wrong MathType answers the wrong question silently. Match MathType to the business question: engagement → `dau` or `wau`; revenue → `sum` or `average` with `math_property`; adoption → `unique`; intensity → `total`.\n\n**Per-user aggregation is a two-stage process** that fundamentally changes the unit of analysis. `per_user='average'` with `math='average'` first computes each user's average, then averages across users. This is NOT the same as a global average — a power user with 1000 events and a casual user with 2 events contribute equally. This is often the right choice (prevents power users from dominating) but it changes the story dramatically:\n\n```python\n# Sweep MathTypes to understand an event from multiple angles\nevent = 'Purchase'  # use a real event name\nprop = 'order_total'  # use a real numeric property\n\nfor mt in ['total', 'unique', 'dau', 'average', 'median', 'sum']:\n    kwargs = {'math': mt, 'last': 30, 'mode': 'total'}\n    if mt in ('average', 'median', 'sum'):\n        kwargs['math_property'] = prop\n    result = ws.query(event, **kwargs)\n    print(f\"{mt:>10}: {result.df['count'].iloc[0]:>12,.2f}\")\n# total = event volume, unique = user reach, dau = daily engagement,\n# average/median = typical transaction, sum = total revenue\n\n# Per-user aggregation: compare global average vs per-user average\nglobal_avg = ws.query(event, math='average', math_property=prop, last=30, mode='total')\nper_user_avg = ws.query(event, math='average', math_property=prop,\n                         per_user='average', last=30, mode='total')\nprint(f\"Global avg: {global_avg.df['count'].iloc[0]:.2f}\")\nprint(f\"Per-user avg: {per_user_avg.df['count'].iloc[0]:.2f}\")\n# If these differ significantly, power users are skewing the global average\n```\n\n**Prefer medians over averages for property distributions** — same principle as funnel time-to-convert. Use `math='median'` instead of `math='average'`, or `math='percentile'` with `percentile_value=50`. Averages include outliers; medians reveal what's typical.\n\n**Silent traps:** `math_property` is silently ignored with non-property MathType (e.g., `math='unique'` discards `math_property`). `rolling` reduces data point count without warning (a 30-day rolling window over 59 days produces ~29 points, not 59). `unit` is silently ignored in `mode='total'`.\n\n#### Funnel Analysis: Windows, Modes, and Time\n\nFunnel queries return time data in step metadata columns (`avg_time`, `avg_time_from_start`) alongside conversion rates.\n\n**Conversion window is the most critical funnel parameter.** It defines the maximum time a user has to complete the funnel from their first step. It affects every other metric — conversion rate, time-to-convert, and segment comparisons all shift dramatically with window size.\n\n**Choosing a window:** Match it to the user journey being measured. Short funnels (ordering food, adding to cart) need tight windows — hours, not days. Long funnels (onboarding, subscription purchase) need wider windows — days or weeks. When unsure, experiment:\n\n```python\n# Try progressively tighter windows to find where signal emerges\nfor window, unit in [(14, 'day'), (7, 'day'), (1, 'day'), (12, 'hour'), (6, 'hour')]:\n    result = ws.query_funnel(steps, last=90,\n        conversion_window=window, conversion_window_unit=unit)\n    final = result.df[result.df['step'] == result.df['step'].max()].iloc[0]\n    print(f\"{window}{unit}: conv={final['overall_conv_ratio']:.3f} \"\n          f\"time={final['avg_time_from_start']/3600:.1f}h\")\n# Look for: conversion stabilizing, time differences appearing at tighter windows\n```\n\n**Conversion counting modes** change what \"conversion\" means:\n- `conversion_rate_unique` (default): unique users who completed. No re-entry — first attempt in the window or out.\n- `conversion_rate_total`: total completions. One user can count multiple times.\n- Combine with `reentry_mode='basic'` or `reentry_mode='optimized'` for multiple attempts. Optimized re-entry picks the best completion path.\n\n**Time-to-convert: prefer medians over averages.** Average time includes outliers — one slow user inflates the average; one fast user pulls it down. Use median or percentiles for true speed trends:\n\n```python\n# Median time via funnel math\nresult = ws.query_funnel(steps, last=90, math='median')\n\n# Compare segments with filtered funnels + tight window\nios = ws.query_funnel(steps, where=Filter.equals('platform', 'iOS'),\n    last=90, conversion_window=6, conversion_window_unit='hour')\nandroid = ws.query_funnel(steps, where=Filter.equals('platform', 'Android'),\n    last=90, conversion_window=6, conversion_window_unit='hour')\n# Compare avg_time_from_start on matching steps\n```\n\n**When comparing segments across funnels:** always try at least 2-3 conversion windows. A difference invisible at 14 days may be stark at 6 hours. This is especially true for speed comparisons — tighter windows filter out noise and reveal which segment completes faster.\n\n**`order` changes what \"conversion\" means.** `'loose'` (default) requires steps in sequence but allows other events between them. `'any'` requires all steps in any order — a user who does C→B→A counts as converting. The difference is dramatic: loose funnels measure sequential workflows; any-order funnels measure feature adoption breadth. When unsure, run both and compare:\n\n```python\n# order='loose' vs 'any' — same steps, fundamentally different questions\nfor ord in ['loose', 'any']:\n    result = ws.query_funnel(steps, last=90, order=ord)\n    print(f\"order={ord}: {result.overall_conversion_rate:.3f}\")\n# If 'any' >> 'loose', users are completing all steps but not in the expected order\n# This often reveals UX issues — users accomplish the goal but not via the designed path\n```\n\n**`holding_constant` isolates cross-step consistency.** Hold a property like `'platform'` or `'device_id'` constant and users who change values between steps (e.g., sign up on iOS, purchase on Android) are excluded. This reveals single-device vs cross-device conversion and is essential for understanding journeys that span platforms. Maximum 3 properties.\n\n**Exclusions disqualify tainted journeys.** `exclusions=[\"Logout\"]` removes users who logged out between funnel steps — unlike flow `hidden_events`, exclusions completely remove users from the funnel. Use for support escalation events, churn signals, or any action that taints the conversion path. Control which steps the exclusion applies to with `Exclusion(\"Logout\", from_step=0, to_step=2)` (0-indexed).\n\n**Per-step filters narrow individual steps without affecting others.** `FunnelStep(\"Purchase\", filters=[Filter.greater_than(\"amount\", 50)])` restricts which Purchase events count, but doesn't filter Signup events. Global `where` filters ALL steps. This distinction is subtle but powerful: filter the population with `where`, filter the definition of a step with per-step filters.\n\n**Session windows are a distinct paradigm.** `conversion_window_unit='session'` constrains the entire funnel to a single engagement session — no multi-session hops. This reveals true in-session conversion behavior, separate from users who spread a journey across days. The third counting mode, `math='conversion_rate_session'`, counts sessions rather than users or events (requires `conversion_window_unit='session'`).\n\n#### Retention Analysis: The Cohort Bucketing Triple\n\n**`retention_unit` + `alignment` + `bucket_sizes` define your entire retention model** — retention's equivalent of the funnel conversion window. `retention_unit` groups users into cohorts (day/week/month). `alignment` anchors cohorts (`birth` = each user's clock starts from their event; `interval_start` = snap to calendar boundaries). `bucket_sizes` sets measurement points. Changing any one reshapes all downstream metrics.\n\n**Match `retention_unit` to your product's natural usage cadence.** Daily products (social, messaging) need `retention_unit='day'`. Weekly products (task management, fitness) need `'week'`. Monthly products (subscriptions, B2B SaaS) need `'month'`. When unsure, experiment:\n\n```python\n# Sweep retention_unit to find natural product cadence\nborn, ret = 'Signup', 'Login'  # use real event names\nfor ru in ['day', 'week', 'month']:\n    result = ws.query_retention(born, ret, retention_unit=ru, last=90)\n    avg = result.average\n    if avg is not None and len(avg) > 1:\n        bucket_1_rate = avg.iloc[1]['rate'] if 'rate' in avg.columns else None\n        print(f\"{ru:>6} retention: bucket 1 = {bucket_1_rate}\")\n# The unit where bucket-1 retention is highest reveals natural usage cadence\n\n# Custom buckets for milestone-based retention (days 1, 3, 7, 14, 30)\nresult = ws.query_retention(born, ret, retention_unit='week',\n    bucket_sizes=[1, 3, 7, 14, 30], unit='day', last=90)\nprint(result.df[result.df['cohort_date'] == '$overall'])\n# Day 1 = activation, Day 7 = habit formation, Day 30 = long-term retention\n\n# Compare alignment modes — can shift results dramatically\nfor align in ['birth', 'interval_start']:\n    result = ws.query_retention(born, ret, retention_unit='week',\n        alignment=align, last=90)\n    print(f\"\\nalignment={align}:\")\n    print(result.average.head() if result.average is not None else \"No data\")\n```\n\n**Be wary of unbounded modes — they inflate retention.** `unbounded_mode='carry_forward'` credits future returns to past buckets — a user who returns only on day 30 gets counted as retained in all buckets from 30 onward. `carry_back` inflates early buckets instead. Useful for \"did they ever engage?\" analysis but distorts standard retention curves.\n\n**`retention_cumulative=True` masks re-engagement gaps.** Cumulative retention creates monotonically increasing curves where each bucket includes all prior buckets. This hides whether users who returned in week 1 ALSO returned in week 2. Standard (non-cumulative) retention reveals re-engagement patterns and true habit formation.\n\n**Counting methodology:** `math='retention_rate'` (% who returned — the default), `math='unique'` (count who returned), `math='total'` (how many times they returned — events, not users). `total` reveals engagement intensity; a user logging in 5 times in bucket 1 counts as 5, not 1. Like funnels, the counting choice changes the question.\n\n#### Flow Analysis: Windows, Cardinality, and Signal-to-Noise\n\n**Cardinality controls signal-to-noise** — the most important flow-specific parameter. Low cardinality (2-3) reveals dominant paths — the main story. High cardinality (10+) reveals edge cases and niche journeys. Start low to find the narrative, then increase to find exceptions.\n\n**`conversion_window` matters for flows too** — identical concept to funnels. Session-based windows (`conversion_window_unit='session'`) reveal in-app behavior within a single engagement. Calendar windows reveal multi-day journeys. A tight window isolates intentional workflows; a wide window captures exploratory meandering:\n\n```python\n# Sweep cardinality to find signal-to-noise sweet spot\nevent = 'Login'  # use a real anchor event\nfor card in [2, 3, 5, 10]:\n    result = ws.query_flow(event, forward=3, cardinality=card, last=30)\n    transitions = result.top_transitions(5)\n    print(f\"\\ncardinality={card}: {len(transitions)} top transitions\")\n    for src, dst, count in transitions[:3]:\n        print(f\"  {src} → {dst}: {count}\")\n# Low cardinality = clear narrative; high cardinality = exhaustive but noisy\n\n# Compare count types (same principle as funnels)\nfor ct in ['unique', 'total', 'session']:\n    result = ws.query_flow(event, forward=3, count_type=ct, last=30)\n    dropoff = result.drop_off_summary()\n    print(f\"\\n{ct}: step 0 dropoff = {dropoff}\")\n# unique = how many people; total = how much activity; session = how many sessions\n\n# Compare collapse_repeated to separate intent from noise\nfor collapse in [False, True]:\n    result = ws.query_flow(event, forward=3, collapse_repeated=collapse,\n                            cardinality=5, last=30)\n    print(f\"\\ncollapse_repeated={collapse}:\")\n    for src, dst, count in result.top_transitions(3):\n        print(f\"  {src} → {dst}: {count}\")\n```\n\n**`collapse_repeated` changes what \"a path\" means.** With `False` (default), A→A→A→B is a distinct path from A→B — repetitive clicks look like distinct journeys. With `True`, consecutive duplicates merge, revealing intent over noise. Toggle this to see both the raw behavior and the simplified user intent.\n\n**`hidden_events` vs `exclusions` — hiding vs disqualifying.** `hidden_events` removes events from display but they still affect path structure and counts. `exclusions` disqualifies users who performed those events entirely — a much stronger operation. Use `hidden_events` for decluttering (e.g., ubiquitous page views); use `exclusions` for removing tainted journeys (e.g., users who churned mid-flow).\n\n**Three modes reveal different stories.** `sankey` shows aggregate flow structure and bottlenecks (where do most users go?). `paths` shows exact user journeys in sequence (what are the top 5 complete paths?). `tree` shows branching decision points (where do users diverge?). Use all three on the same data to build a complete picture.\n\n#### User Profile Analysis: Modes, Aggregates, and Distribution Shape\n\n**`mode` is the most critical user query parameter.** `'profiles'` returns individual user records (one row per user). `'aggregate'` returns a single statistic. These are fundamentally different operations — profiles is a data extraction, aggregate is a calculation. Aggregate is also dramatically faster (single API call vs paginated fetching).\n\n**Sweep aggregate functions to understand distribution shape** before building expensive profile queries. `count` tells you \"how many.\" `extremes` reveals range (min/max). `percentile` at 50 gives median. `numeric_summary` gives mean, variance, and sum-of-squares:\n\n```python\n# Sweep aggregate functions to understand a property's distribution\nprop = 'lifetime_value'  # use a real numeric profile property\nfor agg in ['count', 'extremes', 'percentile', 'numeric_summary']:\n    kwargs = {'mode': 'aggregate', 'aggregate': agg}\n    if agg != 'count':\n        kwargs['aggregate_property'] = prop\n    if agg == 'percentile':\n        kwargs['percentile'] = 50  # median\n    result = ws.query_user(**kwargs)\n    print(f\"{agg:>16}: {result.aggregate_data}\")\n# count = population size, extremes = range, percentile@50 = median,\n# numeric_summary = full distribution stats\n# If mean (from numeric_summary) >> median (from percentile), distribution is right-skewed\n\n# Point-in-time comparison with as_of\ntoday_count = ws.query_user(mode='aggregate', aggregate='count',\n    where=Filter.equals('plan', 'premium'))\npast_count = ws.query_user(mode='aggregate', aggregate='count',\n    where=Filter.equals('plan', 'premium'), as_of='2025-01-01')\nprint(f\"Premium users: {past_count.value} (Jan 1) → {today_count.value} (today)\")\n```\n\n**Prefer medians over averages** — same principle as funnels and Insights. `aggregate='percentile', percentile=50` gives median; `numeric_summary` gives mean. If they diverge significantly, the distribution is skewed and the mean is misleading.\n\n**`as_of` enables temporal analysis** — query profiles as they existed at a past date. Compare population states over time: \"how many premium users existed on Jan 1 vs today?\" Without `as_of`, you always see current state, making growth and churn invisible.\n\n**Inline `CohortDefinition` vs saved cohorts.** Inline cohorts (`cohort=CohortDefinition.all_of(...)`) let you define complex behavioral segments on-the-fly without roundtripping to save/delete in Mixpanel. Much faster iteration for exploratory analysis. Use saved cohorts for production dashboards and monitoring.\n\n#### Analytical Building Blocks: Custom Properties, Cohorts, and Frequency\n\nRaw data is rarely analysis-ready. These three tools transform raw events and properties into analytically useful dimensions, populations, and segments. Recognize when to reach for each — they compose with every query engine.\n\n**Inline Custom Properties — transform data at query time.** When property values are messy, need bucketing, or you need to derive new dimensions, create an `InlineCustomProperty` rather than querying raw values. Key patterns:\n\n- **Bucketing continuous values** for breakdowns (revenue → Low/Medium/High)\n- **Cleaning messy strings** with IFS/REGEX_EXTRACT (campaign names, UTM parameters)\n- **Deriving new dimensions** from arithmetic or date functions (profit margin, days since signup)\n- **Fallback chains** across multiple properties (display_name → username → \"unknown\")\n\n```python\nfrom mixpanel_headless import InlineCustomProperty, PropertyInput, GroupBy, Filter, Metric\n\n# Bucket revenue into tiers for breakdown\nrevenue_tier = InlineCustomProperty(\n    formula='IFS(A < 50, \"Low\", A < 200, \"Medium\", TRUE, \"High\")',\n    inputs={\"A\": PropertyInput(\"revenue\", type=\"number\")},\n    property_type=\"string\",\n)\nresult = ws.query(\"Purchase\", group_by=GroupBy(property=revenue_tier), last=30, mode='total')\n\n# Derive profit margin for aggregation\nmargin = InlineCustomProperty.numeric(\"(A - B) / A * 100\", A=\"revenue\", B=\"cost\")\nresult = ws.query(Metric(\"Purchase\", math=\"average\", property=margin), last=30)\n\n# Clean messy strings for segmentation\ndomain = InlineCustomProperty(\n    formula='REGEX_EXTRACT(A, \"@(.+)$\")',\n    inputs={\"A\": PropertyInput(\"email\", type=\"string\")},\n    property_type=\"string\",\n)\nresult = ws.query(\"Signup\", group_by=GroupBy(property=domain), last=30, mode='total')\n```\n\nUse `InlineCustomProperty` for ad-hoc exploration. When a formula proves valuable, persist it with `ws.create_custom_property()` and reference it via `CustomPropertyRef(id)` across reports.\n\n**Inline Cohorts — define complex populations on-the-fly.** Every analytical question starts with \"among WHICH users?\" Simple property filters (`where=Filter.equals(...)`) answer \"users with attribute X.\" Inline cohorts answer harder questions: \"users who did X at least N times in the last D days AND did NOT do Y AND have property Z.\" Compose criteria with AND/OR logic:\n\n```python\nfrom mixpanel_headless import CohortDefinition, CohortCriteria, CohortBreakdown, CohortMetric\n\n# \"Power users\": purchased 5+ times in 30 days, never contacted support\npower_users = CohortDefinition.all_of(\n    CohortCriteria.did_event(\"Purchase\", at_least=5, within_days=30),\n    CohortCriteria.did_not_do_event(\"Support Ticket\", within_days=90),\n)\n\n# Use inline cohort as a breakdown — no need to save first\nresult = ws.query(\"Login\", group_by=CohortBreakdown(power_users, \"Power Users\"), last=30)\n\n# Use inline cohort as a filter in user queries\nresult = ws.query_user(cohort=power_users, mode='aggregate', aggregate='count')\n\n# Track saved cohort size over time alongside event metrics\nresult = ws.query(\n    [Metric(\"Login\", math=\"unique\"), CohortMetric(saved_cohort_id, \"Power Users\")],\n    formula=\"(B / A) * 100\", formula_label=\"% Power Users Active\", last=90,\n)\n```\n\n**Frequency Breakdown/Filter — segment by behavioral intensity.** `FrequencyBreakdown` answers \"how do users who did X once differ from users who did X ten times?\" `FrequencyFilter` restricts queries to users meeting a frequency threshold. These bridge \"what users did\" with \"who users are\":\n\n```python\nfrom mixpanel_headless import FrequencyBreakdown, FrequencyFilter\n\n# Break down login behavior by purchase frequency\nresult = ws.query(\"Login\", math='unique',\n    group_by=FrequencyBreakdown(\"Purchase\", bucket_size=3, bucket_min=0, bucket_max=15),\n    last=30, mode='total')\n# Reveals: do frequent purchasers also log in more?\n\n# Filter to users who purchased 3+ times in 30 days, then analyze their flow\nresult = ws.query_flow(\"Login\", forward=3,\n    where=FrequencyFilter(\"Purchase\", value=3, date_range_value=30, date_range_unit=\"day\"),\n    last=30)\n# Reveals: what do repeat purchasers do after logging in?\n```\n\n**When to reach for each:**\n- Property values are messy or need derivation → **Custom Property**\n- Population requires behavioral criteria (did X, didn't do Y, frequency thresholds) → **Inline Cohort**\n- You need to segment by event frequency (how often, not just whether) → **FrequencyBreakdown/Filter**\n- You need to compare in-cohort vs out-of-cohort behavior → **CohortBreakdown** with `include_negated=True`\n- You need to track a segment's size as a time series → **CohortMetric** (saved cohorts only)\n\n### Legacy Queries & Counts\n\nThese use older APIs. Prefer the typed query methods above when possible.\n\n```python\ndef segmentation(self, event: str, *, from_date: str, to_date: str, on: str | None = None, unit: Literal['day', 'week', 'month'] = 'day', where: str | None = None) -> SegmentationResult: ...\ndef funnel(self, funnel_id: int, *, from_date: str, to_date: str, unit: str | None = None, on: str | None = None) -> FunnelResult: ...\ndef retention(self, *, born_event: str, return_event: str, from_date: str, to_date: str, born_where: str | None = None, return_where: str | None = None, interval: int = 1, interval_count: int = 10, unit: Literal['day', 'week', 'month'] = 'day') -> RetentionResult: ...\ndef event_counts(self, events: list[str], *, from_date: str, to_date: str, type: Literal['general', 'unique', 'average'] = 'general', unit: Literal['day', 'week', 'month'] = 'day') -> EventCountsResult: ...\ndef property_counts(self, event: str, property_name: str, *, from_date: str, to_date: str, type: Literal['general', 'unique', 'average'] = 'general', unit: Literal['day', 'week', 'month'] = 'day', values: list[str] | None = None, limit: int | None = None) -> PropertyCountsResult: ...\ndef frequency(self, *, from_date: str, to_date: str, unit: Literal['day', 'week', 'month'] = 'day', addiction_unit: Literal['hour', 'day'] = 'hour', event: str | None = None, where: str | None = None) -> FrequencyResult: ...\ndef activity_feed(self, distinct_ids: list[str], *, from_date: str | None = None, to_date: str | None = None) -> ActivityFeedResult: ...\ndef query_saved_report(self, bookmark_id: int, *, bookmark_type: Literal['insights', 'funnels', 'retention', 'flows'] = 'insights', from_date: str | None = None, to_date: str | None = None) -> SavedReportResult: ...\ndef query_saved_flows(self, bookmark_id: int) -> FlowsResult: ...\ndef segmentation_numeric(self, event: str, *, from_date: str, to_date: str, on: str, unit: Literal['hour', 'day'] = 'day', where: str | None = None, type: Literal['general', 'unique', 'average'] = 'general') -> NumericBucketResult: ...\ndef segmentation_sum(self, event: str, *, from_date: str, to_date: str, on: str, unit: Literal['hour', 'day'] = 'day', where: str | None = None) -> NumericSumResult: ...\ndef segmentation_average(self, event: str, *, from_date: str, to_date: str, on: str, unit: Literal['hour', 'day'] = 'day', where: str | None = None) -> NumericAverageResult: ...\n```\n\n### Entity CRUD (App API)\n\nAll entity methods require a workspace ID. Use `python3 $SKILL_DIR/scripts/help.py Workspace.<method>` for full signatures and parameter types.\nUser Guide: `WebFetch(url=\"https://mixpanel.github.io/mixpanel-headless/guide/entity-management/index.md\")`\n\n#### Dashboard (→ `Dashboard`)\n\n`list_dashboards`, `create_dashboard`, `get_dashboard`, `update_dashboard`, `delete_dashboard`, `bulk_delete_dashboards`, `favorite_dashboard`, `unfavorite_dashboard`, `pin_dashboard`, `unpin_dashboard`, `add_report_to_dashboard`, `remove_report_from_dashboard`, `update_text_card`, `update_report_link`\n\n**Blueprints:** `list_blueprint_templates` → `list[BlueprintTemplate]`, `create_blueprint`, `get_blueprint_config`, `update_blueprint_cohorts`, `finalize_blueprint`, `create_rca_dashboard`\n\n**Helpers:** `get_bookmark_dashboard_ids` → `list[int]`, `get_dashboard_erf` → `dict`\n\n#### Bookmark / Report (→ `Bookmark`)\n\n`list_bookmarks_v2`, `create_bookmark`, `get_bookmark`, `update_bookmark`, `delete_bookmark`, `bulk_delete_bookmarks`, `bulk_update_bookmarks`, `bookmark_linked_dashboard_ids` → `list[int]`, `get_bookmark_history` → `BookmarkHistoryResponse`\n\n#### Cohort (→ `Cohort`)\n\n`list_cohorts_full`, `get_cohort`, `create_cohort`, `update_cohort`, `delete_cohort`, `bulk_delete_cohorts`, `bulk_update_cohorts`\n\n#### Feature Flag (→ `FeatureFlag`)\n\n`list_feature_flags`, `create_feature_flag`, `get_feature_flag`, `update_feature_flag`, `delete_feature_flag`, `archive_feature_flag`, `restore_feature_flag`, `duplicate_feature_flag`, `set_flag_test_users`, `get_flag_history` → `FlagHistoryResponse`, `get_flag_limits` → `FlagLimitsResponse`\n\n#### Experiment (→ `Experiment`)\n\n`list_experiments`, `create_experiment`, `get_experiment`, `update_experiment`, `delete_experiment`, `launch_experiment`, `conclude_experiment`, `decide_experiment`, `archive_experiment`, `restore_experiment`, `duplicate_experiment`, `list_erf_experiments` → `list[dict]`\n\n#### Alert (→ `CustomAlert`)\n\n`list_alerts`, `create_alert`, `get_alert`, `update_alert`, `delete_alert`, `bulk_delete_alerts`, `get_alert_count` → `AlertCount`, `get_alert_history` → `AlertHistoryResponse`, `test_alert`, `get_alert_screenshot_url`, `validate_alerts_for_bookmark`\n\n#### Annotation (→ `Annotation`)\n\n`list_annotations`, `create_annotation`, `get_annotation`, `update_annotation`, `delete_annotation`, `list_annotation_tags` → `list[AnnotationTag]`, `create_annotation_tag`\n\n#### Webhook (→ `ProjectWebhook`)\n\n`list_webhooks`, `create_webhook`, `update_webhook`, `delete_webhook`, `test_webhook`\n\n#### Lexicon & Data Governance\n\n**Event/Property Definitions:** `get_event_definitions`, `update_event_definition`, `delete_event_definition`, `bulk_update_event_definitions`, `get_property_definitions`, `update_property_definition`, `bulk_update_property_definitions`, `export_lexicon`, `get_event_history`, `get_property_history`\n\n**Tags:** `list_lexicon_tags`, `create_lexicon_tag`, `update_lexicon_tag`, `delete_lexicon_tag`\n\n**Drop Filters:** `list_drop_filters`, `create_drop_filter`, `update_drop_filter`, `delete_drop_filter`, `get_drop_filter_limits`\n\n**Custom Properties:** `list_custom_properties`, `create_custom_property`, `get_custom_property`, `update_custom_property`, `delete_custom_property`, `validate_custom_property`\n\n**Custom Events:** `list_custom_events`, `update_custom_event`, `delete_custom_event`\n\n**Lookup Tables:** `list_lookup_tables`, `upload_lookup_table`, `download_lookup_table`, `update_lookup_table`, `delete_lookup_tables`\n\n**Schema Registry:** `list_schema_registry`, `create_schema`, `update_schema`, `create_schemas_bulk`, `update_schemas_bulk`, `delete_schemas`\n\n**Schema Enforcement:** `get_schema_enforcement`, `init_schema_enforcement`, `update_schema_enforcement`, `replace_schema_enforcement`, `delete_schema_enforcement`\n\n**Audit & Monitoring:** `run_audit`, `run_audit_events_only`, `list_data_volume_anomalies`, `update_anomaly`, `bulk_update_anomalies`\n\n**Data Deletion:** `list_deletion_requests`, `create_deletion_request`, `cancel_deletion_request`, `preview_deletion_filters`\n\n**Other:** `get_tracking_metadata`\n\n### Business Context\n\nRead and write the markdown documentation that grounds AI assistants in your organization's structure and goals, exposed as a typed Python API.\n\nTwo scopes — `level=\"organization\"` (shared across the whole org) and `level=\"project\"` (per-project). 50,000-character cap enforced **client-side before any HTTP call** so oversize input fails fast. Org-level operations auto-resolve `organization_id` from the cached `/me` response; pass `organization_id=N` to override.\n\nRun `python3 $SKILL_DIR/scripts/help.py search business_context` to see all four methods, two types, and one exception.\n\n```python\nfrom mixpanel_headless import BUSINESS_CONTEXT_MAX_CHARS  # 50_000\n\n# Read\nproject_ctx = ws.get_business_context(level=\"project\")\norg_ctx = ws.get_business_context(level=\"organization\")  # auto-resolves org_id\nexplicit = ws.get_business_context(level=\"organization\", organization_id=42)\n\n# Read both at once (single round-trip via /business-context/chain)\nchain = ws.get_business_context_chain()\nprint(chain.organization.content)\nprint(chain.project.content)\n\n# Write (full-replace; pass \"\" to clear, or use clear_business_context())\nws.set_business_context(\"# About Acme\\n…\", level=\"project\")\nws.set_business_context(\"# Org-wide standards\", level=\"organization\")\nws.clear_business_context(level=\"project\")\n\n# All return BusinessContext with: level, content, organization_id, project_id\n# Plus convenience .is_empty and .character_count properties (Python only)\nprint(f\"{project_ctx.character_count}/{BUSINESS_CONTEXT_MAX_CHARS} chars; \"\n      f\"empty={project_ctx.is_empty}\")\n```\n\n**When to reach for this:**\n\n- User asks \"what's the business context for this project/org?\" → `get_business_context_chain()`\n- User wants to version-control project context as a `.md` file → `ws.set_business_context(Path(\"ctx.md\").read_text(), level=\"project\")` in CI\n- User asks to \"audit which projects have AI context configured\" → iterate `ws.projects()` + `ws.use(project=...)` + `ws.get_business_context(level=\"project\")` and check `.is_empty`\n- User asks to seed a new project from the org default → `chain = ws.get_business_context_chain(); ws.set_business_context(chain.organization.content, level=\"project\")`\n\n**Permissions:** project-scope reads need any project access; project-scope writes need `edit_project_info` on the project. Org-scope writes need `edit_project_info` at the org level (typically OAuth, not service account). The `BusinessContextValidationError` exception is raised client-side BEFORE any HTTP call when content exceeds 50,000 chars, so use it to detect oversize input without burning a round-trip.\n\nUser Guide: `WebFetch(url=\"https://mixpanel.github.io/mixpanel-headless/guide/business-context/index.md\")`\n\n## Key Types\n\nRun `python3 $SKILL_DIR/scripts/help.py types` for the full list of all types. Use `help.py <TypeName>` for fields, constructors, and enum values.\nFull reference: `WebFetch(url=\"https://mixpanel.github.io/mixpanel-headless/api/types/index.md\")`\n\n| Type | Purpose |\n|------|---------|\n| `Filter` | Property filter conditions (`.equals()`, `.contains()`, `.in_cohort()`, etc.) |\n| `GroupBy` | Property breakdown with optional bucketing |\n| `Formula` | Calculated metric expression referencing events by position (A, B, C...) |\n| `Metric` | Event with per-event math/aggregation settings |\n| `CohortMetric` | Track cohort size over time as an event metric |\n| `FunnelStep` | Funnel step with per-step filters, labels, ordering |\n| `Exclusion` | Event to exclude between funnel steps |\n| `HoldingConstant` | Property to hold constant across funnel steps |\n| `RetentionEvent` | Retention event with per-event filters |\n| `FlowStep` | Flow anchor event with per-step forward/reverse configuration |\n| `TimeComparison` | Period-over-period comparison (`.relative(\"month\")`, `.absolute_start(...)`) |\n| `FrequencyBreakdown` | Break down by how often users performed an event |\n| `FrequencyFilter` | Filter by how often users performed an event |\n| `CohortBreakdown` | Break down results by cohort membership |\n| `CohortDefinition` | Inline cohort definition for user queries |\n| `CohortCriteria` | Atomic condition for cohort membership |\n| `CustomPropertyRef` | Reference to a persisted custom property by ID |\n| `InlineCustomProperty` | Ephemeral computed property defined by formula |\n\n**Aggregation enums** (use `help.py <EnumName>` to see all values):\n\n| Enum | Used by | Common values |\n|------|---------|---------------|\n| `MathType` | `query()` | total, unique, dau, average, sum, min, max, percentile, sessions |\n| `FunnelMathType` | `query_funnel()` | conversion_rate_unique, conversion_rate_total, average, median |\n| `RetentionMathType` | `query_retention()` | retention_rate, retention_count |\n\n## Statistical Analysis — numpy, scipy\n\nAll query results produce pandas DataFrames, which integrate directly with numpy and scipy:\n\n```python\nimport numpy as np\nfrom scipy import stats\n\n# Compare two segments\na = result.df[result.df[\"platform\"] == \"iOS\"][\"count\"]\nb = result.df[result.df[\"platform\"] == \"Android\"][\"count\"]\nt_stat, p_value = stats.ttest_ind(a, b)\ncohens_d = (a.mean() - b.mean()) / np.sqrt((a.std()**2 + b.std()**2) / 2)\n\n# Useful scipy.stats tests: ttest_ind, mannwhitneyu, chi2_contingency, pearsonr, spearmanr\n# Useful numpy: np.percentile, np.corrcoef, np.polyfit (trend lines)\n```\n\n## Visualization — matplotlib, seaborn\n\nSave charts to files for the user. Always use a non-interactive backend:\n\n```python\nimport matplotlib\nmatplotlib.use(\"Agg\")\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfig, ax = plt.subplots(figsize=(10, 5))\nresult.df.plot(x=\"date\", y=\"count\", ax=ax)\nax.set_title(\"Daily Logins\")\nfig.savefig(\"chart.png\", dpi=150, bbox_inches=\"tight\")\nplt.close(fig)\n\n# seaborn: sns.lineplot, sns.barplot, sns.heatmap (for retention matrices)\n# Multi-panel: fig, axes = plt.subplots(2, 2) for dashboard-style layouts\n```\n\n## Exceptions\n\nFull reference: `WebFetch(url=\"https://mixpanel.github.io/mixpanel-headless/api/exceptions/index.md\")`\n\n| Exception | When |\n|-----------|------|\n| `MixpanelHeadlessError` | Base for all errors |\n| `ConfigError` | No credentials resolved |\n| `AccountNotFoundError` | Named account doesn't exist |\n| `AuthenticationError` | Invalid credentials (401) |\n| `QueryError` | Invalid query parameters (400) |\n| `BookmarkValidationError` | Params failed validation |\n| `RateLimitError` | Rate limit exceeded (429) |\n| `ServerError` | Mixpanel server error (5xx) |\n| `WorkspaceScopeError` | Workspace resolution error (also raised when org_id can't be auto-resolved for `level=\"organization\"` business-context calls) |\n| `DateRangeTooLargeError` | Date range exceeds API maximum |\n| `OAuthError` | OAuth flow error |\n| `BusinessContextValidationError` | Business context content exceeds 50,000 chars (client-side, before HTTP) |\n"
}

SHA-256: 6b2a938a7beb815bdfd858ef4c3a05a11c5f6a66a1a00340dc7574e8746855a7