← Ad SuperpowersCONTENT HISTORYWHAT CHANGED · RULE-BASED ANALYSIS
Update to Ad Superpowers
Snapshot Sep 30, 2026 · 23:07 UTC · version 2.3.0
Collection source: not recorded for this historical snapshot.
First saved snapshot
No earlier snapshot is available to establish a change.
Compare saved observations
Download comparison JSONFull technical diff · 0 changed fields
Full snapshot data
{
"name": "experiment-design-framework",
"description": "This skill should be used when the user asks to \"design an A/B test for ads\", \"calculate sample size for an experiment\", \"measure incrementality\", \"run a geo lift study\", or mentions \"statistical significance\", \"holdout test\", or \"test duration\". Do NOT use for: creative fatigue analysis (use creative-fatigue-analyzer), general campaign performance review (use platform-specific troubleshooters), or audience strategy (use buyer-persona-framework).",
"included_files": [],
"skill_md_contents": "---\nname: experiment-design-framework\ndescription: \"This skill should be used when the user asks to \\\"design an A/B test for ads\\\", \\\"calculate sample size for an experiment\\\", \\\"measure incrementality\\\", \\\"run a geo lift study\\\", or mentions \\\"statistical significance\\\", \\\"holdout test\\\", or \\\"test duration\\\". Do NOT use for: creative fatigue analysis (use creative-fatigue-analyzer), general campaign performance review (use platform-specific troubleshooters), or audience strategy (use buyer-persona-framework).\"\n---\n\n# Advertising Experiment Design Framework\n\n## Purpose\n\nProvide a rigorous, practical framework for designing and interpreting advertising experiments. Move from \"I think this works\" to \"I know this works, and here's the data.\" Most ad optimization is observational — experiments let you prove causation.\n\n## When to Use This Skill\n\nInvoke when user mentions:\n- **A/B testing:** \"How do I A/B test my ads?\"\n- **Statistical significance:** \"Is this result significant?\"\n- **Sample size:** \"How many conversions do I need?\"\n- **Test duration:** \"How long should I run this test?\"\n- **Incrementality:** \"Is this channel actually driving sales?\"\n- **Holdout test:** \"What would happen if I turned off this campaign?\"\n- **Geo lift:** \"How do I test a campaign's true impact?\"\n- **Multi-variate:** \"Can I test multiple things at once?\"\n- **Learning phase:** \"How do I test without wasting budget?\"\n\n---\n\n## Part 1: Experiment Types\n\n### Overview Matrix\n\n| Experiment Type | Complexity | Cost | Statistical Rigor | Best For |\n|----------------|-----------|------|-------------------|----------|\n| **A/B Test (Split Test)** | Low | Low | Medium-High | Creative, copy, landing pages |\n| **Multi-Variate Test (MVT)** | Medium | Medium | Medium | Multiple creative elements simultaneously |\n| **Holdout Test** | Low | Low-Medium | High | Measuring incrementality of a campaign |\n| **Geo Lift Test** | High | High | Highest | Measuring true channel contribution |\n| **Pre/Post Test** | Low | Low | Low | Rough directional signal only |\n| **Conversion Lift (Meta/Google)** | Medium | Medium | High | Platform-provided incrementality |\n\n### When to Use Each Type\n\n```\nQUESTION: What are you trying to learn?\n\n├── \"Which creative/copy/CTA works better?\"\n│ └── A/B Test (or MVT if testing multiple elements)\n│\n├── \"Is this campaign actually driving incremental sales?\"\n│ └── Holdout Test (simplest) or Conversion Lift Study\n│\n├── \"What's the true ROI of this channel?\"\n│ └── Geo Lift Test (gold standard) or Holdout Test\n│\n├── \"Should I change my bid strategy?\"\n│ └── A/B Test with Campaign Budget Optimization\n│ (run both strategies simultaneously, same audience split)\n│\n├── \"Which audience performs better?\"\n│ └── A/B Test with audience splitting\n│ (Meta: split test feature; Google: experiments)\n│\n└── \"What's the best combination of headline + image + CTA?\"\n └── Multi-Variate Test (need high traffic volume)\n```\n\n---\n\n## Part 2: A/B Test Design\n\n### The 5 Requirements of a Valid A/B Test\n\n| Requirement | What It Means | Common Violation |\n|------------|--------------|-----------------|\n| **1. Single variable** | Change ONE thing between variants | Testing new image AND new copy simultaneously |\n| **2. Random assignment** | Audience randomly split, not self-selected | Showing variant A to one audience, B to another |\n| **3. Sufficient sample size** | Enough data to detect the expected effect | Declaring a winner after 50 conversions |\n| **4. Adequate duration** | Run long enough to capture weekly patterns | Stopping after 3 days because one variant \"looks better\" |\n| **5. Pre-defined success metric** | Decide what \"winning\" means before launch | Switching metric to \"engagement\" when conversion results are flat |\n\n### What to Test (Testing Hierarchy)\n\n**Test big bets first, then refine.** Impact ranking:\n\n| Priority | What to Test | Expected Impact | Minimum Budget |\n|----------|-------------|----------------|----------------|\n| **1 (Highest)** | Offer/Pricing | 50-200% conversion lift | €500 |\n| **2** | Audience/Targeting | 30-100% efficiency gain | €1,000 |\n| **3** | Landing page | 20-80% conversion lift | €500 |\n| **4** | Ad format (video vs image vs carousel) | 20-50% CTR change | €500 |\n| **5** | Creative concept (visual theme) | 15-40% CTR change | €300 |\n| **6** | Headline/Copy | 10-25% CTR change | €300 |\n| **7** | CTA button | 5-15% CTR change | €200 |\n| **8** | Color/font/minor design | 2-10% CTR change | €200 |\n| **9 (Lowest)** | Bid strategy | 5-20% CPA change | €1,000 |\n\n**Rule of thumb:** Don't A/B test CTA button colors when you haven't tested whether video outperforms images.\n\n### A/B Test Setup by Platform\n\n**Meta Ads:**\n- Use the built-in A/B Test feature (Experiments tab)\n- Meta handles randomization and statistical analysis\n- Choose split: audience (default), placement, or delivery optimization\n- Run at ad set level for audience tests, ad level for creative tests\n- Monitor with `meta_get_insights` using ad-level breakdowns\n\n**Google Ads:**\n- Use Campaign Experiments for bid strategy / targeting tests\n- Use Ad Variations for copy / headline tests\n- Experiments split traffic automatically (customizable %)\n- RSA testing: pin different headlines to positions and compare\n- Monitor with `google_ads_run_gaql`:\n\n```sql\nSELECT\n ad_group_ad.ad.id,\n ad_group_ad.ad.name,\n metrics.impressions,\n metrics.clicks,\n metrics.conversions,\n metrics.cost_micros,\n metrics.conversions_value\nFROM ad_group_ad\nWHERE campaign.id = {campaign_id}\n AND segments.date DURING LAST_30_DAYS\nORDER BY metrics.conversions DESC\n```\n\n---\n\n## Part 3: Sample Size & Duration\n\n### Sample Size Calculation\n\n**The formula (simplified):**\n```\nRequired conversions per variant ≈ 16 × (1/MDE²)\n\nWhere MDE = Minimum Detectable Effect (as decimal)\n```\n\n| MDE (Minimum Effect You Want to Detect) | Conversions Needed Per Variant | Total Conversions (2 variants) |\n|----------------------------------------|-------------------------------|-------------------------------|\n| 50% improvement | ~64 | ~128 |\n| 30% improvement | ~178 | ~356 |\n| 20% improvement | ~400 | ~800 |\n| 15% improvement | ~711 | ~1,422 |\n| 10% improvement | ~1,600 | ~3,200 |\n| 5% improvement | ~6,400 | ~12,800 |\n\n**Key insight:** The smaller the improvement you want to detect, the more data you need. Most ad tests should target detecting a 20-30% improvement — smaller effects usually aren't worth optimizing for.\n\n### Duration Guidelines\n\n| Factor | Minimum | Recommended | Maximum |\n|--------|---------|------------|---------|\n| **Calendar time** | 7 days | 14-21 days | 28 days |\n| **Full business cycles** | 1 week | 2 weeks | 4 weeks |\n| **Conversions per variant** | 50 (directional) | 100+ (reliable) | No max |\n| **Confidence level** | 90% (directional) | 95% (standard) | 99% (high-stakes) |\n\n### Duration Calculator\n\n```\nEstimated duration = Required conversions per variant / Daily conversion rate per variant\n\nExample:\n Current daily conversions: 20/day (total campaign)\n Split 50/50: 10/day per variant\n Need 400 conversions per variant (20% MDE)\n Duration = 400 / 10 = 40 days\n\n That's too long. Options:\n a) Accept higher MDE (30% → 178 conversions → 18 days) ✓\n b) Increase budget to get more daily conversions ✓\n c) Use a higher-volume metric (clicks instead of purchases) ⚠️ (less meaningful)\n```\n\n### When You Don't Have Enough Volume\n\n| Daily Conversions | Recommended Approach |\n|------------------|---------------------|\n| > 50/day | Full A/B test, 95% significance, 2-week minimum |\n| 20-50/day | A/B test, 90% significance, 3-week minimum |\n| 5-20/day | A/B test, 90% significance, accept higher MDE (30%+) |\n| 1-5/day | Sequential testing (run A for 2 weeks, then B for 2 weeks) |\n| < 1/day | Don't A/B test conversions. Test higher-funnel metric (CTR, Add to Cart) |\n\n---\n\n## Part 4: Statistical Significance\n\n### What It Means\n\n```\n\"95% statistical significance\" means:\n\nThere is a ≤5% probability that the observed difference between\nvariants occurred by random chance alone.\n\nIt does NOT mean:\n- \"95% chance that variant B is better\" (common misinterpretation)\n- \"Variant B will always outperform A by this margin\"\n- \"The test is 95% accurate\"\n```\n\n### Interpreting Results\n\n| Scenario | Significance | Confidence Interval | Interpretation | Action |\n|----------|-------------|-------------------|---------------|--------|\n| A: 2.1% CVR, B: 2.8% CVR | p = 0.02 (sig.) | B is +20% to +50% better | Clear winner | Implement B |\n| A: 2.1% CVR, B: 2.4% CVR | p = 0.15 (not sig.) | B is -5% to +30% better | Inconclusive | Need more data or accept ambiguity |\n| A: 2.1% CVR, B: 2.2% CVR | p = 0.45 (not sig.) | B is -12% to +18% better | No difference | Either variant works, choose based on other factors |\n| A: 2.1% CVR, B: 1.6% CVR | p = 0.01 (sig.) | B is -15% to -35% worse | Clear loser | Do NOT implement B |\n\n### Common Statistical Mistakes\n\n| Mistake | Why It's Wrong | What to Do Instead |\n|---------|---------------|-------------------|\n| **Peeking and stopping early** | Significance fluctuates early; stopping on a \"good day\" inflates false positives | Set duration upfront, only check at end (or use sequential testing) |\n| **Running until significant** | If you keep running, random fluctuations will eventually reach 95% | Pre-define sample size and duration |\n| **Ignoring negative results** | \"The test didn't work, let's try something else\" loses the learning | Document learnings: what does this tell you about your audience? |\n| **Testing too many variants** | 5 variants = 10 pairwise comparisons = much higher false positive rate | Maximum 3-4 variants; apply Bonferroni correction for multiple comparisons |\n| **Wrong success metric** | Optimizing for clicks when you care about purchases | Define primary metric before launch, ideally closest to revenue |\n| **Novelty effect** | New variant gets initial engagement boost that fades | Run test for 2+ weeks to see past novelty |\n| **Segment cherry-picking** | \"It didn't win overall, but it won with women 25-34!\" | Only analyze pre-defined segments, not post-hoc discoveries |\n\n---\n\n## Part 5: Incrementality Testing\n\n### What Is Incrementality?\n\n```\nIncrementality = Sales with ads - Sales that would have happened anyway (without ads)\n\nExample:\n With retargeting campaign: 100 purchases/week\n Without retargeting (holdout): 75 purchases/week\n Incremental sales: 25/week\n Incrementality rate: 25% (only 25 of 100 sales were truly driven by the ads)\n True ROAS: Reported ROAS × 0.25\n```\n\n### Holdout Test Design\n\n**The simplest incrementality test:**\n\n1. Take your retargeting audience\n2. Randomly split: 90% see ads (treatment), 10% see no ads (holdout)\n3. Run for 2-4 weeks\n4. Compare purchase rate: treatment group vs holdout group\n5. The difference = incremental impact\n\n**Meta implementation:**\n- Create a Conversion Lift study in Experiments\n- Meta handles the random split and measurement\n- Minimum spend: ~€5,000 over the test period\n- Results in 2-4 weeks\n\n**Google Ads implementation:**\n- Use Campaign Experiments with a holdout\n- Or use Google's Conversion Lift measurement (for larger accounts)\n\n### Typical Incrementality by Campaign Type\n\n| Campaign Type | Typical Incrementality | What This Means |\n|--------------|----------------------|----------------|\n| Brand search (own brand) | 10-30% | Most would have found you anyway |\n| Non-brand search (generic terms) | 40-70% | Capturing real intent |\n| Google Shopping | 30-60% | Price comparison, some would buy anyway |\n| Meta prospecting (broad) | 60-85% | Genuine new demand creation |\n| Meta retargeting (all visitors) | 15-35% | Many would have returned anyway |\n| Meta retargeting (cart abandoners) | 20-45% | Some would have completed purchase |\n| TikTok prospecting | 50-80% | Discovery-driven, high incrementality |\n| Google Display | 5-25% | View-through, often low incrementality |\n\n**Implication:** Channels that look \"efficient\" (low CPA, high ROAS) often have low incrementality. Brand search has the best ROAS but the lowest incrementality — those customers were already coming.\n\n### Geo Lift Test Design\n\n**The gold standard for channel-level incrementality:**\n\n1. **Select matched markets:** Pair cities/regions with similar demographics, market size, and baseline sales\n2. **Treatment vs control:** Run ads in treatment markets, no ads (or standard spend) in control\n3. **Duration:** 4-8 weeks minimum (longer for lower frequency categories)\n4. **Measurement:** Compare sales difference between treatment and control, adjusted for baseline\n\n**Market Matching Criteria:**\n\n| Factor | How to Match | Data Source |\n|--------|-------------|-------------|\n| Population size | Within 20% of each other | Census data |\n| Baseline sales | Similar weekly revenue | Shopify/GA4 |\n| Seasonality pattern | Same seasonal trends | Historical sales |\n| Competitive landscape | Similar competitors present | Market research |\n| Media landscape | Similar media costs | Platform data |\n\n**Example Design:**\n```\nTreatment markets: Amsterdam, Rotterdam, Utrecht\nControl markets: Den Haag, Eindhoven, Groningen\n\nRun Meta prospecting campaigns ONLY in treatment markets for 6 weeks.\nCompare sales growth in treatment vs control.\n\nTreatment sales growth: +18%\nControl sales growth: +3% (organic trend)\nIncremental lift: +15%\n```\n\n---\n\n## Part 6: Platform-Specific Testing Features\n\n### Meta Experiments\n\n| Feature | Use Case | Minimum Budget | Duration |\n|---------|---------|---------------|----------|\n| A/B Test | Creative, audience, placement | €500/variant | 7-28 days |\n| Conversion Lift | Campaign incrementality | €5,000+ | 14-28 days |\n| Brand Lift | Awareness/recall measurement | €10,000+ | 14-28 days |\n| Advantage+ Tests | Automated creative optimization | €1,000+ | 14 days |\n\n**Use `meta_get_insights` to monitor tests:**\n- Compare ad-level performance across variants\n- Track primary metric (conversions) AND secondary metrics (CTR, CPC)\n- Check `frequency` to ensure adequate exposure\n\n### Google Ads Experiments\n\n| Feature | Use Case | Setup |\n|---------|---------|-------|\n| Campaign Experiments | Bid strategy, targeting changes | Draft → Experiment (set traffic split) |\n| Ad Variations | Headline, description, URL changes | Find & Replace or custom rules |\n| Video Experiments | YouTube creative testing | Brand Lift integration |\n\n**Use `google_ads_run_gaql` to compare experiments:**\n```sql\nSELECT\n campaign.name,\n campaign.experiment_type,\n metrics.conversions,\n metrics.cost_micros,\n metrics.conversions_value,\n metrics.search_impression_share\nFROM campaign\nWHERE campaign.experiment_type != 'UNSPECIFIED'\n AND segments.date DURING LAST_30_DAYS\n```\n\n---\n\n## Part 7: Testing Calendar & Cadence\n\n### Recommended Testing Cadence\n\n| Business Size | Monthly Ad Spend | Tests per Month | Focus |\n|--------------|-----------------|----------------|-------|\n| Small (<€5K/mo) | €1-5K | 1 test | Creative or audience |\n| Medium (€5-25K/mo) | €5-25K | 2-3 tests | Creative, audience, bid strategy |\n| Large (€25-100K/mo) | €25-100K | 4-6 tests | Full program across platforms |\n| Enterprise (>€100K/mo) | €100K+ | 8-12 tests | Continuous optimization program |\n\n### Annual Testing Roadmap\n\n| Quarter | Testing Focus | Rationale |\n|---------|--------------|-----------|\n| **Q1 (Jan-Mar)** | Audience & targeting tests | Low CPMs, good for testing |\n| **Q2 (Apr-Jun)** | Creative format tests | Prepare winners for H2 |\n| **Q3 (Jul-Sep)** | Landing page & offer tests | Optimize conversion ahead of Q4 |\n| **Q4 (Oct-Dec)** | Minimize testing, run winners | Peak season, don't experiment with high stakes |\n\n### Test Documentation Template\n\nFor every experiment, document:\n\n```\nTest Name: [Descriptive name]\nHypothesis: \"If we [change X], then [metric Y] will [improve/decrease] by [Z%]\n because [reason/insight].\"\nPrimary Metric: [Conversion rate / ROAS / CPA / CTR]\nSecondary Metrics: [Other metrics to watch]\nPlatform: [Meta / Google / TikTok]\nAudience: [Who sees the test]\nVariants:\n - Control (A): [Description]\n - Treatment (B): [Description]\nExpected MDE: [%]\nRequired Sample: [N conversions per variant]\nEstimated Duration: [Days]\nBudget: [€ per variant]\nStart Date: [Date]\nEnd Date: [Date]\nResults: [Fill in after test completes]\nLearning: [What did we learn, regardless of outcome?]\nNext Action: [What do we do with this result?]\n```\n\n---\n\n## Part 8: Common Testing Mistakes & Fixes\n\n### The 10 Most Costly Mistakes\n\n| # | Mistake | Cost | Fix |\n|---|---------|------|-----|\n| 1 | **Not testing at all** | Missing 20-50% efficiency gains | Start with 1 test per month |\n| 2 | **Peeking daily and stopping early** | 30%+ false positive rate | Pre-commit to end date |\n| 3 | **Testing too many things at once** | Can't attribute results | One variable per test |\n| 4 | **Not enough budget per variant** | Inconclusive results, wasted time | Minimum €300-500/variant |\n| 5 | **Running during promotions** | Promotion effect masks test effect | Test during normal periods |\n| 6 | **Ignoring learning phase** | First 3-5 days data is unreliable | Exclude first 3 days from analysis |\n| 7 | **Winner takes all mentality** | Missing nuance in results | Check: does winner work for all segments? |\n| 8 | **Not documenting results** | Repeating failed tests, losing learnings | Maintain a test log (spreadsheet) |\n| 9 | **Testing during seasonality shifts** | Confounding variables | Test within stable periods |\n| 10 | **Copying competitor tests** | Different audience, different results | Test based on your own data and hypotheses |\n\n---\n\n## Part 9: Interpreting & Acting on Results\n\n### Result Interpretation Matrix\n\n| Statistical Sig. | Effect Size | Confidence Interval | Verdict | Action |\n|------------------|------------|-------------------|---------|--------|\n| p < 0.05 | Large (>20%) | Narrow, above zero | Strong winner | Implement immediately |\n| p < 0.05 | Small (5-10%) | Narrow, above zero | Marginal winner | Implement if no downside |\n| p = 0.05-0.10 | Any | Crosses zero | Inconclusive | Extend test or accept ambiguity |\n| p > 0.10 | Near zero | Wide, centered on zero | No difference | Either option works |\n| p < 0.05 | Negative | Below zero | Clear loser | Reject the change |\n\n### What To Do After Every Test\n\n```\n1. DOCUMENT the result (see template above)\n2. SHARE with team (even negative results have value)\n3. DECIDE: Implement, iterate, or reject\n4. GENERATE next hypothesis based on what you learned\n5. QUEUE the next test\n```\n\n### Building a Testing Culture\n\n**The compound effect of testing:**\n```\n1 test/month × 12 months = 12 experiments/year\nAssume 30% have a clear winner with 15% average improvement\n= ~4 winning changes per year\nCompound improvement: 1.15^4 = 1.75x (75% cumulative improvement)\n\nThis is why companies that test systematically outperform those that optimize by intuition.\n```\n\n---\n\n## Part 10: Advanced — Multi-Touch Incrementality\n\n### Beyond Single-Channel Testing\n\nWhen you run ads on Meta, Google, and TikTok simultaneously, turning off one channel affects the others. Advanced incrementality testing accounts for this:\n\n**Media Mix Modeling (MMM):**\n- Statistical model using 2+ years of spend and revenue data\n- Accounts for seasonality, promotions, organic trends\n- Outputs: marginal ROAS per channel, optimal budget allocation\n- Tools: Meta Robyn (open source), Google Meridian (open source), or commercial platforms\n- Best for: €50K+/month spend, 2+ years of data\n\n**Multi-Cell Lift Studies:**\n- Split audience into multiple cells: see all ads, see only Meta, see only Google, see no ads\n- Compare conversion rates across cells\n- Shows interaction effects between channels\n- Requires large audiences (100K+ per cell)\n\n**Practical Alternative for Smaller Budgets:**\n```\nMonth 1: Run all channels normally (baseline)\nMonth 2: Turn off Channel X (measure impact)\nMonth 3: Restore Channel X (confirm recovery)\n\nImpact of Channel X = Baseline revenue - Month 2 revenue\n(Adjusted for seasonality and organic trends)\n\nThis is crude but directional. Better than no incrementality data.\n```\n\n### MCP Tools for Experiment Monitoring\n\n**Pre-test baseline (all platforms):**\nUse `meta_get_insights`, `google_ads_run_gaql`, `tiktok_get_report` to establish 30-day baseline metrics before any experiment begins.\n\n**During test:**\nUse same tools weekly to monitor both variants. Watch for:\n- Sample size accumulation (on track?)\n- Any extreme outliers (data quality issue?)\n- External factors (competitor promotion, PR event?)\n\n**Post-test analysis:**\nPull final metrics from both variants using platform tools, calculate significance, document results.\n"
}SHA-256: c0e5a9ca9419705c437ffe936c19f8a2b3310f7e7ef2578227c9004c0fa9c2bc