← Files Sparkore Market ResearchARCHIVED FILE
skills/mobile-game-market-research/references/validation-playbook.md
4.13 KB · Oct 2, 2026 · 00:35 UTC
# Validation and Commitment Playbook Use before planning an experiment, interpreting evidence from tests, or recommending resource commitments. ## Match test to uncertainty | Test | Can inform | Does not establish by itself | |---|---|---| | Creative / message test | Relative attention and acquisition response in the tested audience/channel | Gameplay quality, retention, LTV or profitable scaling | | Concept interview / survey | Comprehension, motivations, objections, language | Actual installation, repeat play or payment | | Playable prototype / observed playtest | Core experience, comprehension, friction and desire to replay in the sample | Population retention, production throughput or scalable UA | | Store / landing-page test | Response to positioning and observed funnel steps | Retention or payment without actual product cohorts | | Instrumented product cohort / soft launch | Cohort engagement and monetization in tested conditions | Transfer to other markets, channels, seasons or spend levels | Select the cheapest test that can change the next decision. These are options, not a mandatory sequence. A creative success cannot override a failed gameplay hypothesis. ## Experiment specification Record: - Concept version, hypothesis, alternative explanation, and decision at stake. - Target users, geography, platform, recruitment/channel and relevant exclusions. - Treatment, comparator/control where feasible, and exact product/creative version. - Primary metric with event/denominator/cohort/time-window definition; secondary and guardrail metrics. - Thresholds for Pass / Fail, basis from comparable evidence or business requirements, and uncertainty range. - Sample requirement or precision goal, minimum observation period and rationale. Flag exploratory convenience samples. - Inconclusive conditions: insufficient sample, mixed outcomes, tracking errors, recruitment mismatch or confounds. - Budget, person-time and elapsed-time caps; dependencies and accountable owner. - Collection/analysis plan, stopping rule, and decision allowed for each result. Set thresholds before observing results. Do not invent universal CPI, retention, ROAS, or sample-size benchmarks. If no defensible threshold exists, run an explicitly exploratory pilot; define the decision it cannot yet support. Do not repeatedly peek and stop at a favorable fluctuation. ## Interpret results Verify tracking, sample composition, denominator, observation maturity and version consistency before comparing outcomes. Preserve raw results and distinguish observed values, estimates and projections. Show uncertainty; lack of significance is not proof of no effect. Diagnose recruitment or implementation problems before rejecting the underlying need. Report Pass, Fail, or Inconclusive against the original criteria. Explain conflicting metrics and changed assumptions. Any post-result threshold change must be explicit and requires fresh evidence before claiming validation. ## Commitment boundaries - Research further: named unresolved question and research cap. - Approve test: experiment specification with bounded resources. - Approve prototype: scope small enough to test the core thesis, with cost/time cap. - Greenlight production: critical evidence meets the predeclared requirements for that business model and commitment size, feasible staffing/budget is established, and downside/kill criteria are explicit. - Hold / Reject: state why and what new evidence or resource change could reopen the case. A plan to validate critical unknowns is insufficient for production Greenlight. A smaller staged commitment may be recommended if its boundary is explicit. Keep agent recommendation, owner authorization, and measured outcome separate. ## Business scenarios For production decisions, use downside/base/upside scenarios and sensitivity to acquisition cost, retention, monetization and production cost. Record metric definitions, applicable revenue streams, fees, UA, content/LiveOps cost, cash needs and payback horizon. Do not infer LTV from lifetime revenue divided by current downloads. Use cohort-consistent calculations and flag model assumptions. For ad-funded models, missing IAA evidence is a material gap.
SHA-256: f91080d1db7ee2f85f1c9c7989eac9f25ad143c1b879c92dda9bc2b8c24f8e07