← Files Sparkore Market ResearchARCHIVED FILE

skills/mobile-game-market-research/references/evidence-standards.md

5.45 KB · Oct 2, 2026 · 00:35 UTC

↓ Download file

# Evidence Standards

## Evidence hierarchy

Choose sources fit for the claim: publisher materials establish announced features, not neutral proof of commercial success; firsthand play establishes experience, not representative demand. Two articles repeating one dataset count as one evidence lineage.

Prefer the most direct available source:

1. Store listings, developer or publisher materials, product interfaces, patch notes, and firsthand play.
2. Platform, measurement-provider, or research-company reports with disclosed methodology.
3. AppMagic, Sensor Tower, ad-intelligence, and similar estimates, labeled as estimates.
4. Community discussions, reviews, creator coverage, and secondary reporting.

Community evidence is useful for player language, motivation, friction, and hypotheses. It is weak for representative prevalence unless sampled systematically.

## Snapshot discipline

For changeable facts, record:

- metric or observation;
- platform and geography;
- date range and granularity;
- snapshot date;
- source and durable link or source record;
- filters or methodology;
- observed, estimated, calculated, or inferred status.

Do not overwrite older comparable snapshots when trends matter.

## Epistemic layers

- **Fact / observation:** Directly supported by a source or firsthand observation.
- **Estimate:** Model-derived value from a measurement provider.
- **Calculation:** Derived transparently from stated inputs.
- **Interpretation:** Explanation of what evidence may mean.
- **Hypothesis:** Falsifiable explanation or prediction requiring validation.
- **Decision:** Chosen action based on evidence, risk, and constraints.

Avoid causal language when evidence shows only correlation.

## Insight record

```markdown
### Insight title

Observation:
Evidence:
Interpretation:
Alternative explanations:
Implication:
Confidence: High / Medium / Low
What would change our mind:
```

- **High:** Multiple independent, relevant sources converge and alternatives are weak.
- **Medium:** A repeatable pattern exists, but coverage or causality is incomplete.
- **Low:** A useful hypothesis supported by limited or indirect evidence.

## Bias controls

Actively test for:

- survivorship and recency bias;
- review or community selection bias;
- platform and geography mismatch;
- launch spike mistaken for sustained scale;
- downloads mistaken for engagement or profitability;
- revenue/download treated as precise LTV;
- feature presence mistaken for cause of success;
- fake or hybrid creatives mistaken for actual gameplay demand;
- absence of competitors mistaken for validated white space.

Include weak, failed, stagnant, and recent titles where possible. Compare cohorts by release period, platform, geography, and business model when material.

## Performance interpretation

Read curves, not only totals. Look for sustained baseline versus spike, decay after campaigns, repeatable seasonal lifts, divergence between downloads and revenue, geographic concentration, portfolio effects, and meaningful update cadence.

State when evidence cannot distinguish organic growth, paid acquisition, featuring, cross-promotion, or other causes.

## Player research

Preserve source, date, market or language, rating or sentiment, journey stage, topic, statement summary, and interpretation. Normalize around Why Install, Why Stay, Why Pay, Why Quit, Top Love, Top Pain, and Unmet Need.

Quantify theme frequency only with consistent sampling and coding. Otherwise use repeated, occasional, or isolated and report limitations.

## UA and creative research

Record platform, observation date, first-three-second hook, fantasy, mechanic shown, problem, action, transformation, payoff, emotional trigger, authenticity, variants, and visible longevity.

Duration, impressions, spend, or recurrence establish observed exposure or deployment, not profitable performance. Call a creative a performance winner only when a relevant outcome metric, comparator, attribution context, and uncertainty support that conclusion. Otherwise use "long-running", "high observed exposure", or "repeated pattern".

## Missing data

Never fill gaps with plausible-looking values. Request the smallest useful export or screenshot and specify product list, platforms, countries, metric definitions, windows, granularity, and filters. State how missing data limits confidence or the decision.

## Metric comparability and exclusions

Record currency, time zone where material, gross/net definition, included revenue streams (IAP, paid downloads, subscriptions, IAA), fees/tax treatment, coverage limits, and provider revision notes when available. Do not treat app-store spending estimates as total business revenue or profit. Missing IAA makes comparisons with ad-funded games incomplete.

Retain store-specific app IDs and document cross-store unification to avoid double counting. Align dates, markets, platforms, release cohorts and metric definitions. Missing/suppressed is not zero; record incomplete periods and data revisions. Do not silently mix providers or infer historical values from current rankings.

Label weak or failed cases by observable criteria and date: revenue decline, stalled growth, closure, or a stated benchmark. Low store revenue alone does not establish failure, especially when IAA or other channels are unobserved.

Match confidence to the exact claim. A high-confidence competitor download trend does not imply high confidence in a new concept's retention or commercial viability. Sensor Tower and other estimates do not replace direct concept validation.

SHA-256: 244b94b77fb03adc89929608e8a93009b18328ebf50b6186371d3fa6840abda2