Skip to content
    Piiko
    ← All notes

    Note 07 · Benchmarks

    Which mobile app benchmarks should you use?

    Piiko·4 min read·Reviewed

    Look a little closer.

    A magnifying glass examines two app tiles with matching hearts; a different round star token sits outside their frame.
    A useful comparison begins with a good match, not the biggest number.
    The short answer

    Use benchmarks to locate a question worth investigating. Match the population, metric definition, and observation window before treating a number as a useful comparison.

    Choose a comparison that resembles your app.

    A subscription utility, an ad-supported game, and a marketplace do not share one useful definition of “good monetization.” Even within a category, country, platform, pricing, acquisition source, and plan mix can change the comparison.

    Write a comparison brief before looking for a number. For example: paid iOS installs in one country, acquired during the same period, observed through day 30, with net proceeds divided by every install. That is a much narrower question than “What is a good LTV?”

    If the reference cannot match your filters, label the gap. A broad number can still provide context, but it should not become a precise spending target.

    Use each source for the question it can answer.

    RevenueCat’s State of Subscription Apps 2026 is a starting point for subscription-app context. Its customer dataset represents participating businesses, not every app in the market. Check each chart’s category, time window, and definition before using it.

    RevenueCat’s dashboard benchmarks compare store and category peers and distinguish metrics such as initial conversion, paying conversion, and realized value. Initial conversion can include a trial; paying conversion requires a payment. Those are not interchangeable targets.

    For acquisition comparisons, AppsFlyer’s benchmark methodology explains its data and aggregation. Its performance benchmarks average app-level results with equal app weighting after filtering; other sections use aggregated data. It also documents outlier exclusions and geographic fallbacks when a narrower sample is insufficient.

    Read these methodology notes first. An average across apps answers a different question from a rate calculated across all users in those apps.

    A denominator can explain a tenfold difference.

    Imagine a cohort with 1,000 installs, 100 trial starts, and 20 eventual trial conversions. Trial-to-paid is 20 / 100 = 20%. Install-to-paid is 20 / 1,000 = 2%. The business outcome is identical; the denominator changed.

    That is a hypothetical example, not an industry benchmark. Use it as a reminder to record what each percentage means before comparing it with your own dashboard.

    Keep this beside every benchmark
    1. Source, edition, and retrieval date.
    2. Category, country, platform, and acquisition mix.
    3. Numerator, denominator, and cost or revenue basis.
    4. Cohort start, observation window, and maturity.
    5. Sample coverage, weighting, and known exclusions.

    A missing segment is not zero performance. A median is not an average, and a top-quartile value is not a guarantee that changing one setting will get you there.

    Turn the gap into a testable question.

    If comparable apps convert more trial users, investigate whether your product delivers the promised value before the trial ends. If their early LTV is higher, check pricing, country mix, plan mix, and refunds before concluding your paywall is the problem.

    Differences across apps are associations. They do not prove that copying a trial length, price, or onboarding flow will cause the same result in your app.

    Keep your own mature cohort baseline beside the external reference. Define a change, a primary outcome, and a guardrail, then measure. Start with a clearly defined conversion funnel or realized cohort value.

    Sources & further reading

    Examples are hypothetical. They illustrate the method and do not represent Piiko results or industry benchmarks.