← Studio Log
B. 의사결정 (Hỗ trợ quyết định)앱 출시리텐션초기 KPI코호트 분석프로덕트 애널리틱스

The First 90 Days After Launch: What to Actually Measure

The First 90 Days After Launch: What to Actually Measure
by Yeowubie

Why Download Counts Are Not a Launch Result

Downloads in launch week measure attention, not performance. Most of that attention comes from friends, internal announcements, and whatever promotion ran on day one. It says nothing about whether the product is worth using. What the first 90 days must answer is not how many people installed, but how many of them came back.

Downloads look attractive for an obvious reason. The number sits in the store console by default, requires no instrumentation, is comfortably large, and only ever moves up. Because it is cumulative it never falls, which makes it the easiest figure in the world to put in a weekly report. A metric that cannot fall, however, cannot describe a state. The product can get worse while the curve keeps climbing.

The same thousand installs mean completely different things depending on where they came from. Installs driven by an internal notice reflect organizational instruction, not product preference. Installs collected from a QR code at an event are a product of the room, and most of them are gone from the phone by that evening. Installs from paid media arrive carrying whatever expectation the creative set, and when that expectation misses the product, first day drop is steep. Collapse all three into one line and nothing can be read from it.

Treat downloads as a denominator rather than a result. As a denominator they are genuinely useful. To ask what share of this week's installs finished onboarding, or what share returned the next day, you need something on the bottom of the fraction. The moment a denominator is mistaken for a result, the whole team starts optimizing inflow while the leak at the bottom stays exactly where it was.

Even in the first weeks there is one number worth looking at before raw installs: the conversion rate from store listing impression to install. That ratio describes the quality of the listing, not the quality of the product. The icon, the first two screenshots, the title, and the opening line of the description decide most of it. Raising ad spend while that conversion is weak is pouring more water into a leaking bucket.

One more caution. The first two weeks produce the most distorted data you will ever see from your own product. Engineers, internal testers, partners, and people reacting to launch publicity all arrive mixed together. Treating that cohort as a baseline throws off every judgment that follows. Separate internal accounts and test devices from the start, and set your baseline from cohorts that arrive in week three or later.

Why Retention Comes First: What D1, D7, and D30 Tell You

Retention is the metric that tells you whether the product kept its promise. Day one reflects first impression, day seven reflects the seed of repeat use, and day thirty shows whether the product has found a place in someone's routine. Each number points to a different kind of failure, so watching only one leaves you guessing about what to fix.

Agree on definitions before you argue about values. Day one retention means something quite different if it counts people who opened specifically on the following day versus people who returned at any point after install. Tools ship different defaults, and switching tools quietly breaks comparison with your own history. Writing the definition down and confirming that the tracking code and the dashboard use the same one matters more, early on, than moving the number.

Weak day one retention almost always points to the entrance. Forced registration, permission requests stacked before any value appears, a first screen that gives no clue about what to do next, and a gap between what the ad promised and what the app delivers. None of this has much to do with the depth of the product. However good the feature buried three screens in may be, most people never reach it.

Weak day seven retention means there is no reason to open the app again. For a utility that finishes its job in one session, a low number here is natural, and that kind of product should be judged on return interval and completion rate instead. But when a product designed for repeat use loses day seven, the core loop is missing. Teams often read this as a notification problem. Notifications do not manufacture a reason; they only amplify one that already exists.

Weak day thirty retention means the early interest was real and the lasting value was not. Content runs dry, the app is configured once and then has nothing left to ask for, or data accumulates without ever being handed back to the person who produced it. Improvement here usually comes from returning accumulated data to the user in a useful form, not from adding features.

This article will not quote an industry average retention figure. Benchmarks whose source cannot be stated plainly tend to be used for reassurance rather than judgment, and expectations differ fundamentally by product type anyway. Compare yourself to yourself instead. Whether this month's cohort beats last month's, and whether the cohort that arrived after you rewrote onboarding beats the one that arrived before, are far more trustworthy tests.

Different usage rhythms deserve different expectations. An attendance app is supposed to open on every working day, so daily retention is meaningful for it. onSpots, which Yeowubie Interaction builds and operates, is that kind of service. A job service works the other way: it is used intensely during a search and then drifts away naturally once someone is hired. Holding a service like Job Connect VN to daily retention produces the wrong conclusion. Define the normal usage cycle for your product first, then choose a retention window that matches it.

Finally, retention has to be split apart. Split by acquisition channel, by operating system, by device performance tier, and for a multilingual product, by locale. In a structure that serves Korean, Vietnamese, and English side by side, as Langtori does, each locale carries a different acquisition path and a different usage context, so a single blended average hides which side is actually struggling. Averages are very good at concealing two opposite truths at once.

Defining Activation: The Moment Value Lands

Activation is the first moment a user actually feels the core value of the product. Without defining that moment as a single trackable event, you cannot explain why retention is low. The definition can start as a team hypothesis, but it has to be tested against data within 90 days and rewritten if the data disagrees.

The method is simpler than it sounds. Write one sentence describing why the app exists, then pick the smallest behavior that proves the sentence actually happened for someone. For an attendance product, that is the first shift record saved without error. For a service that connects learners and teachers, it is a completed profile followed by a first message sent. For a job service, it is the first application submitted. What these share is that the user received something, not that the company captured something.

Completing signup is therefore not activation. Registration is a procedure that serves our convenience and costs the user effort. The same applies to uploading a profile photo, granting notification permission, and finishing a tutorial. Choose one of those as your activation event and the number will rise easily while retention refuses to follow, and no one on the team will be able to say why.

Once the event is chosen, measure two things together. The first is reach rate: the share of new users who hit the event at all. The second is time to reach it. A low reach rate means the path is blocked; a high reach rate with a long time means the path is bent. Those two problems call for different fixes.

There is also a way to test whether the definition is correct. Compare day seven retention between users who reached activation and users who did not. A wide, clear gap is evidence that the definition landed on a real value moment. A negligible gap means the definition is wrong, and you should nominate a different behavior and run the same comparison again. This test should be run at least once inside the first 90 days.

Break the path into roughly three to five steps. Too few and you cannot see where people stall; too many and noise drowns the signal. Look at drop between steps, but fix one segment, the largest, and read the result in the next cohort. Change several things at once and you lose the ability to attribute any improvement to anything.

The most important technical preparation has to be finished before launch. Settle event naming rules, required properties, and the user identity policy as a written document, and ship with engineering and product looking at the same table. Instrumentation added later does not apply retroactively. The first 90 days of data cannot be regenerated, and the earliest cohorts are never replaced by later ones. Building and operating your own products makes the cost of getting this order wrong very concrete. By the time development is finished, deciding what to measure is already late.

Four Questions to Answer Within 90 Days

The purpose of the first 90 days is learning, not growth. Who stays, what they do before staying, where they leave, and what it cost to bring them in. Answer those four with evidence and the next quarter's decisions stop being guesses. Without answers, every move is a wager.

First, who stays. This is the work of finding shared attributes among surviving users. Split by acquisition channel, region, device, job role, and language, then compare retention across the splits. If one segment clearly outperforms the rest, that segment is the product's real market right now. It may not match the target defined during planning, and if it does not, the data is the one telling the truth. What gets adjusted after 90 days is your target definition, not your users.

Second, what they do before staying. Look at which screens retained users reopen and which features they repeat. A familiar mismatch shows up here often: the feature that absorbed the most engineering effort is missing from the top of the usage distribution, while something built almost as an afterthought sits near the top. Do not explain that signal away. It is the strongest basis you will have for rewriting next quarter's roadmap.

Third, where they leave. Beyond step by step funnel drop, look at the last screen in the last session of users who churned. And pull technical quality inside this question rather than leaving it to engineering hygiene. Crashes, frozen screens, views that are slow only on certain device tiers, and failures under unstable network conditions are not separate concerns; they are direct causes of weak retention. In markets where lower specification devices carry a large share, weight this item more heavily.

Fourth, what it cost. Acquisition cost per channel only becomes meaningful when placed next to retention for that same channel. If a cheap channel loses everyone by the following day, it is the most expensive channel you have. Making that comparison possible requires tagging channel information onto the cohort at install time, which is another piece of preparation that must be done before launch.

What not to do in the same window is equally clear: large scale paid campaigns and sweeping feature expansion. Pushing a lot of traffic through before the measurement system stands makes results uninterpretable, and shipping many features at once permanently destroys attribution. Ninety days is a period for fixing your coordinates, not for scaling.

Set an operating rhythm as well. Review cohorts once a week, on the same day, in the same format, and cap the routinely watched metrics at three to five. Build several dozen dashboards and nobody opens any of them. A team that looks at the same three numbers in the same place every week learns considerably faster than one that admires a different impressive screen each Monday.

Without Metrics, You Cannot Spend a Marketing Budget

A marketing budget is a tool for telling more people about a good product quickly. It is not a tool for making an unretentive product retentive. Spending before retention and activation have been validated amounts to paying to acquire people who were always going to leave. Instrumentation comes first, scale comes after.

The order is not complicated. Instrument, validate the activation definition against data, confirm that retention holds steady across cohorts, test channels in small separated slices, and scale only the ones that survive. Skip the order and budget becomes consumption rather than learning. Keep it, and the same budget lasts considerably longer.

Early on you cannot know lifetime value, simply because the data has not accumulated yet. That means you cannot set a ceiling on acquisition cost either, so in practice teams judge with proxies. Activation reach rate by channel and survival at the thirty day mark do that job. When a channel brings people to activation reliably and they are still present a month later, that is grounds for spending more there even without a precise lifetime value figure.

The fit between creative and product can also be managed as a metric. When the ad copy overpromises, clicks and installs rise while first day retention falls. The moment those two numbers appear on the same table, marketing and product start having one conversation instead of two. In organizations that watch only cost per install, that conversation never happens, and the familiar loop repeats: the product team blames the quality of acquired users, the marketing team blames the state of the product.

Metrics need an owner. One person pulls the numbers each week under the same definition, shares them in the same format, and records the date whenever a definition changes. Leave that role empty and within a few weeks a single metric name will carry two different values, and meetings will be spent reconciling numbers instead of interpreting them.

What survives the 90 days is not the download graph. What survives is one answer, backed by data, about who this app serves and what it does for them. With that answer, where to spend, what to build next quarter, and which requests to decline all become much easier to settle. It is also why Yeowubie Interaction builds and operates its own products while developing for clients at the same time. The questions that matter after launch are not the questions that mattered during development, and only a team that has run something in production designs for them in advance.

If you are about to launch, the list to review right now is not the feature list but the event list. If you have already launched, try writing your activation moment as a single sentence as of today. If that sentence does not come out immediately, you have also just identified the most urgent item in your first 90 days.

Related posts