← Studio Log
C. 신뢰 구축 (Niềm tin)수익화가격 실험무료 유료 전환제품 기획사용자 신뢰

How to Design the Experiment That Moves a Product from Free to Paid

How to Design the Experiment That Moves a Product from Free to Paid
by Yeowubie

Monetization Is Not Picking a Price, It Is Designing an Experiment

Turning a free product into a paid one is not a decision about a number. It is a design task that verifies what can be sold, who is willing to pay, and under what conditions a wallet opens. A price list is the output of that verification, not its starting point. Reverse the order and weeks disappear into argument.

Pricing meetings run long mostly because of missing information. Nobody in the room knows where users actually get stuck, so everyone presents intuition as if it were evidence. Once the conversation mixes what a competitor charges, how many months the build took, and what the team feels it deserves, no conclusion arrives. All three describe our situation rather than the user's.

Change the question and the meeting shortens. Instead of asking what to charge, ask first what is keeping the people who are still here. That question can be answered by observation. Look at which screens they return to, which task they repeat several times a week, and at what moment they leave. Price is the last layer placed on top of that observation.

Many teams treat the free period as a loss. The real purpose of keeping something open, though, is not only acquiring users but gathering the evidence a price will stand on. You need accumulated detail about which features people repeat and where they get frustrated before you can pick a candidate to charge for. Without that record, any plan you build is a plan built on guesses.

Designing it as an experiment means writing four things down in advance. What you are trying to verify, who the group is, what you observe and for how long, and which result means continue while which result means stop. Without those four lines it is not an experiment, it is a feature release. A feature release has no criteria for interpreting the outcome, which means any story can be attached to it afterward.

Yeowubie Interaction intends to work in the same order. The plan is to keep the products already published to the stores, onSpots, Langtori and Job Connect VN, open first, observe how they are actually used, and attach paid experiments on top of that. Because measurement instrumentation is not yet connected, we have no conversion figures to publish. What we can describe at this stage is the sequence, not the outcome. Stating the sequence up front is how we intend to earn the right to talk about results later.

Choosing What to Charge For: Frequency and Urgency

Candidates for paid features are selected along two axes: how often the user reaches for it, and how much trouble they are in without it. Something used frequently that causes real difficulty in its absence is the first candidate. The fact that a feature took a long time to build, or was technically hard, is not a selection criterion.

Start with frequency. A task repeated on a weekly rhythm becomes a habit, and a habit carries a high replacement cost. A feature used once a month, however impressive, has a weak reason to be paid for. The user feels the need only in that moment and forgets it afterward. When reading frequency, do not look at the overall average but at the usage interval of the people who genuinely keep coming back. Averages get dragged down by everyone who tried once and left.

The urgency axis is measured by the cost of failure. Write down what the user has to do instead when the feature is not there. If they have to write it on paper again, ask each supervisor one by one, or if missing it leaks money or opportunity, urgency is high. If the absence is merely mildly inconvenient, it is low. Urgency is not a feeling, it shows up as the substitute labor the user has to perform.

Overlay the two axes and the judgment gets simple. Frequently used and painful when missing is the paid candidate. Frequently used but survivable without stays free, because it keeps new users arriving. Rarely used but critical when missing tends to fit per-use charges or an add-on line. Rarely used and survivable without is not a monetization target at all, it is a cleanup target.

Applied to products, the split becomes concrete. In onSpots, the attendance service, the urgent work sits with the administrator who compiles, checks and exports records as evidence, rather than with the daily act of logging. In Job Connect VN, the recruitment service for the Vietnamese market, the side under time pressure is the hiring side handling listing visibility and applicant organization, not the candidate browsing. Langtori provides matching and mapping between people learning Korean and people teaching it, so the point where value can attach is the quality and durability of the connection, not the learning material itself. Teaching on Langtori is carried out by the individual tutors and institutions participating on the platform. Including Winnie, the platform supporting small merchants, this point differs product by product, so we do not copy one product's criteria onto another.

One common mistake deserves a direct mention: deciding to charge because something was expensive to build. Users do not know our development cost and have no reason to. Payment is decided by comparing what they gain against what they give up. If an expensive feature happens to also score high on frequency and urgency, good. When the two disagree, drop the cost argument and follow the user.

Cutting the Experiment Into Small Enough Units

A good experimental unit consists of one question, one target group, and one observation window. If any of those three contains more than one item, the result cannot be interpreted. A large change such as introducing a full pricing tier is not an experiment, it is closer to a bet placed while several variables move at once.

Consider what happens when a pricing plan lands all at once. The price level, the feature boundary, the wording of the notice, the checkout flow and the announcement timing all change together. If the outcome is bad you cannot tell which element caused it, and if the outcome is good you cannot reproduce it next time. Reversing is also hard. The announcement that withdraws a pricing plan costs far more than the one that introduced it.

There are three main ways to cut. Cutting by audience means applying the change only to new signups, or only to one region, or only to one usage type. Cutting by feature means leaving the plan alone and moving a single item behind the boundary to watch the reaction. Cutting by time means opening something for a defined window and returning to the previous state when it closes. One of the three is enough, and for a first attempt, cutting by audience is the safest.

Designing for reversibility is partly an engineering matter. If experiment membership is driven by configuration rather than hard-coded, you can roll back without a deployment when something goes wrong. Once the experiment reaches a stage where real charges occur, the refund path and the wording for that case belong inside the experiment design too. An experiment with no prepared way back is better left unstarted.

Keep the window short, but not too short. In a product whose usage cycle is weekly, three days of observation tells you nothing. Allow at least two or three turns of the natural cycle, and beware that stretching much further lets market conditions shift until the experimental conditions themselves are unstable. Before setting the window, identify the product's natural rhythm.

Finally, write the decision criteria down before starting. Record which number above which line means expand, and which line below means stop. If the criteria are constructed after seeing the result, people generally move them toward what they hoped to see. A single paragraph written in advance blocks that drift. That document also becomes the starting point when the same experiment is run again later.

Asking About Willingness to Pay Without Losing Users

Willingness to pay can be probed before a checkout screen exists. Describing a paid item and collecting interest, showing options when someone reaches a usage limit, and taking an actual payment are three different experiments. Often the answer from the earliest stage is enough to determine the next decision.

The lowest pressure method is to collect interest first. Describe clearly what would be paid, then ask whether the person wants to be notified. What you learn here is not precise demand but the distribution of interest. Which items attract signups, and which usage types respond, sets the direction of the next experiment. Nothing is charged at this stage, so the user loses nothing by answering.

How existing free users are treated is what decides trust. When a feature someone already relies on suddenly locks, they do not experience a lost feature, they experience a broken promise. So when redrawing a boundary, preserve the old terms for existing users, or at minimum give enough grace and advance notice. A grace period is not better for being longer, it needs to be long enough for the user to prepare an alternative. What matters is not the length but the fact that warning was given.

There is a line worth stating plainly. We do not hide the cancel control where it is hard to find, do not charge automatically at the end of a trial without announcing the date, and do not build a flow where signing up takes two taps while cancelling takes five steps. These designs lift a short-term number and leave behind refund requests, negative reviews, and users who never return. If the goal is a relationship that outlasts the contract, remove those options at the start. Payment terms and the method of cancellation should appear at the same type size on the screen immediately before checkout.

Rejection is data too. Give people who viewed the paid notice and closed it a place to say why, and you can separate a price problem from a feature problem from a timing problem. The three call for entirely different responses. If it is price, redraw the boundary. If it is the feature, it was not ready to be sold. If it is timing, move where the notice appears. That question should be short and skippable.

The wording of the notice is also fair game for testing. When changing wording, though, do not change the facts. Stating the same terms more clearly and making unfavorable terms less noticeable are different acts. The first is an improvement, the second returns later as a cost. When polishing the copy, judge it by whether anything is left to surprise the user after they pass that screen.

Reading the Results: What Counts as Signal and What Is Noise

Reading results requires two rules. First, use exactly the criteria set before the experiment. Second, accept that early experiments rarely reach statistical confidence, and read qualitative material alongside the numbers. Break either rule and what you are reading is not the data but the expectations of whoever is interpreting it.

Start with the risk of small samples. In an early product the observed group is often in the tens or low hundreds, and at that scale a few days of ordinary variation exceeds any real difference. Comparing figures past the decimal point is meaningless there. It is safer to treat only wide gaps as signal and to defer judgment on narrow ones. Repeating the experiment and checking whether the direction holds is frequently more practical than a statistical test.

Separating words from behavior matters just as much. Answering a survey that you would pay, tapping the notice, reaching the checkout screen, and completing a payment are four signals of different strength. Reliability rises toward the end while sample size falls. Concluding from the earliest indicator alone produces a picture more optimistic than reality. Waiting only for the last indicator means never deciding, so the workable approach is to read every stage while weighting them differently.

Sources of noise are mostly predictable. A campaign or an outside mention pulls in visitors with a different profile and conversion drops. Traffic arriving through a single channel lets that channel's tendencies color the whole result. Cycles already present in the market, holidays, academic terms, accounting close, shake the numbers as well. Recording what happened during the observation window keeps you from misreading the same figures months later.

Handling a failed experiment should also be a rule. Missing the threshold confirms that this item cannot be sold under current conditions, not that it can never be sold. Leave a single page describing what was tried, what came out, and what would have to change before trying again. As those pages accumulate the team stops repeating the same experiment, and someone who joins later can read the context behind past decisions.

Finally, a monetization experiment is an experiment about the product and an experiment about the relationship at the same time. Users remember how we spoke and how we stepped back more than they remember what we tried to sell. If you gave notice, asked, and accepted the refusal, that user is prepared to hear the next proposal. If you charged quietly and made cancellation difficult, you keep the first month's number and lose the relationship. In the passage from free to paid, the thing that actually needs designing is not the price list but the way this relationship is run.

Related posts