Many projects do not fail because the team executed badly. They fail because too much was invested before critical assumptions were tested.
A team can spend six months building a Product before discovering that customers do not consider the problem important enough. Marketing can commit a large Media Budget before learning which Message generates qualified demand. A business can open a full location before knowing whether demand in that area will repeat.
A Small Experiment is a way to buy information for a decision before buying a much more expensive commitment.
But doing something small does not automatically make it a useful Experiment. A useful test starts with clarity about what remains unknown, what evidence would reduce that uncertainty and what result would change the decision.
A Small Experiment does not remove risk. It helps you learn before being wrong becomes expensive.
Before a major investment, a business usually relies on several unproven assumptions: customers have the problem, they will pay for the solution, the Channel can reach them, or the Business Model can generate acceptable economics.
Small Experiments turn these assumptions into testable questions: Assumption → Hypothesis → Experiment → Evidence → Decision
You do not need to build the complete Product to answer every question. If the main uncertainty is Demand, you may start with Customer Interviews, a Landing Page, Prototype, Pre-order or Pilot and increase the commitment as the evidence becomes stronger.
Strategyzer recommends identifying assumptions that are both critical and weakly supported, then defining the Hypothesis, Test, Metric and Success Threshold before the experiment is run.
Before experimenting, identify the risk you are trying to reduce
Suppose a business is considering a new Subscription Service.
Before investing, the idea may depend on several assumptions:
- Customers genuinely experience the problem
- The problem is important enough to solve
- Customers understand the Proposition
- Customers will sign up
- Customers will pay 499 baht per month
- The service can be delivered at an acceptable Cost
- Customers will continue using it after the first month
These are different risks.
Strategyzer groups important Business Idea assumptions into dimensions including Desirability, Feasibility and Viability and recommends prioritizing Critical Assumptions that currently have weak supporting evidence.
The first question should therefore not be: “What experiment should we run?”
It should be: “What must be true for this investment to work, and which of those assumptions do we know the least about?”
Small Experiments reduce Decision Uncertainty; they do not eliminate all risk
Imagine the full project requires an investment of 5 million baht.
The team is still uncertain whether customers want the service.
One option is to build everything and launch.
Another is to create a Landing Page and Prototype to test whether Target Customers understand the Proposition, take an observable action and fit the expected customer profile.
A Landing Page cannot prove that the 5-million-baht project will succeed.
It may, however, answer specific questions before the business commits 5 million baht.
That is the value of an Experiment.
It does not need to create 100% certainty.
It needs to reduce uncertainty that matters to the next decision.
Start with the Critical Assumption, not the easiest thing to test
Teams often choose questions that are easy to test, such as: “Which Logo do people prefer?”
while the real risk may be: “Does anyone need this Product?”
The Logo Test can produce valid data without reducing the most important uncertainty.
A more useful approach is to list the assumptions and ask:
- How damaging would it be if this assumption were wrong?
- How much evidence currently supports it?
High-impact + Low-evidence assumptions should receive priority.
Strategyzer uses this logic in Assumptions Mapping, where teams identify high-risk, low-evidence hypotheses before allocating resources to experiments.

Turn the Assumption into a testable Hypothesis
Assumption: “Customers will probably be interested in this service.”
This is difficult to test because “interested” is vague.
A stronger Hypothesis might be: “At least 15% of Target Visitors who see the Proposition will register for a trial within 14 days.”
Now the test has:
- Target
- Behavior
- Metric
- Threshold
- Timeframe
Strategyzer's Test Card makes four elements explicit before an experiment:
- What must be true?
- How will we test it?
- What will we measure?
- What does success look like?
Defining the threshold before seeing the data matters because it reduces the temptation to redefine “success” after the result is known.
A useful Experiment does not need to be large, but it must match what you need to learn
Suppose the Hypothesis is: “SME owners struggle to combine Sales Data from multiple Channels.”
Customer Interviews may be useful early evidence for understanding whether the problem exists, when it occurs and how customers currently solve it.
Now change the Hypothesis: “SME owners will pay 1,500 baht per month for a Tool that solves this problem.”
An Interview asking: “Would you be interested?”
provides relatively weak evidence for Willingness to Pay.
The team may need evidence closer to behavior:
- Pricing Test
- Landing Page with Price
- Demo Request
- Trial Sign-up
- Pre-order
- Paid Pilot
- Actual Purchase
Strategyzer describes evidence as generally becoming stronger as experiments move closer to Real-world Behavior and meaningful customer commitment, while faster and cheaper tests are useful earlier for exploration.
Evidence Strength should increase with the size of the commitment
Before spending 20,000 baht to learn, Directional Evidence may be sufficient.
Before committing 20 million baht, stronger evidence should normally be expected.
The question is therefore not simply: “Do we have Data?”
Ask: “Is this evidence strong enough for the size of the decision?”
An illustrative Evidence Ladder is:
Level 1: Customer Says - Interview / Survey / Stated Interest
Level 2: Customer Acts - Click / Sign-up / Request Demo / Join Waitlist
Level 3: Customer Commits - Deposit / Pre-order / Paid Pilot / Contract Intent
Level 4: Market Behavior - Actual Purchase / Repeat Purchase / Retention / Real Usage
Level 5: Controlled Evidence - Treatment vs Control / Randomized Experiment where appropriate
Higher levels are not necessary for every question, and the levels are not perfect substitutes for one another.
The point is that: “Customers say they are interested”
is not equivalent to: “Customers committed real money.”
A Small Experiment does not always mean an A/B Test
Experimentation in business covers different types of learning.
Exploratory Tests can include:
- Customer Interview
- Prototype Test
- Landing Page
- Fake Door Test
- Concierge Test
- Pilot
- Pre-order
Controlled Experiments address stronger causal questions such as: “Did changing the Message increase Conversion?”
When appropriately designed, A/B Tests with Random Assignment can separate a Treatment Effect from environmental changes more effectively than a simple Before-After comparison. Microsoft Research describes randomized A/B experimentation as an important method for causal inference because Treatment and Control groups can be made comparable at the start of the experiment.
Therefore: Prototype Test ≠ Randomized Experiment
Both are useful. They answer different questions.
Do not confuse Before-After patterns with evidence of Cause
Suppose a retailer changes its Promotion on June 1.
Conversion Rate:
Before = 3.2%
After = 4.1%
FACT:
Conversion increased after the Promotion changed.
But we should not automatically conclude: “The new Promotion caused a 0.9 percentage-point increase.”
During the same period, other factors may have changed:
- Seasonality
- Payday
- Traffic Mix
- Competitor Promotions
- Stock Availability
- Media Campaigns
- Website Experience
Before-After Analysis can identify a Pattern.
It does not automatically eliminate Alternative Explanations.
If the decision requires Causal Evidence and the context allows it, an appropriate Comparison Group or Randomized Control can provide stronger evidence.
Example 1: Before opening a new location, test demand without immediately making a long-term commitment
A full location may require:
- Rent
- Renovation
- Equipment
- Staff
- Inventory
- Marketing
A Critical Assumption might be: “This area has sufficient recurring Demand from our Target Customers.”
A smaller test could involve:
- Pop-up
- Temporary Booth
- Delivery-only Test
- Weekend Pilot
- Area-specific Pre-order Campaign
None perfectly replicates a permanent location.
But if a Pilot shows very weak Demand under reasonably favorable conditions, learning that before signing a long lease may be far less expensive than discovering it afterwards.
Example 2: Before committing a large Media Budget, test the Message
Suppose the full Campaign Budget is 3 million baht.
The team has three Messages:
A: Save Money
B: Save Time
C: Easy to Use
Instead of choosing in a meeting, the business can use a smaller budget to test Creatives under comparable Audience and Delivery Conditions.
But the Primary Metric must be defined before the test.
CTR?
Qualified Lead?
Conversion?
Revenue?
If the Business Goal is Qualified Leads but the team chooses CTR after seeing that it produces the most attractive result, the experiment may optimize the wrong outcome.
A strong experiment defines the Metric before it runs.
Example 3: Before building a Feature for four months, test whether users actually want it
Suppose a new Feature requires four months of development.
The team is not yet sure users want it.
Early tests might include:
- Prototype
- Clickable Mockup
- Fake Door
- Manual Concierge Service
- Limited Beta
If many users click the Fake Door, there is evidence of Interest in the Feature entry point.
It does not yet establish:
- Actual usage after launch
- Repeat usage
- Willingness to Pay
- Retention
- Long-term Product Value
- A Click is evidence of Interest.
It is not evidence of long-term Product Value.
This illustrates a fundamental rule of experimentation:
The Metric must match the Claim.
Do not measure what is easy if it does not answer the Decision
Business Question: “Should we invest in this Subscription Service?”
Metric:
Instagram Likes
Likes may measure Content Engagement.
They provide limited evidence about Willingness to Pay or Retention.
Metrics closer to the decision could include:
- Landing Page Conversion
- Trial Activation
- Paid Conversion
- Repeat Usage
- Cancellation
- Contribution Margin
An Experiment is not useful because it generates a lot of Data.
It is useful when the Data can change the Decision.

Define the Decision Rule before seeing the result
Suppose the Primary Metric for a Pilot is Paid Conversion.
Before the test, the team defines:
GO ≥ 12%
REVISE 7–11.9%
STOP / Rethink < 7%
These figures are illustrative, not universal Benchmarks.
Real thresholds should reflect:
- Economics
- Existing Baseline
- Strategic Requirements
- Cost
- Risk
- Minimum Viable Outcome
Defining a Decision Rule in advance clarifies how Evidence will affect the decision and reduces post-hoc interpretation.
Strategyzer's Test Card similarly requires a Success Criterion or Threshold to be specified as part of the test design.
A Hypothesis that is not supported does not mean the Experiment failed
Suppose the Hypothesis is: “20% of Landing Page Visitors will Request a Demo.”
Actual result: 6%
If the test was well designed, Traffic matched the Target and Measurement was reliable, the 6% result is useful information.
Possible explanations include:
- The Proposition is unclear
- The Need is not strong enough
- The Audience is wrong
- The Offer is weak
- The Price is too high
The result alone does not tell us which explanation is correct.
More evidence may be needed.
An Experiment that refutes a Hypothesis is therefore not necessarily a Failed Experiment.
It may be Cheap Learning that prevents an Expensive Failure.
Poorly designed Small Experiments can create False Confidence
Saying: “We tested it.”
does not guarantee trustworthy evidence.
Check:
- Did the Participants match the Target?
- Was the Stimulus realistic enough?
- Did the Metric match the Hypothesis?
- Was the Sample adequate for the Claim?
- Was there Selection Bias?
- Could Seasonality or External Factors explain the result?
- Was Measurement working correctly?
- Was the Test Duration appropriate?
- Were Metrics selected after seeing the results?
Strategyzer emphasizes that experiments need clear, testable hypotheses and appropriate participants and test artefacts to produce evidence strong enough to reduce uncertainty.
Do not make the Experiment so small that it cannot answer anything
Small Experiment does not always mean:
Small Sample
Short Duration
Low Budget
“Small” should mean: A commitment smaller than full-scale investment, but still large enough to generate evidence useful for the decision.
If the test is so small that Noise dominates the result or too few Target Customers participate, the experiment can increase confusion instead of reducing uncertainty.
Experiment size should reflect:
Decision Risk
Expected Effect
Variability
Required Precision
Available Traffic / Sample
Cost of Being Wrong
As risk increases, move from Cheap Tests toward Stronger Evidence
Testing does not need to happen in one step.
A possible sequence is:
Stage 1: Explore
Customer Interview
Question: Does the Problem exist?
Stage 2: Test Interest
Landing Page
Question: Does the Proposition generate Action?
Stage 3: Test Commitment
Pre-order / Paid Pilot
Question: Will customers commit money or time?
Stage 4: Test Experience
Limited Pilot
Question: Do customers receive Value and return?
Stage 5: Validate Impact
Controlled Experiment or appropriate Market Comparison
Question: Does the Intervention generate Incremental Outcomes?
Stage 6: Scale
Increase investment as Evidence becomes stronger
Strategyzer recommends sequencing fast, inexpensive Discovery Tests before moving toward experiments with stronger evidence and larger commitments.
An Experiment should end with a Decision, not just a Report
The team should not stop at: “Conversion = 11.4%.”
Ask:
What did this Evidence change?
Which Critical Assumption remains uncertain?
Should we invest more?
What needs to change first?
What should we test next?
The decision may be:
GO Evidence is strong enough for the next level of commitment.
REVISE There is a promising Signal, but the Proposition, Product or Execution needs adjustment.
STOP Evidence challenges a Critical Assumption strongly enough that the current approach should not receive further investment.
TEST AGAIN Evidence remains unclear or the Experiment has important limitations.
Strategyzer emphasizes that the purpose of the Testing Cycle is to use evidence to decide whether to persevere, pivot or stop, rather than measuring progress by the number of experiments completed.
8 questions to ask before a major investment
- What must be true for this investment to work?
- Which Assumption would cause the greatest damage if it were wrong?
- What Evidence do we already have, and how strong is it?
- What is the smallest Experiment that can answer the Critical Question?
- Which Metric actually matches the Hypothesis?
- What are the Success Threshold and Decision Rule?
- If the result contradicts our expectations, are we genuinely prepared to change the decision?
- Is the current Evidence strong enough for the amount of money and risk we are about to commit?
If the answer to Question 7 is “no,” the Experiment may be serving to justify a decision already made rather than helping the team learn before deciding.
The takeaway: Invest in learning before investing in scale
Small Experiments cannot make every business decision safe.
Businesses will always make decisions under uncertainty.
What Experimentation can do is test important assumptions before the Cost of Being Wrong becomes larger.
A useful sequence is: Business Decision → Critical Uncertainty → Hypothesis → Smallest Useful Experiment → Evidence → Decision Rule → Next Investment
Fast and inexpensive tests can be appropriate during Exploration. As the size of the decision increases, Evidence Strength should also increase. And when the claim is causal, a suitable Experimental Design is stronger than relying on a Before-After Pattern alone.
The question before approving a major investment should therefore not only be: “How confident are we in this idea?”
A more useful question is: “What is the most important assumption behind this idea, and can we buy stronger evidence with an experiment that costs less than making the full investment?”

A Small Experiment is not designed to prove that an idea is “right.” Its purpose is to reduce uncertainty before a larger commitment. Start by identifying the Critical Assumption that could make the investment fail, then define the Hypothesis, Test, Metric and Success Threshold before seeing the result. Choose the smallest and fastest experiment that can still produce evidence strong enough for the next decision. As the size and risk of the investment increase, the required Evidence Strength should increase as well.
Sources
- Strategyzer. Validate Your Ideas with the Test Card. Framework for defining Hypothesis, Test, Metric and Success Threshold before scaling a Business Idea.
- Strategyzer. How to Test Your Idea: Start With the Most Critical Hypotheses. Explains testing Critical Assumptions across Desirability, Feasibility and Viability before committing to implementation.
- Strategyzer. How Strong Is Your Innovation Evidence? Discusses Experiment Speed, Evidence Strength and moving closer to Real-world Behavior as validation becomes more important.
- Strategyzer. Designing Strong Experiments. Covers Testable Hypotheses, participant selection and Experiment Design for producing stronger evidence.
Microsoft Research. Experimentation and the North Star Metric. Explains the distinction between Association and Causality and the role of randomized A/B Experiments in estimating Causal Impact. - Microsoft Research. Online Experimentation at Microsoft. Discusses Controlled Experiments and Randomization for evaluating how Product Changes affect Customer Behavior.
.png)



.png)
.png)
.png)
.png)






