Write the causal relationship
“Test a new landing page” is not a hypothesis. A useful hypothesis states what will change, for whom, why and which observable result should follow. For example: reducing the number of required fields for qualified enterprise visitors will increase completed requests without reducing downstream qualification. That structure forces the team to expose the mechanism rather than jumping directly to a tactic.
Define the decision criterion before the test starts
Teams often reinterpret results after seeing them. A better discipline is to write the rule in advance: what result is strong enough to scale, what result is weak enough to stop and what ambiguous outcome requires another test. The rule can include not only a statistical metric but also business quality: qualified leads, gross profit, retention, meeting rate or another downstream event.
Prioritisation should not create false precision
Scoring systems are useful for comparison, but a number with two decimal places does not make uncertainty disappear. Treat priority scores as a way to structure discussion around impact, evidence gap, test cost, speed and reversibility. High-impact assumptions with weak evidence often deserve attention first because the cost of being confidently wrong is high.
Manage a portfolio, not a queue of ideas
A strong experiment system balances different types of learning: acquisition, offer, pricing, experience, retention, sales process and measurement. If every test is a creative variant, the organisation may optimise the surface while preserving a weak business model underneath. A portfolio view also forces explicit trade-offs: what are we not testing this cycle and why?
The system starts with the right to close hypotheses
Some teams keep weak ideas alive because too much effort or status is attached to them. Learning requires permission to conclude that an assumption is wrong. Closing a hypothesis is not failure when the test prevents a larger investment. This is why the experiment owner must have access to the decision, not only to the analytics.
Not every test needs statistical perfection
The standard should match the decision. A high-volume checkout experiment may justify formal statistical testing. A small B2B market with eighty target companies may rely on a different evidence design: repeated buyer conversations, pilot commitments, observed stage movement and economic feasibility. What matters is whether the evidence is strong enough for the cost and reversibility of the next decision.
Practical case: turning an idea list into a decision portfolio
A growth team has forty ideas. Instead of executing the loudest requests, it rewrites them as hypotheses, attaches the current evidence, scores the cost of being wrong and identifies the next proof that would change the decision. Half of the ideas disappear because they cannot articulate a mechanism. Several are merged. The remaining portfolio becomes smaller but more valuable: each test has an owner, a decision date and a clear next action.
30-day protocol
- Rewrite the current backlog as causal hypotheses, not tactics.
- Attach current evidence and the cost of error to each hypothesis.
- Choose a balanced set of 3–5 tests for the next cycle.
- Write scale / stop / continue rules before launch.
- At the end of the cycle, record what the team learned and close hypotheses that no longer deserve attention.
How to assess experiment-program quality
- Share of tests with a pre-written decision rule.
- Share of hypotheses closed or materially updated.
- Time from hypothesis to decision.
- Percentage of tests connected to a downstream business metric.
- Repeated testing of the same assumption without new evidence.
- Amount of experiment debt: old tests without a documented conclusion.
A good hypothesis card
A practical card contains: business question, target segment, causal mechanism, current evidence, expected observable change, primary metric, quality guardrail, test cost, decision deadline, scale condition and stop condition. The purpose is not paperwork. It is to make the reasoning visible enough that another person can challenge it before money is spent.
Decision note
A strong analysis makes its assumptions visible, connects evidence to a decision and defines the next observation that can confirm, weaken or close the hypothesis.
The purpose of this note is not to make uncertainty disappear. It is to make the assumptions visible, connect them to a decision and define the next evidence step.