Product hypotheses: the roadmap that leads to market fit
A feature roadmap keeps you busy for a long time. It does not tell you whether you are heading in the right direction. The hypothesis roadmap answers that question: it lists what you need to prove to reach market fit, in order, and works out from that what to build.
The problem with a feature roadmap
Most product managers share a feature roadmap with their teams: an ordered list of things to build, often fed by user stories. It has one advantage: it provides work for a long time. And one major flaw: all that time, head down and pedalling hard, you do not know whether you are pedalling in the right direction.
A user story says what a feature delivers in an ideal world. It does not say whether that world exists.
Until you reach market fit, you are working on hypotheses, whether you write them down or not. The only question is whether they are explicit and tested, or implicit and expensive.
What is a product hypothesis?
A product hypothesis is a statement about your market, your users or your ability to deliver, which must be true for the product to succeed and which can be checked. A simple format:
We believe that [this target group] [will do this] because [this reason]. We will know when we observe [this measurable signal].
Hypotheses traditionally fall into three families:
- Desirability: do users want the solution, and will they change their habits?
- Viability: does the business model hold up, at what price, with what margin?
- Feasibility: can we build and run it, at scale?
The hypothesis card
One sentence is not enough to run a test. For every important hypothesis, I keep a short card, the same for all of them. It fits on half a page and is filled in before the test, except for the last two lines.
| Field | What goes in it |
|---|---|
| Statement | The “We believe that… We will know when…” sentence |
| Family | Desirability, viability or feasibility |
| What depends on it | The hypotheses and initiatives that fall if this one is false |
| Current evidence | What you already know: interviews, data, precedents. “None” is an acceptable answer |
| Test | The cheapest experiment that could prove it wrong |
| Signal and threshold | The measurement, and the value below which the hypothesis is invalidated |
| Decision date | The day you decide, even if the data is incomplete |
| Cost of the test | In working days and in money |
| Owner | One person, who presents the result |
| Result | What was observed, with figures |
| Decision | Continue, adjust, abandon, and what that changes on the roadmap |
The “What depends on it” line is the one people forget most. Yet it is the one that lets you order the hypotheses.
Building the hypothesis roadmap
At SipScience, I built my first hypothesis roadmap three months after joining the team, once I had understood the situation. The method has four steps.
- Start from the end goal. Not “launch V2”, but a business outcome: a revenue level, a presence in a number of markets, profitability.
- Work out the necessary conditions. What has to be true for the goal to be reached?
- Order them by phase and by risk. The hypotheses everything else depends on come first.
- Define the test for each one: an experiment, a measurement or a feature to build, with the expected signal.
An example
A fictional example: an app for booking a restaurant table at a discount, connected to the restaurant’s point-of-sale software. Goal: to be present and profitable in several major European cities.
| Phase | Hypothesis | How to test it |
|---|---|---|
| 1 | Our integration works with our pilot restaurants’ point-of-sale software. | Connect five restaurants, measure synchronization errors. |
| 1 | Restaurants agree to offer a discount in exchange for predictability. | Canvass one neighbourhood, measure the sign-up rate. |
| 1 | Users book and come back. | Eight-week retention in the first neighbourhood. |
| 2 | Acquiring restaurants and users can be repeated in a second city. | Same method, compare acquisition costs. |
| 3 | The data collected is valuable to partners. | Interviews and letters of intent. |
Notice that the first hypotheses are about the technology and the offer, not about marketing. There is no point investing in growth while the integration does not hold up. From these hypotheses you derive the items on the product roadmap: payment methods, compatibility with point-of-sale systems, restaurant onboarding. Every feature has an explicit reason to exist.
Ordering by risk, on the same example
Saying “the hypotheses everything depends on come first” is not enough when three hypotheses look equally important. A simple score helps: for each one, rate from 1 to 3 the impact if it is false, then from 1 to 3 the current uncertainty. Multiply the two to get the risk. At equal risk, the cheapest test goes first.
| Hypothesis | Impact if false | Uncertainty | Risk | Cost of the test | Order |
|---|---|---|---|---|---|
| Point-of-sale integration | 3: the whole model depends on it | 3: never tested | 9 | Medium | 1 |
| Restaurants accept the discount | 3: no offer, no product | 2: a few favourable interviews | 6 | Low | 2 |
| Users book and come back | 3 | 2: usage known from competitors | 6 | High: the offer has to exist first | 3 |
| Repeating in a second city | 2: you can stay local for longer | 2 | 4 | High | 4 |
| Data valuable to partners | 1: extra revenue | 3 | 3 | Low | 5, in the background |
Two remarks. First, the scoring is not a science: it exists to make disagreement visible. If the engineering team rates the integration’s uncertainty at 1 and the product manager at 3, that is the conversation to have. Second, the last hypothesis carries low risk but a very cheap test: you can run it in parallel without tying up the team.
The card, filled in
For the first hypothesis: statement, “we believe our connector syncs bookings with the pilot restaurants’ point-of-sale systems; we will know when fewer than 2% of bookings show an error over four weeks”. What depends on it: all the others. Current evidence: none in real conditions. Test: five restaurants, two different point-of-sale systems. Decision date: end of the second month. Owner: the engineering lead. The 2% threshold is a team choice, discussed with the pilot restaurants, not a standard.
Common mistakes, and how to spot them
- Hypotheses nothing could contradict. “Users will love the app.” Warning sign: no test result could make you change your mind.
- A threshold set after the fact. Warning sign: the threshold appears in the results presentation, not on the original card.
- Tests that drag on. Warning sign: the decision date slips at every review. A test that cannot conclude within a few weeks is often badly designed.
- Testing only feasibility. Warning sign: every card belongs to the same family, the one the team knows best how to test.
- Keeping an invalidated hypothesis “just to see”. Warning sign: initiatives that depended on it stay on the roadmap.
Small team or large organization
In a startup, the hypothesis roadmap fits on one page, and the founder or product manager keeps the cards. Three to five active hypotheses are enough. In a large organization, it becomes a governance tool: every funded initiative must point to the hypotheses it tests, and the quarterly review looks at results before budgets. It is also the best antidote to projects that carry on because they have already cost a lot.
What you gain
The hypothesis roadmap sits between the business model canvas and the product roadmap. It lets you:
- stay focused on strategic goals;
- make rational choices you can defend in front of the team and leadership;
- build only what helps validate something;
- know when and why to pivot, when an important hypothesis falls.
It is not set in stone. It evolves, but only at the pace of validations and invalidations. That is what sets it apart from a roadmap that shifts with every request. It also frames the MVP: the MVP is simply the test of the phase 1 hypotheses.
In hindsight
I wrote this article in 2022, still shaped by my work at SipScience. Of this series, it is probably the one whose substance I would change least, and the one I use most today, including outside product work.
What I would keep
The hypothesis format with a measurable signal, and ordering by risk. I have since seen teams lose whole quarters because the riskiest hypothesis was also the most uncomfortable to test, and had been pushed back.
What I would nuance
I presented the hypothesis roadmap as a tool for finding market fit. It goes well beyond that: for an established organization, every transformation project also rests on hypotheses, and they are rarely written down. Today I use it when scoping AI projects far more often than at startups.
I would also nuance the place of feasibility. In 2022, I readily put it first, as in the restaurant example. For an AI project, that is no longer always right: a working prototype is quick to build, and the real risk shifts to adoption and cost at scale.
What AI assistants have changed
The cost of testing has dropped. A landing page, a clickable prototype, an analysis of a hundred interviews take a few hours. So you can test more hypotheses, earlier, and that is real progress. For a company torn between an off-the-shelf tool and a custom build, the guide off-the-shelf software or a custom application applies the same test-before-you-commit logic.
But a quick test is still a test, with its biases. A prototype generated in an hour can charm people in an interview without proving anything about real use. And a summary of interviews produced by an assistant must be checked against the interviews themselves; otherwise the assistant is validating your hypotheses for you. Cheaper does not mean less rigorous.
Frequently asked questions
How many hypotheses should you test at once?
As few as possible: one or two per cycle, the ones everything else depends on. Testing ten hypotheses in parallel makes the results impossible to interpret.
When should a hypothesis be considered invalidated?
When the signal set in advance is not reached within the planned timeframe. The threshold must be written down before the test; otherwise you always end up finding a reason to carry on.
How is this different from OKRs?
OKRs set objectives and key results to achieve. The hypothesis roadmap explains what has to be true to achieve them, and in what order to check it. The two complement each other.
How do you order hypotheses that all look like priorities?
Rate each one from 1 to 3 for the impact if it is false and for the current uncertainty, and multiply. At equal risk, test first the one whose test costs least. The scoring is mostly useful for surfacing disagreements.
Who decides that a hypothesis is validated?
The owner named on the card presents the result, but the decision follows the threshold written before the test. If the threshold is met, you carry on; if not, you adjust or abandon, and the roadmap reflects it.
Where does AI fit in?
AI projects are, by nature, bets on hypotheses. Will the model reach sufficient quality on your real data? Will the teams use it? Will the cost per request hold at real volume? Most AI POCs only test the first question, and on a favourable sample.
Apply the same method. Write down the desirability, viability and feasibility hypotheses for your AI project, with a measurable signal for each. The metrics I detail in measuring an AI support assistant are concrete examples. And order them by risk: often the riskiest hypothesis is not technical but adoption by the teams. That is what I call the organizational 70%, and it is the subject of from AI POC to production.
If the system falls into a sensitive area, add a compliance hypothesis from phase 1: “this system does not fall into the high-risk categories, or we know how to meet the obligations”. You test it by reading Annex III, as I describe in classifying your AI systems under the AI Act.