Use a two phase framework. Phase 1 sizes the market against a floor, scores the rating gap, then buys 200 units for $5,000 to $10,000 to test up to four candidates at once. Phase 2 scales only once rating, conversion rate, and cost of acquisition are proven on live sales.
The short version
- Two million dollars a year is the market floor. Below that there is not enough revenue to capture profitably once acquisition cost is paid.
- Research ends in a purchase order, not a spreadsheet. Desk work narrows the field, live sales decide.
- Test four candidates, not one. Comparison is what makes a result readable.
- Phase 1 is a cost line, so budget it. Between $5,000 and $10,000 per candidate, and it buys evidence rather than revenue.
- The framework is only as good as its rejections. If nothing gets cut, you are not running a framework.
Start with the number that disqualifies most ideas
Two million dollars in annual category revenue is where I stop looking. Under that line, the arithmetic stops working before you make a single mistake: the addressable pool is too shallow to pay for the customers you have to buy, and a realistic share of a small market does not cover the fixed costs of running a brand.
Work it through. Take a category doing $800,000 a year. A strong new entrant might take eight percent of it inside a year, which is $64,000 in revenue. Strip out the referral fee, fulfillment, landed cost, storage, and the advertising it took to get there, and what is left will not pay for the photography, let alone the inventory. There is nothing wrong with the execution in that example. The market was too small to reward it.
Above the floor the picture inverts. A market at $6 million a year does not need you to dominate it. A modest slice, sold at a defensible margin, pays for stock, ads, and the next product. Almost every product research framework you will read starts with the product. Start with the size of the room instead.
The framework, phase by phase
Phase 1: buy evidence
The output of Phase 1 is not a decision to launch. It is a set of three numbers you cannot get from any tool: what the product is rated after real buyers use it, what percentage of visitors convert on your listing, and what one customer actually costs to acquire.
To get them you place a deliberately small order. Two hundred units, landed between $5,000 and $10,000, on up to four candidates at once. Four small answers are more useful than one expensive guess, because they tell you which of your assumptions traveled and which did not.
Phase 2: scale what earned it
Phase 2 begins only when all three signals point the same way. Rating holding, conversion rate stable past the noise of the first fortnight, and acquisition cost inside the margin. Then you order properly, widen the traffic, and build the catalog around the winner.
If the three signals disagree, that is a fix conversation or a kill conversation, and it is worth naming which one. A weak rating with strong conversion is a product problem. Strong rating with weak conversion is a listing and pricing problem. Both healthy but acquisition cost too high usually means the market is more expensive than your desk research suggested.
What the research phase actually costs
| Line | Typical range | What it buys you |
|---|---|---|
| Desk research and market sizing | Your time, or a paid analyst's | The floor test and the shortlist |
| Patent and category checks | Low, mostly time | Legal room before you commit tooling |
| Samples from three suppliers | A few hundred dollars | Proof the fix is manufacturable |
| Phase 1 inventory, 200 units | $5,000 to $10,000 landed | Rating, conversion rate, acquisition cost |
| Photography and listing build | Varies by category | A listing that can be fairly judged |
| Phase 1 advertising | Budget for a real test window | A readable acquisition cost |
Read that table as the price of information. Every line exists to remove one guess, and a single production run committed on a guess costs more than all six lines combined.
The rule that keeps the framework honest
Write your pass mark down before you look at candidates, and record every rejection with the gate it failed. A framework that never rejects anything is a preference dressed up as a process.
This is also the part that gets skipped when research is outsourced. At Flapen we will not put a quote in front of a seller before the market has been sized, because a quote implies the work is worth doing and we do not know that yet. Ask any firm you are considering to size your market before they price their service. The order of those two steps tells you what you are buying.
What most agencies will not tell you
Most product research sold as a service is a tool export with a logo on it. The data underneath is available to anyone with a subscription, which means the recommendation you receive is the same recommendation several hundred other sellers can generate this week. The value was never in the numbers. It is in the rejections, the supplier quote against a specific complaint, and the willingness to say the category is too small.
The second thing: a research service paid per report has no stake in whether the product works. A firm that will later have to launch the thing behaves differently, because it inherits the consequence of a bad pick. When you compare providers, ask who owns the outcome twelve months from now. That question changes the answers you get.
Related answers
- Product criteria checklist for Amazon private label
- How to validate Amazon product demand fast
- Top ways to brainstorm Amazon product ideas
- Recommend tools to estimate Amazon startup costs
- Amazon seller roadmaps and capital: the complete guide
We run this framework for every brand we take on, and you can see the method at Flapen.

