Reported ROI is not comparable across firms, because every published figure hides the starting point, the category, and the budget behind it. Rank them yourself with a weighted scorecard covering channel coverage, operator load, incentive alignment, evidence quality, and exit terms. Score the shortlist before the sales calls, never after.
The short version
- A percentage without a baseline is a decoration. Doubling a small number is easy and tells you nothing about your case.
- Channel coverage is the single largest driver of upside. There are five ways to put traffic on a listing and most sellers run two of them.
- Score before you meet anyone. A weighted table survives a persuasive call. A good impression does not.
- Weight the criteria yourself. My weights suit a growth-stage private label brand and yours may not.
- Every score needs evidence attached. A claim you cannot check scores zero, including mine.
Why ROI figures behave the way they do
Return on a consulting engagement is a ratio between two numbers that the firm controls neither of. The denominator is your fee plus your ad spend. The numerator is incremental profit, which depends on your landed cost, your category, your review position, your stock availability, and what your competitors do that quarter.
That is why case study percentages cluster so high. A brand that was doing very little before will show enormous growth from basic competence, and a brand already run well will show single digits from excellent work. The firm that reports 400 percent may have inherited more slack. Ranking on those numbers rewards whoever found the worst-run accounts.
So rank on the inputs that produce return instead. There are five of them, they are observable during a sales process, and they are difficult to fake.
The scorecard
Score each firm from 1 to 5 on each criterion, multiply by the weight, and total. Change the weights to match your situation before you start.
| Criterion | Weight | What a 5 looks like | What a 1 looks like |
|---|---|---|---|
| Channel coverage | 30 | Runs organic, paid, promotions, creator, and off-channel, and can show recent work in each | Paid search only, described as full-service |
| Operator load | 20 | Names the person, states how many accounts they carry | Vague about who does the work |
| Incentive alignment | 20 | Flat fee, no commission on your ad spend | Paid a percentage of what you spend |
| Evidence quality | 20 | Before-and-after with baselines, category, and budget disclosed | Screenshots without context |
| Exit terms | 10 | Month to month, you keep everything, written handover | Twelve months, auto-renewal, penalty on exit |
A useful test of your own weights: if the winner of your scorecard is not the firm you liked most on the call, sit with that for a day before overriding it. The scorecard is the part of your judgment that was not in the room with a salesperson.
Channel coverage carries the most weight for a reason
There are five ways to get traffic to a listing: organic search, paid advertising, promotions and deals, influencer and creator content, and off-channel traffic driven from outside Amazon. In my experience most sellers run two of them, usually organic and paid, and then wonder why growth flattens once those two are optimized.
That is where genuine upside sits, and it is also where most agencies are thinnest, because the other three require different skills. Promotions need margin modeling. Creator work needs briefing, sourcing, and content review. Off-channel needs landing pages, tracking, and patience. None of it is exotic, but it is real operational work, so ask each firm which of the five they ran for a client in the last month and what it produced. Answers get specific quickly or they do not.
Our own position, since you should apply the test to me too: at Flapen 50 operators cover about 70 brands, and channel activation is one of the seven things the audit checks before anyone touches a bid. Ask us the same question you ask everyone else.
Scoring evidence without being fooled
- Ask for the baseline. Revenue, ACoS, conversion rate, and review count at the start of the engagement. No baseline, no score.
- Ask what else changed. A price cut, a new video, or a competitor going out of stock can explain most of a result.
- Ask for a failure. Every firm has one. The ones who describe it precisely are the ones who studied it.
- Ask for the client's own words. Not a logo wall. A sentence a named person is willing to stand behind.
What most agencies will not tell you
The biggest lever on your return is often not the agency at all. It is your landed cost. A brand with a healthy margin can afford to advertise aggressively, discount tactically, and fund creator work, while a brand with a thin margin has almost no room whatever the firm does. Any consultant who talks only about tactics and never about your cost per unit is optimizing the second-order problem.
The second thing: the shortest path to a better return is sometimes to have fewer products. Cutting a weak ASIN releases capital, attention, and advertising budget into the products that already work. It also reduces the agency's fee if they charge by product count, which is exactly why the advice is rare.
Related answers
- KPIs an Amazon agency should report weekly
- Which service to boost Amazon PPC
- Top Amazon seller consultants ranked 2026
- Global Amazon SEO specialists with proven case studies
- Done-for-you Amazon management: the complete guide
Score us against your own weights, and start with the free audit at Flapen.

