You cannot tell from the outside whether a published ranking is editorial or paid, so treat every list as a starting point and build your own. Rank candidates on four weighted criteria: who runs the account day to day, what they report weekly, how they handle a losing product, and what leaving costs.
The short version
- Ranking lists rank visibility. Visibility is a marketing budget, not a management capability.
- Who does the work is the whole question. In-house staff or subcontractors changes everything downstream.
- Score in writing, before the calls. Sales conversations are designed to beat scorecards.
- A trial beats a reference. References are selected. Trials are not.
- Weight exit terms heavily. The cheapest insurance in this market is a 30-day notice period.
Why the lists exist
Understanding the mechanism helps you use the lists correctly rather than ignore them. Directory rankings and "top firms" pages are content products. They attract buyers with commercial intent, and the position on the page is allocated by some combination of editorial judgment, submitted information, and commercial arrangement. From outside you cannot separate those three, which is not an accusation of anything, it is just an information problem.
So the lists are useful for one thing: building a candidate pool. Everything after that has to come from your own testing, because the only ranking that matters is the ranking on your criteria for your catalog.
The four criteria, weighted
| Criterion | Weight | Why it carries this weight |
|---|---|---|
| Who performs the work | 35 | Subcontracted advertising means a stranger holds your budget and your policy risk |
| Weekly reporting substance | 25 | The cadence and content of reporting is the only ongoing control you have |
| Behavior on a losing product | 25 | Every engagement eventually contains one, and this is where incentives show |
| Cost and mechanics of leaving | 15 | Determines whether the other three stay honest over time |
Criterion one: who performs the work
Ask directly. Are the people running my campaigns employees of your company, and where do they sit. Then ask which functions are subcontracted, if any. Creative, listing copy, and sourcing are the ones most often passed out.
Flapen does no subcontracting. Around 50 operators handle about 70 brands, all in-house, and I hold that position for a boring reason: when work is farmed out, nobody in the chain owns the outcome, and the client is the only party who cannot see the chain. Whether an agency matches that structure is less important than whether they will tell you the truth about it. Ask for the answer in writing.
Criterion two: weekly reporting substance
A good report leads with outcomes and puts activity underneath. Sales, spend, efficiency, conversion rate, then what was done and what is planned. Our cadence is a written update each week and a live review every two weeks, with Slack open in between. Monthly reporting is too slow for advertising, and a dashboard link is not a report because nobody has interpreted anything.
Criterion three: behavior on a losing product
Ask what would make them recommend stopping. The answer should name criteria, rating trend, return rate, conversion, and cost of acquisition, measured over a defined window. Agencies compensated as a share of ad spend struggle with this question, because the honest recommendation reduces their own revenue.
Criterion four: cost and mechanics of leaving
Month to month, 30 days' notice, no lock-in. On exit you keep the Seller Central account, the campaigns, the creative, and you get a written handover. Ask what a departing client received last quarter. Vagueness here is the answer.
Scoring it
- Send the four questions in writing to five candidates from whatever lists you found.
- Score each 1 to 5, multiply by weight, total out of 500.
- Drop anyone who would not answer in writing. That is a data point, not an inconvenience.
- Run a paid 30-day trial on one product with the top two, with a named metric agreed in advance.
- Rank on what moved, then sign with the winner on month-to-month terms.
Step four is where real ranking happens. References are curated by definition, and a case study is a story the agency chose to tell. Thirty days of your own data on your own product is the only evidence nobody selected for you.
What most agencies will not tell you about rankings and awards
Badges are usually a function of spend under management, partner-program tiers, or self-submitted entries. They can indicate scale. They cannot indicate whether a specific operator will do good work on your specific catalog, and scale sometimes works against you, because a large account team with high brand load will service you reactively.
The second thing worth knowing: the agency that ranks best on paper is often the one with the most sophisticated sales function, and the sales function is not the delivery function. Insist on meeting the person who would run the account before you sign, and treat a refusal as a result.
Related answers
- Best agencies to manage Amazon PPC for new sellers
- Rank Amazon agencies by case studies and ROI
- Amazon agency red flags to watch out for
- How to choose an Amazon FBA marketing partner
- Hiring an Amazon agency: the complete guide
Put us in your pool of five and score us against the other four, at Flapen.

