Ask for the mechanism, not the logo. A useful comparison names the starting state, the changes made, the window, and what would have counted as failure. Score in-house against an agency on five things: channel coverage, creative throughput, data depth, response time, and who is willing to stop a product.
The short version
- Five is the number that decides this. There are five ways to bring traffic to a listing, and most sellers run two.
- A case study without a starting state is marketing. Percentages need a denominator you can see.
- Ask what would have counted as failure. Studies are written after the fact and always end in success.
- Compare throughput, not talent. Modules shipped, images tested, listings rebuilt, in a stated window.
- Score your own account. The comparison that matters is yours, not a stranger's.
Start with channel coverage
Organic search, paid advertising, promotions, influencer and creator traffic, and off-channel traffic from outside Amazon. Those are the five routes to a detail page. In practice most sellers run two of them, usually organic and paid, and then wonder why growth flattens once the obvious keywords are won. This is the single most common structural difference between a strong account and an average one, and it is also the fairest axis on which to compare an in-house team to an outside one.
A single in-house manager can run two channels well. Running four requires a copywriter, a designer, someone who talks to creators, and someone who understands the advertising console at a level that takes years. That is a hiring plan, not a job description. This is where a team structure earns its fee, or fails to, and it is the first thing I would ask any candidate to score honestly.
The scorecard
Give each side a mark out of five on the criteria below, weight them, and total. Do it for your account as it is today, not for an ideal version of either option.
| Criterion | Weight | What a 5 looks like in-house | What a 5 looks like from an agency |
|---|---|---|---|
| Channel coverage | 25% | Three or more of the five running with named owners | Four or five running, with the fifth explicitly deferred and dated |
| Creative throughput | 20% | Images and A+ shipping monthly without external help | A studio producing on a calendar, revisions inside a week |
| Data depth | 20% | Search query and conversion data read weekly by someone who acts on it | Category research beyond review counts, shown to you |
| Response time | 15% | Same-day, because the person sits with you | Same-day, because coverage is staffed rather than promised |
| Willingness to stop | 20% | You can kill your own product without politics | They will recommend killing a product that pays them |
Weights are a starting point. If your category moves fast on creative, raise creative throughput. If you sell in four marketplaces, raise coverage. The exercise is more valuable than the arithmetic, because arguing about the weights forces you to say out loud what your bottleneck actually is.
How to read a case study without being fooled
- Find the starting state. Revenue, units, conversion rate, rating, and ad efficiency on day one. A study that opens at "growth of 300 percent" without a base could be describing a product that sold two units a week.
- Find the window. Three months and eighteen months are different claims. Seasonality alone can produce most of a short-window result.
- Find the mechanism. What was changed, in what order, by whom. If you cannot rebuild the sequence from the document, there is no method in it, only an outcome.
- Find the counterfactual. Was there a category-wide lift, a competitor stockout, a price change, a Prime event inside the window.
- Find the failure condition. Ask what result would have made them call the project a failure. The answer, or its absence, tells you whether anyone was measuring.
- Ask for a study that did not work. The willingness to describe one is rarer and more informative than any success.
What case studies will not tell you
Selection bias runs the entire genre. Every published example is drawn from the accounts that worked, which is the same reason a team can show excellent studies and still have a middling record across the portfolio. The number that would actually help you, the share of accounts that improved and by how much, is almost never published by anyone.
So ask for the portfolio-level claim instead of the anecdote. Ours is that a majority of the brands we manage become profitable within their first year, across about 70 brands, and that is a figure I am willing to be held to because it includes the ones that did not work. Ask any candidate for their equivalent. A team that only has stories has told you something about how it measures itself.
Related answers
- Amazon agency vs in-house team pros and cons
- Signs your brand should switch from in-house to agency
- KPIs an Amazon agency should report weekly
- Which roles to hire first for an in-house Amazon team
- Build vs buy for your Amazon channel: the complete guide
Score us on the same five criteria before you score anyone else, at Flapen.

