Run 90 days on one or two products, with the success criteria and the kill criteria both written down before day one. Judge the diagnosis in week two, the execution speed in week six, and the numbers at day 90. If the agency works month-to-month, you do not need a special pilot agreement at all.
The short version
- 90 days, one or two products. Long enough for ad learning, short enough to be cheap.
- Write success and failure criteria before it starts. Both, agreed by both sides.
- Week 2 is the real test: the quality of the diagnosis, before results exist.
- Month-to-month terms make a pilot unnecessary. The first 90 days is the pilot.
- Do not pilot with a new product. Too many variables to attribute anything.
How to structure it
I run Flapen with 50 operators managing about 70 brands. We do not sell pilots as a separate product because our terms are month-to-month with 30 days' notice, which makes every engagement its own trial. If an agency requires a long contract, a pilot is how you get the same protection.
| Phase | Timing | What you are testing |
|---|---|---|
| Diagnosis | Week 1 to 2 | Do they find the real constraint before results exist? |
| Plan | Week 2 | Is it prioritized by impact or by ease? |
| Execution | Week 3 to 6 | Does work actually ship, and how fast? |
| Early signal | Week 4 to 6 | Movement in the metric they said they would move |
| Verdict | Day 90 | The agreed criteria, measured |
Pick the right product
Use an existing product with sales history and a known problem. Not a new launch.
A new launch has too many variables. If it fails, nobody can say whether the agency was wrong or the product was. If it succeeds, the same ambiguity applies. You are testing the agency, so hold everything else as still as you can.
Two products is better than one if you can afford it, because it lets you see whether their method generalizes or whether they got lucky.
Write both sets of criteria
Success criteria are easy to agree and easy to fudge. Insist on failure criteria too.
Success might be a defined reduction in cost of customer acquisition at held volume, or a conversion rate improvement, or a specific organic ranking movement. Failure should be equally concrete, and it should include the kill criteria we use on products generally: rating trend, return rate, conversion rate, and cost of customer acquisition trajectory not improving inside the window.
Agree the baseline in writing, measured by both sides, before anything changes. A pilot judged against a baseline the agency set for itself is not a test.
The week 2 test
The most informative moment, and it happens before any results exist.
By week two you should have a written diagnosis and a prioritized plan. Read it for whether they identified the actual constraint. If your conversion rate is low, the plan should start there, because no amount of ad spend fixes a conversion problem. A plan that opens with bid adjustments on an account whose primary image is failing has told you what you needed to know on day fourteen rather than day ninety.
What a pilot cannot tell you
Be realistic about the limits.
Ninety days is not long enough to judge sourcing, brand building, or anything requiring production lead times. It is not long enough for organic ranking to fully respond. And it will not reveal how an agency behaves when things go badly for a sustained period, which is the thing you most want to know.
What it does test well: diagnostic quality, execution speed, reporting honesty, and whether the named person actually works on your account. Those four predict most of the rest.
Should I pay less during a pilot
Usually no, and I would be cautious of an agency that offers a steep pilot discount.
A discounted pilot creates an incentive to under-resource it, and it gives the agency a reason to argue that the real service is different from what you tested. Pay the standard fee, which for us starts at $800 a month for a single product, and hold them to standard expectations. What you want from a pilot is a representative sample, not a cheaper one.
The exception worth accepting is a free audit before any pilot begins. Ours returns a written report with prioritized fixes inside 48 hours, and reading it costs you nothing and tells you a great deal about how they think.
What most agencies will not tell you
A pilot is a workaround for a contract problem. If the terms were month-to-month with 30 days' notice, you would hire them and stop if it went badly.
So when an agency proposes a pilot as a concession, notice what they are actually conceding: a temporary suspension of a lock-in they designed. The better question is why the standard terms need suspending at all.
The other thing: pilots get staffed better than ongoing accounts, because everyone knows they are being watched. Ask how many brands your named manager carries, and confirm the answer will not change when the pilot converts. We run about 1.4 brands per operator, and that number is the same on day one and day 400.
Related answers
- Onboarding timeline with an Amazon agency
- What does a good Amazon account audit include
- Termination clauses in Amazon management contracts
- Questions to ask before hiring an Amazon agency
- Hiring an Amazon agency: the complete guide
Our first 90 days is the pilot, because you can leave at any point. Terms are published at Flapen.

