We scraped more than 2 billion keyword ranks for Amazon retail media optimisation (PPC) software. Here's what I learned.
I started this as a data provider. By the end I'd almost become a PPC GOD, because to serve the best real-time data I had to learn every edge case these tools live and die by. Here are the ones that stuck.
⦿ Organic rank is THE datapoint every PPC software should track
Not impressions, not ACOS in isolation — organic rank. It's the honest scoreboard for every move you make.
⤷ What happened to your ranks the next day / week / month after that bid, budget or image change?
⤷ Did organic rank improve? Stay the same? Get worse?
If your PPC tool can't answer that, it's optimising blind.
⦿ One PPC move on a single keyword can move thousands of others
Keywords aren't isolated. Push one and the ripple hits the rest of the catalog.
⤷ How many DIFFERENT keywords' organic rank did this one PPC change impact?
⤷ Which single keyword has the most impact on the rest of your keywords' organic rank?
Find that keyword and you've found your lever.
⦿ Ask if you can hold the rank without PPC at all
⤷ If yes — cut the PPC and see if it stays on top.
⤷ If it doesn't — what's the minimum PPC budget you can keep to maintain organic rank at that spot?
Most people never test this. They keep paying for a position they already own.
⦿ Watch for ad cannibalization
If both your organic AND sponsored rank are 1 (both at the top):
⤷ You're wasting money on a sale you'd have gotten organically.
⤷ This is called ad cannibalization.
⤷ STOP PPC aggressively for keywords once they reach this state.
⦿ Sponsored ranks are competitive HOURLY, let alone daily
⤷ If you do day parting, check the ranks in advance.
⤷ Don't blindly trust yesterday's data — every day is a new fight.
⤷ Spot the hour with the least PPC competition where you can make the most sales.
Why this needs accurate, real-time data
Every one of these edge cases falls apart if the underlying rank data is slow, stale or incomplete. Sponsored placements load dynamically and disappear on a single scrape, so you need retry logic that captures ALL of them. Ranks change by the hour, so you need speed. And a rank is meaningless if it's read from the wrong location, so zip code has to be a locked control variable.
That's the whole reason I obsess over accuracy, residential IPs (not datacenter), and retry mechanisms — unlike clueless general scrapers that won't serve the same accuracy, speed or reliability. When rank is tied directly to dollars, wrong data is worse than no data.
