Collection that does not break
Most scrapers break when a site changes its markup or starts blocking robots. Ours are built so they don't, and this page is exactly how — tier by tier, with what each one costs.
6
tiers, climbed only as far as needed
10M+
products a week at peak, on our own infrastructure
50+
cloud machines in that fleet

Six tiers, and the discipline to stop early
Each tier costs more than the one below it, so the skill is stopping as soon as the data comes back clean. Starting at the top is how collection budgets get burned.
IP and user-agent rotation
Residential and datacenter addresses in rotation with honest pacing. This alone handles roughly 80% of targets, and it is where every engagement starts.
Headless and headful browsers
Real browsers on cheap horizontal infrastructure, for sites that need a page to render before there is anything to read.
Scalable browser grids
For rendered applications and infinite scroll, where the data arrives after the page and the session has to stay alive to get it.
Fingerprint rotation
Purpose-built anti-detect browsers varying screen, timezone, CPU class, fonts, plugins, even GPU and sound-card signals — because a fingerprint that never changes is itself a signal.
Paid challenge solving
Where a challenge is unavoidable it is answered through a paid solver rather than pretended away. It costs money per solve and we say so when we scope it.
The AI tier
Real machines on residential connections running agents that behave like a person browsing, for the hardest targets. Costly, but it scales and it is predictable.
Getting the data once is a demo
Getting it every day for a year is the engagement.
Parsers that adapt when a target rewrites its markup overnight; jobs that resume from the last processed record rather than restarting after a crash; validation before anything is indexed, so a silently empty field is caught as a fault instead of stored as a fact. On one delivered engine that discipline held 98% scrape success across 500,000 indexed products.
Collection is the first half
The price-intelligence pipeline
Our own product: price, stock, ratings and reviews from roughly twenty major retailers on a daily cycle, validated, indexed and served by API.
See it →Data engineering
The foundations underneath — queues, validation, indexing and the recovery that keeps it all running when a source misbehaves.
See it →Send us the site that keeps blocking you
This is our AI engineering practice
It is real work and it is where our four products came from. But what Cognilium leads with is narrower: optimization apps that run in tandem with Microsoft Dynamics 365, computing the decisions the ERP records but does not derive — the optimal price, the optimal pick path, the optimal stock level. See the optimization apps · How we build inside the ERP.