Collection that does not break
Most scrapers break when a site changes its HTML or starts blocking robots. Ours are built so they don't — and this page is exactly how, tier by tier, including what each tier costs.
6
tiers of escalation, used in order
10M+
products a week at peak, on our own infrastructure
50+
cloud scraping machines in that fleet

There are two kinds of site, and the price differs
Before anyone quotes a number, the targets get sorted. Which bucket a site falls into decides the tier it needs, and the tier decides what it costs to run every day for a year.
Simple
Static HTML, or an application backed by an API you can read directly. Rotation and pacing are usually the whole solution, and roughly 80% of targets sit here.
Complex
Rendered applications, infinite scroll, sessions and logins, and sites that actively look for robots. These need browsers, fingerprint discipline, and sometimes the top tier.
Six tiers, climbed only as far as needed
Each tier costs more than the one below it, so the discipline is to stop as soon as the data comes back clean. Starting at the top is how collection budgets get burned.
IP and user-agent rotation
Residential and datacenter addresses in rotation, with honest request pacing. This alone handles roughly 80% of targets, and it is where every engagement starts.
Headless and headful browsers
Real browsers on cheap horizontal infrastructure, for sites that need a page to actually render before there is anything to read.
Scalable browser grids
For JS-rendered applications and infinite scroll, where the data arrives after the page does and the session has to stay alive to get it.
Browser-fingerprint rotation
Purpose-built anti-detect browsers varying screen, timezone, CPU class, fonts, plugins, even GPU and sound-card signals — because a fingerprint that never changes is itself a signal.
Paid CAPTCHA solving
Where a challenge is unavoidable, it is answered through a paid solver rather than pretended away. It costs money per solve, and we say so when we scope it.
The AI tier
Real laptops on residential connections running generative agents that behave like a person browsing, for the hardest targets. Costly, but it scales and it is predictable.
Sites change their HTML overnight
Getting the data once is a demo. Getting it every day for a year is the engagement.
Our parsers adapt when a target rewrites its markup, jobs resume from the last processed record rather than starting over after a crash, and validation runs before anything is indexed — so a silently empty field is caught as a fault rather than stored as a fact.
What has actually been built
Four collection systems, described the way the record allows — the engineering named, the clients not.
An auto-parts search engine
500K+ products indexed with 98% scrape success, Redis queues, async retries and JSON sanitisation throughout.
An e-commerce data pipeline
Millions of URLs moving through an S3 → Redis → OpenSearch path, built so a crash resumes rather than restarts.
Marketplace ASIN tracking
Rate-limit-aware batching against a marketplace data API, because the limit is part of the design rather than an error to handle later.
A storefront ingestion bot
Scrape, clean the listing with a model, then publish through the storefront's own API — collection and publication as one pipeline.
Built with
Send us the site that keeps blocking you
This is our AI engineering practice
It is real work and it is where our four products came from. But what Cognilium leads with is narrower: optimization apps that run in tandem with Microsoft Dynamics 365, computing the decisions the ERP records but does not derive — the optimal price, the optimal pick path, the optimal stock level. See the optimization apps · How we build inside the ERP.