Data engineering · Our own infrastructure

Collection that does not break

Most scrapers break when a site changes its HTML or starts blocking robots. Ours are built so they don't — and this page is exactly how, tier by tier, including what each tier costs.

6

tiers of escalation, used in order

10M+

products a week at peak, on our own infrastructure

50+

cloud scraping machines in that fleet

Dew on a spider web
How we scope

There are two kinds of site, and the price differs

Before anyone quotes a number, the targets get sorted. Which bucket a site falls into decides the tier it needs, and the tier decides what it costs to run every day for a year.

Simple

Static HTML, or an application backed by an API you can read directly. Rotation and pacing are usually the whole solution, and roughly 80% of targets sit here.

Complex

Rendered applications, infinite scroll, sessions and logins, and sites that actively look for robots. These need browsers, fingerprint discipline, and sometimes the top tier.

The escalation

Six tiers, climbed only as far as needed

Each tier costs more than the one below it, so the discipline is to stop as soon as the data comes back clean. Starting at the top is how collection budgets get burned.

01

IP and user-agent rotation

Residential and datacenter addresses in rotation, with honest request pacing. This alone handles roughly 80% of targets, and it is where every engagement starts.

02

Headless and headful browsers

Real browsers on cheap horizontal infrastructure, for sites that need a page to actually render before there is anything to read.

03

Scalable browser grids

For JS-rendered applications and infinite scroll, where the data arrives after the page does and the session has to stay alive to get it.

04

Browser-fingerprint rotation

Purpose-built anti-detect browsers varying screen, timezone, CPU class, fonts, plugins, even GPU and sound-card signals — because a fingerprint that never changes is itself a signal.

05

Paid CAPTCHA solving

Where a challenge is unavoidable, it is answered through a paid solver rather than pretended away. It costs money per solve, and we say so when we scope it.

06

The AI tier

Real laptops on residential connections running generative agents that behave like a person browsing, for the hardest targets. Costly, but it scales and it is predictable.

The part that matters at month six

Sites change their HTML overnight

Getting the data once is a demo. Getting it every day for a year is the engagement.

Our parsers adapt when a target rewrites its markup, jobs resume from the last processed record rather than starting over after a crash, and validation runs before anything is indexed — so a silently empty field is caught as a fault rather than stored as a fact.

The record

What has actually been built

Four collection systems, described the way the record allows — the engineering named, the clients not.

An auto-parts search engine

500K+ products indexed with 98% scrape success, Redis queues, async retries and JSON sanitisation throughout.

An e-commerce data pipeline

Millions of URLs moving through an S3 → Redis → OpenSearch path, built so a crash resumes rather than restarts.

Marketplace ASIN tracking

Rate-limit-aware batching against a marketplace data API, because the limit is part of the design rather than an error to handle later.

A storefront ingestion bot

Scrape, clean the listing with a model, then publish through the storefront's own API — collection and publication as one pipeline.

The stack

Built with

PlaywrightPuppeteerBeautifulSouplxmlCrawleeResidential proxiesDockerKubernetesRedisPythonTypeScript

Send us the site that keeps blocking you

A working session on your actual targets: we sort them into simple and complex, tell you which tier each one needs, and what running that every day would honestly cost. If a target is not worth collecting, we will say so.
Our second practice

This is our AI engineering practice

It is real work and it is where our four products came from. But what Cognilium leads with is narrower: optimization apps that run in tandem with Microsoft Dynamics 365, computing the decisions the ERP records but does not derive — the optimal price, the optimal pick path, the optimal stock level. See the optimization apps · How we build inside the ERP.