Service · Web scraping & price intelligence

Collection that does not break

Most scrapers break when a site changes its markup or starts blocking robots. Ours are built so they don't, and this page is exactly how — tier by tier, with what each one costs.

6

tiers, climbed only as far as needed

10M+

products a week at peak, on our own infrastructure

50+

cloud machines in that fleet

A fishing net with orange mesh and coloured rope
The escalation

Six tiers, and the discipline to stop early

Each tier costs more than the one below it, so the skill is stopping as soon as the data comes back clean. Starting at the top is how collection budgets get burned.

01

IP and user-agent rotation

Residential and datacenter addresses in rotation with honest pacing. This alone handles roughly 80% of targets, and it is where every engagement starts.

02

Headless and headful browsers

Real browsers on cheap horizontal infrastructure, for sites that need a page to render before there is anything to read.

03

Scalable browser grids

For rendered applications and infinite scroll, where the data arrives after the page and the session has to stay alive to get it.

04

Fingerprint rotation

Purpose-built anti-detect browsers varying screen, timezone, CPU class, fonts, plugins, even GPU and sound-card signals — because a fingerprint that never changes is itself a signal.

05

Paid challenge solving

Where a challenge is unavoidable it is answered through a paid solver rather than pretended away. It costs money per solve and we say so when we scope it.

06

The AI tier

Real machines on residential connections running agents that behave like a person browsing, for the hardest targets. Costly, but it scales and it is predictable.

The part that matters at month six

Getting the data once is a demo

Getting it every day for a year is the engagement.

Parsers that adapt when a target rewrites its markup overnight; jobs that resume from the last processed record rather than restarting after a crash; validation before anything is indexed, so a silently empty field is caught as a fault instead of stored as a fact. On one delivered engine that discipline held 98% scrape success across 500,000 indexed products.

Send us the site that keeps blocking you

We sort your actual targets into simple and complex, tell you which tier each one needs, and what running that every day would honestly cost. If a target is not worth collecting, we will say so.
Our second practice

This is our AI engineering practice

It is real work and it is where our four products came from. But what Cognilium leads with is narrower: optimization apps that run in tandem with Microsoft Dynamics 365, computing the decisions the ERP records but does not derive — the optimal price, the optimal pick path, the optimal stock level. See the optimization apps · How we build inside the ERP.