Discrete, process & regulated manufacturing

Inference on the plant floor, never on the PLC

We build inference systems that read from your MES, SCADA and historian. Rockwell FactoryTalk, Siemens Opcenter, Ignition, AVEVA InTouch, PI — over OPC-UA and MQTT Sparkplug B, sitting DMZ-side of your level 2 control network.

We never write to the PLC.

3

ISA-95 levels we work across: 2 read, 3 infer, 4 train

6

plant-floor workloads we build

0

synchronous writes into a control loop

Stacked steel pipe
What we build

Six workloads, wired into the systems you already run

Each one reads from the historian and MES and hands its answer back to a person or a work order. None of them closes a loop on a regulated process without human approval.

01

Predictive maintenance with RUL prediction

Vibration spectra at 1-25.6 kHz from accelerometers, motor-current signatures, bearing pass-frequency detection. A hybrid model — physics-based fault classifier plus a survival-analysis RUL head — lands work-order recommendations in the MES, never in the PLC.

02

Visual quality inspection on conveyor lines

Two-stage detection (presence, then localisation), tuned to the recall floor the line requires on safety-critical defects with false-reject held to a rework lane. Shadow mode for two to four weeks before the model is ever given the trigger.

03

Golden-batch deviation across ISA-88 records

Phase-segmented scoring against the recipe model: multivariate PCA and PLS for regulated pharma, auditable and validatable under GAMP 5; a variational autoencoder for non-regulated chemicals where the envelope is richer.

04

OEE root-cause from PLC tag history

Six-big-losses decomposition from raw tags and shift logs, with downtime causes auto-classified from the operator comment field by a model trained on that plant's own vocabulary.

05

Multi-tier traceability and genealogy

A graph linking serial to work-order to operator to machine to process tags to component lots to supplier lots — so when a defect spikes, the blast radius is a query rather than a week.

06

Operator-assist agents

Grounded in your own standard operating procedures, surfaced in the HMI operators already trust, and read-only against the validated systems of record.

Where the line is

Level 2 stays untouched

Inference belongs at level 3. Training belongs at level 4. Real-time control belongs to the people who already own it.

A predictive-maintenance model reads the historian and publishes a work-order recommendation; a yield model runs against ERP data and posts a suggestion back. We do not co-locate inference on the PLC and we do not write synchronously into control loops. That boundary is checkable, which is what makes it worth more to your OT team than any assurance.

The data layer

Pipelines modelled to ISA-95 levels 2 to 4

B2MML payloads where you want them, IEC 62443 segmentation throughout, and an honest audit of what your historian is actually keeping before anyone promises a model.

OPC-UA + MQTT Sparkplug B

OPC-UA south of the DMZ with Basic256Sha256 and mutual X.509; Sparkplug B over an MQTT broker with mTLS for high-cardinality edge-to-cloud — unified namespace, birth and death certificates, tagged ACLs per gateway.

MES and ERP read-side integration

FactoryTalk ProductionCentre, Siemens Opcenter, GE Proficy, AVEVA System Platform, Plex; SAP S/4HANA PP/EWM/QM via CDC or BAPI/OData, Oracle Fusion via REST — write-back only into MES work orders, never to a PLC.

Historian access

PI via the Web API or AF SDK, AVEVA Historian, Ignition Tag Historian — with an information-content audit before we quote, so historian compression is not quietly hiding the frequencies your model needs.

Time-series and waveform processing

Parquet on object storage for tag history, a fast store for hot windows, edge capture at 1-25.6 kHz for vibration, and matrix-profile or transformer encoders for anomaly detection.

Genealogy, BOM and batch records

A multi-tier supply graph, BOM hierarchies from the PLM system, and ISA-88 batch records as B2MML for golden-batch modelling.

IEC 62443 zone-and-conduit deployment

Training at level 4 in your own cloud account, inference at level 3, one-way replication from level 2 historians, OT security alerts piped to your SIEM, and no public-internet egress for the inference stack.

A CPU socket on a motherboard

Engineering questions

Asked by plant and OT teams, answered plainly

Inference belongs at level 3 (MES / plant operations), occasionally bridged through the level 3.5 DMZ to a level 4 (ERP / cloud) training cluster. Level 2 (PLC / SCADA real-time control) stays untouched - we never co-locate inference on the PLC and never write back synchronously to control loops. A predictive-maintenance model reads from the historian (level 3), publishes a work-order recommendation to FactoryTalk ProductionCentre or Opcenter (level 3), and lets the CMMS push it to operators. A demand-forecast or yield-optimisation model runs at level 4 against the S/4HANA or Oracle Fusion data and posts adjusted MRP suggestions back. Anything that closes a loop on a regulated process passes through human approval first.

OPC-UA is the right choice when you need rich type information, historical access (HA service set), and explicit security contexts - we run Basic256Sha256 with mutual X.509 certs issued from the plant CA, with audit logging on the OPC-UA server. MQTT Sparkplug B is the right choice for wide-area, low-bandwidth, high-cardinality edge-to-cloud flows - its Birth/Death certificates and unified namespace (UNS) topology make it easier to handle thousands of intermittently-connected devices. We typically run OPC-UA south of the DMZ (level 2 to level 3) and bridge to Sparkplug B for cloud-side analytics, with TLS 1.2+ and mTLS on the broker (HiveMQ or EMQX) and tagged ACLs per edge gateway. Both modes are wrapped in IEC 62443 zone-and-conduit policy.

IEC 62443 zone-and-conduit thinking maps cleanly onto the Purdue model. We place training infrastructure at level 4 (or in the customer VPC), the inference and feature store at level 3, and a data-diode or one-way replication into the level 3.5 DMZ from level 2 historians. Conduits are firewalled, certificate-pinned, and monitored - typically Claroty or Nozomi on the OT side feeds anomaly alerts into the same SIEM. Models never call out to the public internet; updates arrive as signed offline bundles, scanned at the boundary, and registered in a private MLflow. Latency budget for advisory inference is in the 100-500ms range, which fits the segmentation without contortion.

Any AI output that contributes to a GMP decision - golden-batch pass/fail, in-process release, deviation flag - becomes a Part 11 electronic record. That means ALCOA+ (attributable, legible, contemporaneous, original, accurate, complete, consistent, enduring, available) with audit trails on every change, secure e-signatures, and validated change control. We pin model versions, hash the feature vectors, and write an immutable audit log into the eBR system - MasterControl, Veeva Vault QualityDocs, or a plant historian configured for Part 11. The model itself is validated as a Computer System under GAMP 5: risk-categorised, IQ/OQ/PQ executed, with a continued process verification (CPV) plan tied to FDA Process Validation Stage 3. Pharma never auto-acts; the system recommends and a qualified person approves.

Vibration-based bearing health and RUL prediction needs raw or near-raw waveform data, not aggregated historian tags - typically 1-25.6 kHz sampling from accelerometers, captured by an edge box (NI cDAQ, OSIsoft PI Connector for high-speed, or an Ignition Edge with the Sparkplug module). For motor-current signature analysis, 10-100 Hz is enough. The PI / Ignition tags that most plants log at 1 Hz with compression are fine for OEE and process anomaly detection, but useless for incipient bearing-fault detection - you cannot see a 2 kHz bearing pass frequency in 1 Hz data. We do an information-content audit before quoting a predictive-maintenance project: if the historian compression is dropping the relevant frequencies, we install or commission edge capture first and start the model later.

It is a Pareto curve, not a single number. We start by labelling 5-10k images per defect class with the plant QA team, train a two-stage detector (a fast classifier for present/absent, then a localiser for defect type and bounding box), and calibrate the threshold against a labelled holdout. On a typical electronics or food-packaging line we target recall tuned to the floor the line requires on safety-critical defects (cracked seal, foreign object) - that comes at the cost of a false-reject rate around 1-3%, which is acceptable because rejects go to a rework lane, not the bin. For cosmetic defects we tune the other way: lower recall, near-zero false-reject, because over-rejecting blows up the scrap rate and operators will disable the system within a week. The model goes live in shadow first - it flags, the human inspector still decides - for two to four weeks before we hand over the gate.

A golden-batch model learns the multivariate envelope of the recipe phases - temperatures, pressures, agitation, addition timings - from a curated set of historically good batches. We pull the batch genealogy from the ISA-88 batch executor (DeltaV Batch, Rockwell PlantPAx, Siemens SIMATIC BATCH) as B2MML or via the OPC-UA historical access service, segment by recipe phase, and fit either a multivariate PCA / PLS model (the regulated-pharma default - explainable, easy to validate under GAMP 5) or a more flexible variational autoencoder for non-regulated chemicals. At runtime each new batch is scored phase-by-phase; deviations beyond the Hotelling T-squared or SPE control limits surface to the batch-record reviewer with the exact tags driving the deviation. For commercial pharma the model is an aid, not an arbiter - release-by-exception still requires a qualified-person signature against the eBR.

A traceability graph links every finished serial number back through its work-order, the operator and machine that built it, the in-process parameters from the historian, the component lots consumed (per BOM), and from there to the upstream supplier lots - across however many tiers the customer can give us data for. When a defect is found in the field, the query is 'which serials share this combination of supplier lot, machine, shift, and process-parameter band?' - instead of recalling six months of production we recall the affected sub-population, often two orders of magnitude smaller. We build the graph in Neo4j or TigerGraph for relational manufacturers, or in a columnar store with materialised joins where the volumes are higher (automotive Tier 1, semiconductor). For regulated industries the graph is read-only from validated source-of-truth systems (MES, ERP, LIMS) - we do not become the system of record, we become the system of inference on top of it.

Bring the tag history and the historian audit

Our delivered manufacturing engagement is generative AI and optimization work inside an automotive-parts manufacturer's Dynamics 365 Supply Chain environment; the plant-floor stack above is capability we build to, not a list of shipped rollouts, and we will say which is which on the call. Indicative delivery windows: an MES integration is 6 to 12 weeks; an OPC-UA or historian connector with a feature store, 3 to 5 weeks; a conveyor defect-detection model from shadow to live, 8 to 14 weeks; a golden-batch model with a GAMP 5 validation pack, 10 to 16 weeks.
Our second practice

This is our AI engineering practice

It is real work and it is where our four products came from. But what Cognilium leads with is narrower: optimization apps that run in tandem with Microsoft Dynamics 365, computing the decisions the ERP records but does not derive — the optimal price, the optimal pick path, the optimal stock level. See the optimization apps · How we build inside the ERP.