Predictive Systems in Regulated Environments Need Evidence, Not Just Accuracy

Published on 2026-05-25

A correct prediction may still be indefensible

A model that correctly identifies an emerging equipment fault, a suspect batch condition or a clinically relevant pattern has demonstrated something useful. It has not necessarily demonstrated that its output can support a regulated decision.

In controlled environments, the relevant question is not only whether a prediction was accurate when assessed against a test set. It is whether the organisation can establish what data was used, how that data was obtained and processed, which model version generated the result, what controls were in effect, and whether the output was used within its validated intended purpose.

This distinction matters after an adverse event, inspection, deviation investigation or product-quality review. If a system recommends an intervention and the organisation cannot reconstruct the basis of that recommendation, it may be impossible to determine whether the issue arose from the underlying process, a sensor fault, an unapproved data change, a model defect or inappropriate operator use. A plausible output is not evidence of a dependable process.


Accuracy is a limited measure

Accuracy is often an incomplete metric even in conventional machine-learning work. For imbalanced events, a model can appear highly accurate while missing the uncommon cases that matter most. False-negative and false-positive rates, calibration of predicted probabilities, detection latency, confidence intervals and performance across operating conditions may be more consequential.

In a regulated application, accuracy also says little about the integrity of the path from physical process to model output. Consider a predictor using temperature, pressure and vibration data to identify developing process instability. The apparent model input may be a clean numerical time series, but its suitability depends on further conditions:

  • The sensor must be appropriate for the measured range and installed so that it represents the process of interest.
  • Calibration status, measurement uncertainty and any applicable traceability chain must be understood.
  • Timestamps must be reliable enough to preserve the relationship between signals, events and reference outcomes.
  • Data acquisition, buffering and transfer must detect loss, duplication, corruption and out-of-sequence records.
  • Transformations such as filtering, aggregation and feature extraction must be specified and versioned.

Without these controls, a model can be accurately predicting an artefact of the data pipeline rather than a property of the process. A drift in a sensor or a change in a historian tag may leave the algorithm operational while invalidating the meaning of its input.


Provenance must extend from measurement to decision

Provenance is often treated as a record of where a dataset originated. For a predictive system, it needs to be broader. The organisation should be able to associate a material output with its source records and the sequence of processing steps that produced it.

A practical evidence chain may include raw observations, acquisition time and source identity, calibration and maintenance status where relevant, data-quality flags, transformation software versions, feature definitions, model identifier, model configuration, inference environment, output timestamp and the identity of the receiving system or reviewer. It should also preserve relevant contextual information, such as operating mode, asset configuration, batch or lot association, and exclusions applied to the data.

This does not mean retaining every transient internal calculation indefinitely. Retention and review arrangements should be proportionate to intended use, risk and applicable record requirements. It does mean defining in advance the evidence needed to reproduce or explain a material output. Reconstruction should not depend on an engineer locating a retired notebook, an overwritten container image or an undocumented database query.

The security boundary is part of this design. Source data, model artefacts and decision records need controlled access, integrity protection and appropriate segregation of duties. If a user can alter training labels, feature logic or thresholds without independent review and a durable record, the model's apparent performance has little governance value.


Validation concerns the intended use, not the algorithm in isolation

A technically capable algorithm is not automatically fit for a regulated workflow. Validation needs to address the system's intended use and the decisions it informs.

For an advisory tool, the evidence may need to show that outputs are presented accurately, uncertainty and limitations are intelligible, users can review supporting information, and an appropriate human decision process remains in place. For a system that automatically initiates a process action, the required assurance is likely to be more stringent because a failure can directly affect product, safety or compliance.

Validation should exercise realistic operating conditions rather than only curated development data. This includes missing or delayed inputs, values outside the trained range, sensor substitutions, communications loss, unusual operating modes and degraded dependencies. The system should have defined behaviour when it cannot produce a reliable result: for example, suppressing the prediction, raising a data-quality alert or reverting to an established manual procedure.

Acceptance criteria should be specified before testing where practicable. They should cover more than predictive performance, including data integrity, access control, auditability, interface correctness, alarm behaviour and recovery. The applicable regulatory framework and organisational quality system determine the specific requirements; they should be confirmed against authoritative sources for the particular sector and jurisdiction.


Model change is a controlled change

Predictive systems change more readily than conventional rules-based software. A revised feature calculation, retraining dataset, changed decision threshold or updated dependency can alter behaviour materially without changing the user interface.

Each change therefore needs impact assessment. The assessment should determine whether prior validation remains applicable, what regression testing is required, whether the model's intended use has changed, and whether training and operating procedures require revision. Promotion between development, test and production environments should use controlled, identifiable artefacts rather than informal copies of files or live edits to configuration.

Online learning and automatic retraining require particular caution. Their appeal is that the model can adapt to new conditions. Their risk is that production behaviour changes faster than the organisation can review its evidence. In many regulated applications, a bounded update process with held-out evaluation data, independent approval and a defined rollback route is more defensible than continuous autonomous change.

Monitoring remains necessary after release. Performance must be assessed against reliable reference outcomes where available, but monitoring should also cover input drift, missingness, prediction distributions, data-pipeline failures and use outside approved conditions. A stable accuracy statistic can conceal a deteriorating measurement system or a narrowing population of reviewed cases.

Predictive capability becomes suitable for regulated use when it is treated as part of an engineered system of measurement, software, people and controls. The objective is not to eliminate uncertainty. It is to make the system's limits visible, its changes controlled and its material outputs supported by evidence that can withstand review.

Copyright © 2026 Obsidian Reach Ltd.

UK Registed Company No. 16394927

3rd Floor, 86-90 Paul Street, London,
United Kingdom EC2A 4NE

020 3051 5216