Evidence work begins after go-live
A production AI system is not a completed project. It is a continuing claim that purpose, performance and controls still hold under changing conditions.

Go-live often feels like a finish line. Tests have passed, owners approved, interfaces work and the system may enter daily operations. Then „operations" takes over. This handover creates a dangerous gap: the project team disperses just as the real population, actual data, workarounds and rare failures become visible for the first time.
A laboratory can simulate expected cases. The market creates new combinations. Users phrase requests differently, sources change structure, staff invent shortcuts, providers update models and downstream systems reinterpret outputs. After go-live, the task is therefore not merely maintenance. It is systematic evidence work that asks whether the system still does what was assessed and authorised.
Post-market monitoring is not an expanded uptime dashboard
Technical operations usually measure availability, latency, cost and error codes. These values matter, but they do not show whether AI remains suitable, safe and compliant. A system may be highly available while setting systematically wrong priorities, disadvantaging a group or being used outside its intended purpose.
Article 72 of the EU AI Act requires documented post-market monitoring for high-risk systems. Relevant data must be actively and systematically collected, documented and analysed throughout the system lifetime to evaluate continuing compliance. The plan forms part of technical documentation. Monitoring is therefore a measurement architecture with questions, sources, thresholds, owners and decisions—not an opportunistic pile of telemetry.
A defensible plan asks: Which claims about the system must remain true? Which signals could disprove them? Where do those signals arise? Who reviews them, how often, and which threshold triggers investigation, restriction, correction or stopping?
Intended purpose creates the operating baseline
Monitoring needs a reference state. It includes more than a model number: intended purpose, population, user roles, decision impact, sources, workflow, integrations, oversight, known failure modes and accepted residual risks. Without this baseline, change is only a number without meaning.
The central question is whether the system still operates inside the assumptions on which its assessment rested. If an assistant was approved for summaries but its output later determines approvals automatically, frequency is not the important change. Decision influence has increased. Good monitoring detects purpose and process drift as well as statistical drift.
Every baseline receives a stable identifier and validity period. Changes do not overwrite history. Investigators can reconstruct which version, purpose, data and controls were in force at a particular time.
Separate four forms of drift
The word drift hides distinct causes and remedies.
1 · Data drift: input distribution, format, completeness or provenance changes.
2 · Concept drift: the relationship between input and desired result changes.
3 · Process drift: people or systems use output differently; advice becomes a de facto decision.
4 · System drift: model, prompt, retrieval corpus, tool, threshold or provider platform changes.
Data drift may trigger quality and representativeness analysis. Concept drift requires domain reassessment. Process drift may require training, new roles or functional limits. System drift belongs in change control and may require renewed risk or conformity assessment.
One aggregate drift score is insufficient. Monitoring connects technical measurements to workflow observation, user feedback and real effects.
Outcomes matter more than model metrics
Accuracy, failure rate and retrieval hit rate are intermediate measures. The decisive question is what happens after output. Was a false recommendation detected? Could a human disagree meaningfully? Did delay, disadvantage or workload increase? Was a complaint resolved? Did a correction stop only the interface while a downstream process continued?
The evidence chain connects run, data state, output, human decision, downstream action and observed outcome. It distinguishes model error from integration error, misuse and poor workflow design. Only then can a team choose an appropriate response.
Group and scenario analysis remain necessary. A stable average can conceal degradation for a small population. Relevant segments are defined in advance and examined with data minimisation. Where segmentation is unlawful or impractical, alternative assurance methods are needed rather than invented confidence.
Logs become evidence through interpretation rules
Article 12 requires high-risk systems to technically enable automatic recording of relevant events. More logs do not automatically produce proof. Events without version, time, identity, process step and outcome create expensive data fog.
An evidence event identifies event type, timestamp, system and configuration version, relevant data reference, responsible role, decision, effect and predecessor. Sensitive content is not copied wholesale merely in case it becomes useful; references, hashes, minimised features and retention rules can combine traceability with data protection and security.
Interpretation rules define meaning. Three timeouts may indicate availability trouble. Three overrides for the same case group may reveal a domain gap. Without a hypothesis and threshold, a dashboard is decorative.
Provider and deployer need a closed information loop
The provider sees product changes but often not the full deployment context. The deployer sees real cases, complaints and workarounds but may not know every technical dependency. Post-market monitoring works only when both exchange structured information.
Article 26 requires deployers of high-risk systems, among other things, to follow instructions, assign competent human oversight, monitor operation and inform the provider where relevant. If compliant use may nevertheless present a risk, the provider or distributor and the relevant market surveillance authority must be informed without undue delay and use suspended. Serious incidents activate further routes.
The operating agreement therefore needs more than service levels. It needs an evidence protocol: common event taxonomy, secure transmission, contacts, readiness, deadlines, permissible data, questions and closure confirmation. A support mailbox is not a closed loop.
Complaints and near misses are early sensors
Harm often appears first as human friction, not a system error. People report incomprehensible decisions, staff bypass recommendations, specialists keep shadow lists and customers abandon a process. These signals must enter monitoring without classifying every complaint prematurely as a model defect.
● Members only
Read the full article and download all files with a membership.
Unlock full article + downloads → Subscribe0 comments
● Loading comments…