SAKIZLI AI
Article28 Jul 2026 · 16 min read22 / 23Members · Subscription

Evidence work begins after go-live

A production AI system is not a completed project. It is a continuing claim that purpose, performance and controls still hold under changing conditions.

AI ActGovernanceObservabilityRisk
FFurkan SakızlıAI researcher & tutor · independent
A bright modular AI system operates beyond a transparent launch gate inside a continuous blue monitoring loop; measurement points, version markers and one amber incident signal lead into a controlled corrective loop
A production AI system is a continuing claim — not a completed project

Go-live often feels like a finish line. Tests have passed, owners approved, interfaces work and the system may enter daily operations. Then „operations" takes over. This handover creates a dangerous gap: the project team disperses just as the real population, actual data, workarounds and rare failures become visible for the first time.

A laboratory can simulate expected cases. The market creates new combinations. Users phrase requests differently, sources change structure, staff invent shortcuts, providers update models and downstream systems reinterpret outputs. After go-live, the task is therefore not merely maintenance. It is systematic evidence work that asks whether the system still does what was assessed and authorised.

Post-market monitoring is not an expanded uptime dashboard

Technical operations usually measure availability, latency, cost and error codes. These values matter, but they do not show whether AI remains suitable, safe and compliant. A system may be highly available while setting systematically wrong priorities, disadvantaging a group or being used outside its intended purpose.

Article 72 of the EU AI Act requires documented post-market monitoring for high-risk systems. Relevant data must be actively and systematically collected, documented and analysed throughout the system lifetime to evaluate continuing compliance. The plan forms part of technical documentation. Monitoring is therefore a measurement architecture with questions, sources, thresholds, owners and decisions—not an opportunistic pile of telemetry.

A defensible plan asks: Which claims about the system must remain true? Which signals could disprove them? Where do those signals arise? Who reviews them, how often, and which threshold triggers investigation, restriction, correction or stopping?

Intended purpose creates the operating baseline

Monitoring needs a reference state. It includes more than a model number: intended purpose, population, user roles, decision impact, sources, workflow, integrations, oversight, known failure modes and accepted residual risks. Without this baseline, change is only a number without meaning.

The central question is whether the system still operates inside the assumptions on which its assessment rested. If an assistant was approved for summaries but its output later determines approvals automatically, frequency is not the important change. Decision influence has increased. Good monitoring detects purpose and process drift as well as statistical drift.

Every baseline receives a stable identifier and validity period. Changes do not overwrite history. Investigators can reconstruct which version, purpose, data and controls were in force at a particular time.

Separate four forms of drift

The word drift hides distinct causes and remedies.

1 · Data drift: input distribution, format, completeness or provenance changes.

2 · Concept drift: the relationship between input and desired result changes.

3 · Process drift: people or systems use output differently; advice becomes a de facto decision.

4 · System drift: model, prompt, retrieval corpus, tool, threshold or provider platform changes.

Data drift may trigger quality and representativeness analysis. Concept drift requires domain reassessment. Process drift may require training, new roles or functional limits. System drift belongs in change control and may require renewed risk or conformity assessment.

One aggregate drift score is insufficient. Monitoring connects technical measurements to workflow observation, user feedback and real effects.

Outcomes matter more than model metrics

Accuracy, failure rate and retrieval hit rate are intermediate measures. The decisive question is what happens after output. Was a false recommendation detected? Could a human disagree meaningfully? Did delay, disadvantage or workload increase? Was a complaint resolved? Did a correction stop only the interface while a downstream process continued?

The evidence chain connects run, data state, output, human decision, downstream action and observed outcome. It distinguishes model error from integration error, misuse and poor workflow design. Only then can a team choose an appropriate response.

Group and scenario analysis remain necessary. A stable average can conceal degradation for a small population. Relevant segments are defined in advance and examined with data minimisation. Where segmentation is unlawful or impractical, alternative assurance methods are needed rather than invented confidence.

Logs become evidence through interpretation rules

Article 12 requires high-risk systems to technically enable automatic recording of relevant events. More logs do not automatically produce proof. Events without version, time, identity, process step and outcome create expensive data fog.

An evidence event identifies event type, timestamp, system and configuration version, relevant data reference, responsible role, decision, effect and predecessor. Sensitive content is not copied wholesale merely in case it becomes useful; references, hashes, minimised features and retention rules can combine traceability with data protection and security.

Interpretation rules define meaning. Three timeouts may indicate availability trouble. Three overrides for the same case group may reveal a domain gap. Without a hypothesis and threshold, a dashboard is decorative.

Provider and deployer need a closed information loop

The provider sees product changes but often not the full deployment context. The deployer sees real cases, complaints and workarounds but may not know every technical dependency. Post-market monitoring works only when both exchange structured information.

Article 26 requires deployers of high-risk systems, among other things, to follow instructions, assign competent human oversight, monitor operation and inform the provider where relevant. If compliant use may nevertheless present a risk, the provider or distributor and the relevant market surveillance authority must be informed without undue delay and use suspended. Serious incidents activate further routes.

The operating agreement therefore needs more than service levels. It needs an evidence protocol: common event taxonomy, secure transmission, contacts, readiness, deadlines, permissible data, questions and closure confirmation. A support mailbox is not a closed loop.

Complaints and near misses are early sensors

Harm often appears first as human friction, not a system error. People report incomprehensible decisions, staff bypass recommendations, specialists keep shadow lists and customers abandon a process. These signals must enter monitoring without classifying every complaint prematurely as a model defect.

A complaint record connects channel, case reference, system version, impact, relevant group, investigation, response and correction. Repeated patterns are aggregated. Near misses—cases where oversight or luck prevented harm—remain visible. Counting only realised harm discards the most valuable prevention data.

Incident response starts before an incident

Article 73 establishes tiered reporting deadlines for serious incidents. Generally, reporting follows immediately after a causal link or its reasonable likelihood is established and no later than 15 days after awareness. Certain particularly severe cases have shorter maximum periods; an incomplete initial report may preserve timeliness and be completed later.

These periods are not targets for internal action. Detection, triage, evidence preservation and activation must happen earlier. A runbook asks: Who can disable the system? Who preserves logs without contaminating investigation? Who assesses impact on people? Who informs provider, deployer, authorities and other parties? Who documents why a signal was not reportable?

Uncontrolled changes must not compromise later causal evaluation. Configurations, versions and relevant artefacts are frozen or reproducibly preserved while people are protected and further harm prevented.

Every change needs an evidence gate

AI systems change more frequently than many conventional products. A provider updates the base model, a retrieval corpus grows, a prompt is adjusted, a tool gains permission or a source is added. Each change may alter performance, risk, transparency, oversight and purpose.

Article 17 explicitly includes modification procedures, testing before, during and after development, monitoring, incident reporting, record keeping and accountability within the quality management system. A ticket saying „minor prompt improvement" is insufficient. The gate identifies affected claims, new failure paths, data changes, regression tests, documentation, rollback and approval.

Impact, not file size, determines the change class. One new tool permission may be riskier than a large internal refactor. A model version enters production only after comparison data, boundary cases, oversight and reversal have been tested.

Corrective action is more than a patch

Where a high-risk system is not conforming, Article 20 requires immediate appropriate corrective action. Depending on circumstances, this may mean bringing it into conformity, withdrawal, disabling or recall. Relevant distributors, deployers, representatives, importers, authorities and possibly notified bodies must be informed.

A corrective-action record describes the issue, scope, affected versions and deployments, immediate protection, root-cause analysis, lasting action, validation, communication and closure criteria. It also checks downstream consequences: Must decisions be reviewed, people informed or data products regenerated? A software patch does not heal an effect that already occurred.

The smallest useful monitoring architecture

A small provider does not need a giant control room. A workable minimum contains six linked registers:

Baseline: purpose, version, population, data, controls and residual risk.

Signals: metrics, thresholds, complaints, drift and near misses.

Changes: modification, impact, tests, approval and rollback.

Incidents: triage, evidence, reporting, investigation and protection.

Actions: correction, owner, deadline and effectiveness review.

Review calendar: daily alerts, periodic reviews and event-triggered reassessment.

The technology can be simple. Stable identifiers and relationships matter most. A complaint must lead to the run, version, change and corrective action.

Method: BASELINE → SIGNAL → TRIAGE → CHANGE → VERIFY

BASELINE fixes the valid system claim. SIGNAL gathers technical and human observation. TRIAGE assesses impact, urgency and reporting. CHANGE applies a bounded correction or adaptation. VERIFY determines whether action and documentation actually address the cause.

The final step closes the loop. Without effectiveness review, monitoring becomes a ticket factory. With it, operations becomes a controlled learning system that converts negative evidence into better boundaries and decisions.

Worksheet: Design a post-market monitoring plan

1. Define purpose, version, population and decision impact.

2. State five claims that must remain true in operation.

3. Assign technical, human and outcome signals to each claim.

4. Define thresholds for review, pause, correction and incident triage.

5. Design the information loop between provider and deployer.

6. Specify the change and regression-test gate.

7. Connect complaints, near misses and logs to actions.

8. Plan review cadence, ownership and effectiveness checks.

Reflection: Which system claim can your current dashboard not test? Which real change could currently enter production unnoticed?

All materials to download — the topic overview and the worksheet:

Scope: This article presents a professional operating and evidence model, not legal advice. The cited requirements particularly concern high-risk AI systems; risk-proportionate elements may be useful voluntarily elsewhere. Sectoral, data protection, employment, product and national law require separate assessment. Editorial review date: 17 July 2026.

Members only

Read the full article and download all files with a membership.

Unlock full article + downloads → Subscribe

0 comments

Loading comments…

Sign in to comment · become a member →