Governance
GPT-Red: Governing AI That Changes Itself
Self-improving models break the assumption behind every AI register entry and DPIA: that the system you assessed is the system running today. A practical change-control approach for compliance teams.
Almost every piece of AI governance documentation in circulation describes a system as it behaved on the day somebody looked at it. The register entry, the data protection impact assessment, the vendor due-diligence pack and the sign-off email are all snapshots. A system marketed as self-improving — one retrained on its own outputs, tuned against automated adversarial testing, or silently updated by a supplier between your review cycles — quietly invalidates that snapshot on a schedule you do not control. The question for a risk owner is therefore not "is this model robust?" but "robust as measured when, by whom, against what, and would anyone here notice if it stopped being?"
For most mid-size UK and EU organisations this is not an abstract concern, because you are rarely the party doing the improving. You are a deployer of somebody else's model. The improvement happens on their infrastructure, on their timetable, described in release notes you may not receive. What you need is not a better one-off model evaluation. It is a change-detection and change-response process that treats model behaviour as a live thing to be monitored, rather than a fact to be documented once and filed.
What "self-improvement" actually covers
The label attached to a particular system — GPT-Red is one of several names given to this design direction — matters far less than the mechanism underneath it. In practice, teams encounter at least five distinct things bundled under the same phrase, and they carry very different governance weight.
- Automated adversarial testing feeding back into training. The system is attacked by another system, failures are collected, and the model is hardened against them. Genuine robustness gains, but the test set is generated rather than agreed, so nobody outside the vendor knows what "hardened" now means.
- Continuous or online learning. The model updates from live interaction data. Rare in commercial deployments, and the highest-risk category, because behaviour drifts without any release event to hang a control on.
- Vendor-side version changes. By far the most common. The weights behind an endpoint change; your integration code does not. This is the case most registers handle worst.
- Configuration and scaffolding changes. A revised system prompt, an added tool, a new guardrail, a changed refusal threshold. No model change at all, yet the observable behaviour of the system your staff use can shift substantially.
- Retrieval and knowledge changes. The corpus the model draws on is updated. Outputs change, provenance changes, and the personal data in scope may change with them.
Only the first two involve the model improving itself in any meaningful sense. All five change what a user experiences, and all five can move a system across a compliance boundary. Registers that track only a model version number will miss three of them entirely.
Why a changing system breaks the paperwork you already hold
Map each change type to a detection route, the evidence it undermines, and a proportionate response. Most organisations find the middle column is where the real gap sits.
| Change | Usually detected by | Evidence it undermines | Proportionate response |
|---|---|---|---|
| Vendor model version update | Nothing, unless subscribed to release notes | Accuracy and bias testing; DPIA behavioural findings | Contractual notice period plus a standing regression set |
| Retraining on your own data | Internal change request | Lawful basis and purpose limitation analysis | DPIA review before, not after, the retrain |
| Automated adversarial hardening | Vendor disclosure only | Robustness claims in the assurance pack | Ask for method and coverage, not a pass mark |
| System prompt or guardrail edit | Often no change control at all | Human oversight design; intended-purpose statement | Version control and approval on the prompt itself |
| Retrieval corpus change | Data owner, if anyone | Records of processing; transfer analysis | Treat the corpus as a documented data source |
What the rules actually expect
The EU AI Act does not treat risk management as a document. For high-risk systems it requires a continuous, iterative process maintained across the entire lifecycle, and it requires that systems achieve an appropriate level of accuracy, robustness and cybersecurity and perform consistently in those respects throughout that lifecycle. It also explicitly addresses systems that keep learning after being placed on the market, including the duty to mitigate feedback loops where a system's own outputs shape its future inputs.
Crucially, the Act provides a legitimate route for change rather than forbidding it. Changes a provider has pre-determined and set out in the technical documentation at the point of conformity assessment do not, on that basis, count as a substantial modification. That is the mechanism worth understanding: planned, documented change is anticipated; unplanned, undocumented change is what reopens the assessment. The Act also sets out when a deployer can inherit provider obligations — for example by putting its own name on a system, or by modifying it substantially or changing its intended purpose. If your team is fine-tuning and rebranding a supplier's model, that boundary deserves a legal read rather than an assumption.
Deployer duties are more modest but concrete: use the system in line with the provider's instructions, assign human oversight to people with the competence and authority to exercise it, monitor operation, inform the provider of risks and serious incidents, and retain automatically generated logs where they are under your control for at least the period the Act specifies. Providers of general-purpose models judged to carry systemic risk face additional expectations including model evaluation and adversarial testing. On penalties, the Act's maximums are up to 7% of global turnover for prohibited practices and up to 3% for most other obligations; UK and EU GDPR sit at up to 4%. Those ceilings are rarely the operative risk for a mid-size organisation, but they set the seriousness of the regime.
Data protection law does independent work here. The accuracy principle applies to personal data a changing model produces about people, not only to the data you fed it. A DPIA is expected to be reviewed when the risk represented by the processing changes — and a model that materially alters its behaviour is precisely that. Where outputs feed decisions with legal or similarly significant effects, the automated decision-making provisions and the associated safeguards apply regardless of how robust the vendor says the model has become.
In the UK there is no single cross-sector AI statute. Existing regulators apply existing law, which means UK GDPR, the ICO's guidance on AI and data protection, and — for regulated firms — sector expectations on outsourcing, operational resilience and model governance. If you also serve EU customers or place systems on the EU market, plan to the stricter of the two. On the standards side, ISO/IEC 42001 expects a management system with defined lifecycle controls, impact assessment, monitoring and internal audit; the NIST AI Risk Management Framework's Measure and Manage functions cover the same ground in a less certifiable form.
A checklist for governing a model that changes
- Does the register record the exact model identifier, version, and the date behaviour was last tested — not just the product name?
- Is there a contractual right to advance notice of model changes, and a defined notice period?
- Do you hold a standing regression set: real prompts from your own use cases, with expected and unacceptable outputs, that can be re-run on demand?
- Is the system prompt under version control, with a named approver for edits?
- Is there a documented rollback position — a pinned version, an alternative supplier, or a manual process — and has it been exercised?
- Are guardrails enforced in code after generation, rather than only instructed in the prompt?
- Do logs capture enough to reconstruct a specific decision months later, and are they retained for a defined period?
- Does someone own the question "has behaviour changed?" as a named responsibility with a review cadence?
- Have you asked the vendor how robustness is tested, what the coverage is, and what is explicitly out of scope — rather than accepting a summary score?
- Is there a trigger written down that forces a DPIA review, and does "material model change" appear in it?
- Do the people exercising human oversight have the authority and the time to reject an output, and is rejection recorded?
Where this guidance stops
This covers the deployer position: an organisation using AI systems built by others, with obligations centred on monitoring, oversight and documented change control. It does not cover what is required if you develop or substantially modify a high-risk system and take on provider obligations, and it does not tell you whether a given system is high-risk — that classification is a legal analysis of your specific use and intended purpose, not a judgement to be made from a checklist.
It also sits outside sector-specific regimes. Financial services, healthcare and medical devices, employment screening, education and law enforcement each carry additional expectations from their own regulators that can be more demanding than the general position described here. The EU AI Act's obligations phase in over a staged timetable, with national implementation and supervisory arrangements still settling in places, so confirm the position that applies to you rather than working from a general summary.
Take specialist legal advice where a decision has legal or similarly significant effects on individuals, where you are considering branding or fine-tuning a third-party model as your own, where international transfers are involved, or where a change has already occurred and you are assessing whether it should have been notified. Nothing above is a substitute for that. What it should do is make sure that when the question arrives, you can answer it with records rather than recollection.
- EU AI Act
- Model risk
- Change control
- Red teaming
- Robustness
- DPIA