Tag: synthetic data gdpr compliance

  • Synthetic Data, Real Decisions: How UK Firms Are Training AI Without Touching Customer Records

    Synthetic Data, Real Decisions: How UK Firms Are Training AI Without Touching Customer Records

    There is a problem sitting at the centre of almost every serious AI project in the UK right now. The models need data. Enormous quantities of it, labelled carefully, representative of edge cases, and reflective of real-world complexity. But the data that would do the job best is also the data that regulators, customers and common sense say you should not be feeding into a training pipeline. Patient records. Transaction histories. Purchase behaviour linked to identifiable individuals. For a long time, this felt like an impasse. It is not one any more.

    Synthetic data AI training in the UK has moved from a niche academic concept to something financial services firms, healthcare technology companies and large retailers are deploying in production. The pitch is straightforward: generate artificial datasets that carry the statistical properties of real data without containing any actual personal information. The model trains on the synthetic version, performs almost as well as it would on the real thing, and the ICO has far less to complain about. I have spent time talking to people actually building these systems, and the reality is more textured than the pitch suggests, but the commercial logic is genuinely solid.

    Server racks in a data centre representing synthetic data AI training UK infrastructure
    Photo by panumas nikhomkhai on Pexels

    Why the ICO’s data minimisation rules are forcing the issue

    The ICO’s guidance on AI and data protection is explicit: you should not collect or process more personal data than is strictly necessary for your stated purpose. Training a fraud detection model on five years of raw, personally identifiable transaction records when a synthetic equivalent would serve almost as well is exactly the kind of thing that makes an ICO audit uncomfortable. The ICO published its guidance on AI and data protection and it is not ambiguous on the data minimisation obligation under UK GDPR. Firms building AI products need to document why they are using personal data at all, and increasingly the answer “because we had it” is not going to cut it.

    This is where synthetic data enters as a technical compliance mechanism, not just a nice-to-have. If your training corpus is synthetic, the data minimisation argument largely resolves itself. You are not processing personal data in the training phase at all. The thornier question, which I will come to, is how you generate the synthetic data in the first place without that generation process becoming a data protection issue in itself.

    What financial services firms are actually building

    UK financial services firms are arguably the furthest along. Banks and insurers have spent years accumulating transaction data at scale, and they have equally spent years being told by the FCA and ICO that they cannot simply do whatever they like with it. The use case that comes up most often is fraud detection. Training a classifier to spot anomalous transactions requires examples of fraudulent behaviour, but real fraud examples are rare and involve real victims. Synthetic generation lets teams oversample rare events, creating artificial fraud scenarios with the right statistical signature without exposing a single actual customer.

    Lloyds Banking Group and Barclays have both referenced synthetic data programmes publicly in the context of their AI development work. The approach is not just about compliance; it also solves the class imbalance problem that plagues fraud models, where genuine fraud cases might represent 0.1% of your dataset. You can generate as many synthetic fraudulent transactions as you need to train a robust model. That is a genuinely useful engineering property, not just a regulatory fig leaf.

    Credit risk modelling is another active area. Lenders want models that reflect the behaviour of underserved customer segments, people who have thin credit files or non-standard income patterns. Real data on these groups is scarce by definition. Synthetic generation, trained on the limited real data that does exist, can expand the training set to cover edge cases the model would otherwise never encounter.

    Healthcare tech: where the stakes are highest

    In healthcare technology, the tension between data utility and data sensitivity is sharper than anywhere else. NHS patient records are among the most valuable datasets in the world for training diagnostic AI, and they are also subject to strict access controls, Caldicott Guardian oversight and a public that has, fairly or unfairly, grown suspicious of how its health data gets used.

    Several UK health tech firms working with NHS trusts are now using synthetic patient data generated from real cohorts to train triage algorithms and predictive models for readmission risk. The process typically involves training a generative model, often a variational autoencoder or a GAN variant, on a legitimately accessed, properly consented real dataset, then using that generative model to produce a much larger synthetic population. The synthetic population retains the correlations that matter clinically, age-related comorbidity patterns, demographic risk factors, treatment response distributions, without being traceable to any real individual.

    Sensyne Health, before its restructuring, was an early UK example of this approach. Smaller firms like Featurespace and Babylon (in its NHS work) have also grappled with the same challenge. The ICO has not yet published specific guidance on synthetic health data, but the general principle that synthetic data used correctly sits outside the definition of personal data gives firms a defensible position. I would caveat that by saying the “used correctly” part requires genuine technical rigour; a poorly generated synthetic dataset can leak real records through membership inference attacks, and that is not a hypothetical risk.

    Retail’s quieter adoption

    Retail is a less obvious context but a growing one. Large UK retailers, Tesco, ASOS and Marks and Spencer among those who have discussed AI investment publicly, are building recommendation engines, demand forecasting models and customer lifetime value predictors. All of these benefit from rich behavioural data. All of them also involve processing data about identifiable individuals.

    The specific use case I find most interesting is synthetic data for model testing. Rather than running A/B tests on live customers with a half-baked model (which carries both commercial and regulatory risk), retailers can generate synthetic customer cohorts that mirror their actual customer base and stress-test model behaviour against them first. This is less about training from scratch and more about safe iteration, but it represents the same underlying logic: keep real customer data out of the development loop wherever you can.

    My take is that retail adoption is being driven less by regulatory pressure and more by practical data access problems. Loyalty card data lives in one system, online browsing in another, in-store transaction data in a third. Stitching these together in a way that is both technically and legally clean is genuinely hard. A synthetic unified customer record, generated from the real datasets but not itself a personal record, sidesteps a lot of that complexity.

    The generation problem you cannot ignore

    The uncomfortable bit that advocates for synthetic data sometimes gloss over is this: you cannot generate high-quality synthetic data without first processing real data. That processing step is still subject to UK GDPR. You still need a lawful basis, you still need to minimise retention, and you still need to document the purpose. The synthetic generation is not a magic circle that makes everything that happened before it disappear from a regulatory standpoint.

    This is why I think the firms getting this right are treating synthetic data generation as a distinct data processing activity with its own data protection impact assessment, rather than as an end-run around compliance. The ones getting it wrong are treating it as a loophole, and when the ICO looks more carefully at AI pipelines, those are the firms that will have problems. For a deeper look at what ICO auditors are actually asking about in AI contexts, the piece on what UK businesses are being asked to prove during ICO AI audits is worth reading alongside this.

    There is also the question of synthetic data quality. A dataset that does not accurately reflect the statistical properties of real-world data will produce a model with poor generalisation. The validation step, checking that a model trained on synthetic data performs comparably on real data, requires access to real data for evaluation. You have not escaped the need to handle personal information; you have just moved it to a different stage of the pipeline.

    Tools UK teams are actually using

    The tooling landscape has matured considerably. Mostly AI, a UK-rooted company now operating internationally, produces one of the more widely used synthetic tabular data platforms. Gretel.ai has a UK customer base across financial services. For health data specifically, NVIDIA’s MONAI framework includes synthetic generation capabilities that some NHS-adjacent teams have adopted. Open-source options like SDV (Synthetic Data Vault) from MIT are popular with engineering teams who want full control over the generation process.

    The choice between proprietary platforms and open-source tools tends to come down to the same build-versus-buy question facing UK engineering teams across the board. Given the sensitivity of health and financial data, some teams are uncomfortable sending even the training data for the generative model outside their own infrastructure, which pushes them toward self-hosted open-source solutions. That is a reasonable position. I covered the broader GDPR considerations for UK firms using synthetic data to train AI in more detail previously if you want the full regulatory picture. And if your firm is at the earlier stage of thinking about how financial modelling and data feeds into investor-facing work, the piece on how UK founders are using financial modelling tools to satisfy data-hungry investors gives useful context on where data discipline starts to matter commercially.

    Synthetic data is not a universal answer. But for UK firms trying to build AI products that work in regulated sectors, it is fast becoming a serious and necessary part of the technical architecture. The firms treating it that way, rigorously and with proper documentation, are building something defensible. The ones treating it as a shortcut will find out soon enough that it is not one.

    Frequently Asked Questions

    What is synthetic data and how does it differ from anonymised data?

    Synthetic data is artificially generated data that mirrors the statistical properties of a real dataset without containing any actual records from real individuals. Anonymised data starts with real personal data and has identifying information removed; synthetic data is created fresh by a generative model trained on real data, so no real record is ever directly included in the output.

    Does using synthetic data mean UK firms can ignore GDPR requirements entirely?

    No. The process of training a generative model to produce synthetic data still involves processing real personal data, which is subject to UK GDPR. Firms need a lawful basis for that initial processing, a data protection impact assessment, and documented retention limits. The synthetic output itself sits outside the definition of personal data, but the generation pipeline does not.

    How accurate are AI models trained on synthetic data compared to real data?

    Accuracy depends heavily on the quality of the synthetic data generation. Well-generated synthetic tabular data can produce models within 2-5% of the performance achieved on real data for tasks like fraud detection or credit risk scoring. The validation step, testing the model on real data, remains essential and is where quality issues tend to surface.

    Which UK sectors are using synthetic data AI training most actively right now?

    Financial services is the most mature adopter, particularly for fraud detection and credit risk modelling. Healthcare technology firms working with NHS data are also active, using synthetic patient cohorts to train diagnostic and triage algorithms. Retail is growing, especially for recommendation and demand forecasting models.