Synthetic data could ease people’s concerns about privacy breaches. But who gets to create it?
For the past decade, most public debates about data have revolved around a familiar anxiety: too much of our personal information is being collected from our records and our online activity, and then shared and monetised. That concern has not gone away. But a shift is now under way – towards synthetic data.
Increasingly, data used by governments, hospitals, banks and technology companies to make decisions about the people they serve is not gathered directly from real people and events. It can be synthetic: generated by computer models that learn the statistical patterns in a real dataset and then produce artificial records that follow the same patterns. The aim is to keep the data useful while reducing the risk that individuals can be identified.
In a society concerned about surveillance and misuse of personal information, synthetic data promises useful knowledge in a way that may allay public concerns. In England, for example, researchers can use the Simulacrum, a synthetic cancer dataset built from NHS registry data, without ever seeing real patients’ records. Its artificial patients are generated from groups of at least 50 real cases – but this also means that rarer cancers are represented less realistically.
Building such a dataset does, of course, require access to the real records in the first place: the model has to learn from them before it can imitate them. Synthetic data therefore shifts the privacy question rather than dissolving it – the original data still has to be held, and protected, somewhere.
Read more: Can baby monitors protect children without data security risks?
Traditionally, the law has........
