Sovereignty, sustainability, and single points of failure

I write this sitting in an airport on the way to Germany, where I will be meeting with leaders several European research agencies to discuss how we can ensure the resilience of infrastructure in an increasingly tumultuous world

Marcus Povey

Perhaps more than most biological disciplines, data is of critical importance for Structural biology, and data management has always been prioritised by members of the field through repositories and databases.

However, a PDB deposition is not the same thing as “the data is safe”, a deposition is the polished output, not the working material. Raw data, from both successful experiments and unsuccessful ones, matters, now, perhaps more than ever. Reprocessing methods improve, and the old raw data can yield new scientifically useful results years later; you lose that option if only the final model survives. Validation and reproducibility requires the raw data, as it is what lets someone actually check a result. With the adoption of AI, models need volume and diversity of training data, not just final curated depositions. Finally, this data needs to be placed in context, from the original scientific proposal, through raw data, to final deposition – a “cradle to grave” provenance chain (the kind we’re building at Instruct) – not just an isolated binary gathering digital dust on a disk somewhere.

Data storage is asymmetric with processed structures backed up in multiple locations (PDB mirrors globally), raw data usually exists with the user that did the research and, if you’re lucky, at the facility that performed the experiment. For some technologies, raw data is of small enough size (e.g. NMR) that mirrored storage is possible, for others (e.g. Cryo-EM or Xray) the sizes start getting problematic very quickly.

The Problem to the East, and the Problem to the West

As I write this, I am currently sat waiting for a flight to Germany, where I will be meeting with representatives from a number of other research infrastructures in order to discuss matters of data sovereignty, sustainability of critical infrastructure, and the single points of failure our communities have. The goal is to start putting together a coherent plan for addressing them, going into the next few years.

This is will be an open forum, and I’ve no doubt we will have many interesting discussions.

Specifically about raw data, I have been mulling how the current status quo of holding this data at the facility represents a single point of failure that is not often thought about. People typically didn’t give much thought to raw data, only the processed output, but now with AI these data are starting to become more valuable. Indeed, with instrument time coming at an expensive premium, reuse and harvesting existing data for new insights is becoming much more of a priority.

Recent history has brought about its own challenges that someone in my position, thinking about such high octane things as scientific data management, ever thought to have on their bingo card.

In June 2025, one of our facilities in Israel was hit by an Iranian ballistic missile strike, which destroyed years’ worth of research data. Member and partner facilities in the Baltics, Finland and Poland, are all watching the havoc wrought by Russia in Ukraine with stone faces, since the chance that the front line may come their way is far from zero. To the west, formerly stalwart partners have become increasingly erratic over a number of administrations, and while I’m not going to get drawing into the politics of it all, it is clear that the funding decisions made in foreign capitals can have very real world impacts closer to home.

Europe finds itself sandwiched between forces which stresses the need for increased redundancy, strategic depth and sovereignty within its borders. Redundancy that will take more than a single artillery shell, or spreadsheet deletion, to remove.

Bringing a buddy

The data volumes involved in some cases become substantial, often running into tens or even hundreds of TB per experiment in some cases, so getting someone to pay for everyone else’s storage is probably unfeasible for the time being. But perhaps storing some data for a friend?

A concept I’ve been pondering, and I should stress that this doesn’t in any way reflect Instruct’s policy position, is the idea of buddying. Imagine, a facility in Germany or France “buddies” with a similar facility in the Baltics and agree to mirror each other’s data (or at least the portion of it that is Instruct funded). This might be done on a similar technology basis with similar data quantities (X-Ray to X-Ray for example), or it might be done for wildly different technologies and leverage that disparity. Specifically, while an NMR or a Mass Spectrometry facility may produce a “lot of data” that storing it becomes a problem, this volume of data may be a rounding error when partnered with a synchrotron facility’s storage.

So long as the location and provenance of the data is tracked, any number of these bilateral agreements could be put in place, with money coming from national resilience budgets among other places. There’s also the question of consent, since not every facility will be comfortable with a partner holding a mirror of data that hasn’t been published yet, and that’s a policy problem to solve before it’s a technical one. ARIA already tracks data back to the original proposal, and FandanGO integration could be extended as a mechanism for cross-facility indexing, or existing pilot activities with EOSC and EOSC Beyond to broker storage provided by one or more EOSC nodes using this mechanism, already proves that in principle this could work from a technology point of view, so all that is needed is the policy decision to be layered on top.

In closing…

Raw data is becoming increasingly relevant for new discoveries, with mining data for insights becoming equally as important as producing new data (which is often time consuming and expensive). Raw data is needed to verify and reproduce results, and often contains information that is missing from a final structure placed in the PDB. The use of raw inputs for AI training is new, and the importance should not be understated.

Previously, I’m not sure the raw data preservation was given much consideration; partly for practicality of storage, partly because nobody thought of their facility as being a single point of failure.

With this data becoming increasingly valuable, relying on a single institution or a single funding line for sustainability, is probably not enough.

Leave a Reply