The hidden gap between available and usable
Europe has built one of the world’s most ambitious environmental data infrastructures. The Water Framework Directive (WFD) alone carries an estimated compliance cost of €89 billion for the 2022-2027 cycle, and a growing share of this investment flows into sensor networks, earth observation programmes, and centralized data portals. Tens of thousands of water-related datasets are now published online. Yet a critical question has remained largely unexamined: can this data actually be used? As water governance shifts toward automated analytics, AI-assisted workflows, and cross-border data integration, datasets must be more than technically available to human users; they must be machine-ready. A dataset whose format is undeclared, whose license is ambiguous, or whose access link is broken cannot be ingested into a flood forecasting model, cited in a compliance report, or merged with a neighbouring country’s records. Our recent study addressed exactly this distinction, and the results are sobering.
A framework for measuring Operational Readiness
Our team has developed the EU Water Dataset Observatory, a reproducible framework that evaluates metadata quality not against generic checklists but against the specific demands of three core management tasks mandated by EU water legislation: automated early warning, WFD compliance reporting, and cross-border coordination. Each task activates a distinct, weighted set of metadata dimensions, nine in total, ranging from temporal recency and spatial precision to license openness and multilingual availability, with weights derived from operational guidance documents rather than subjective judgment [[1],[2]]. Every dataset receives a sufficiency score from 0 to 1 for each task, and a complementary impact proxy identifies where metadata improvements would deliver the greatest operational return. We applied the framework to 2,176 water-related datasets harvested from data.europa.eu [[3]], the EU’s primary federated data portal, and validated the results through manual inspection of 25 real datasets from major EU publishers, including the European Environment Agency, the Joint Research Centre, Eurostat, and Copernicus.
What We Found
Mean sufficiency scores across all three management tasks ranged from just 0.16 to 0.24 on the 0–1 scale, and zero percent of datasets, as represented through the federated metadata layer, achieved operational readiness for any task. This conclusion proved fully robust to sensitivity testing. The failures are strikingly uniform: 100% of harvested records lacked mapped machine-readable format declarations and standardized open license metadata, while 99.4% lacked controlled vocabulary keywords and multilingual descriptions. Manual verification deepened the picture: 48% of curated dataset URLs were broken or retired, and not a single one of the 25 datasets carried explicit license metadata on its portal page, even though all were freely accessible resources published under open-data mandates.
Two further findings carry particular weight. First, five of the seven water-management domains we queried, including water quality and WFD metrics, returned zero results through the federated portal, despite such datasets demonstrably existing on national and specialized EU platforms. They are, in effect, invisible to the cross-border users the portal is designed to serve. Second, an exploratory country-level analysis found that metadata quality does not correlate with national wealth, digital maturity, or water stress. The Netherlands, among Europe’s most digitally advanced countries, scored lowest in our sample. The problem, in other words, is institutional rather than capacity-driven.
From diagnosis to policy
These findings arrive at a decisive moment. The European Commission launched a targeted revision of the WFD in 2026, creating a rare opportunity to embed metadata quality standards directly into EU water law. The most consequential failures we identified, missing license and format declarations, are also the cheapest to fix: they can be populated through administrative action and enforced through automated validation at the point of portal ingestion, at negligible cost relative to the monitoring investments they protect. Mandatory machine-readable metadata fields for datasets reported under WFD obligations, persistent identifiers with active link monitoring, and task-specific metadata templates would together transform the operational value of data Europe already owns.
Broader lessons for the data we share
The implications extend well beyond water. As public administrations embrace AI-assisted decision-making, the bottleneck is shifting from data collection to data readiness. General-purpose quality frameworks such as the FAIR principles remain necessary [1], but they are no longer sufficient: fitness must be assessed against the concrete tasks that data are meant to support. Metadata, long treated as an afterthought of cataloguing, has become critical infrastructure for evidence-based governance. The lesson of our study is ultimately optimistic: the barriers separating Europe’s vast data holdings from genuine operational use are tractable, inexpensive, and clearly identifiable. What is required is not more data, but the governance discipline to make existing data work.
[1] Wilkinson MD, Dumontier M, Aalbersberg IJ, et al (2016) The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 3, 160018.
[2] Neumaier S, Umbrich J and Polleres A (2016) Automated Quality Assessment of Metadata across Open Data Portals. Journal of Data and Information Quality 8(1), 2.
[3] Publications Office of the European Union (2024) Open Data Maturity Report 2024. data.europa.eu.
