Two weeks ago, we launched Data Foundations for AI, a new collaboration between the Open Data Institute and the Global Partnership for Sustainable Development Data. Our starting proposition was straightforward. Governments are moving quickly to explore what AI can do for public services, but without the right data foundations in place, much of that potential will be difficult to realise.
Much of our focus so far has therefore been on the foundations governments need to put in place, such as effective governance and policies, robust technical infrastructure, and the skills and institutional capabilities to manage and use data well.
But one of the most interesting conversations during our launch in New York made us think about whether we need to understand those foundations even more broadly, and particularly about the role of what we might call community infrastructure.
When we talk about data foundations, it is easy to picture interoperable systems, common standards, good metadata and high-quality AI-ready datasets. The discussion in New York highlighted another set of foundations that can be less visible, including the relationships, networks and mechanisms through which people participate in decisions about data and AI. Participants challenged us to think more deeply about how citizens, including marginalised communities, can shape decisions about how data is shared and used.
Governments need ways for citizens and communities, particularly those most likely to be excluded from or affected by AI, to have a meaningful voice in those decisions. That relies partly on trusted relationships between government, civil society, researchers and the private sector, as well as institutions capable of listening to different perspectives, negotiating competing interests and building legitimacy around new uses of data.
These things can sound less tangible than a data platform or technical standard, but they can have just as much bearing on whether data and AI can be used effectively and legitimately. The question for us is therefore not whether they count as infrastructure, but how explicitly community participation, trusted relationships and mechanisms for voice should form part of what we mean by Data Foundations for AI. After all, AI systems do not operate in a vacuum. Decisions about what data is collected, whose experiences are represented, who can access that data, what it can be used for and what happens when things go wrong are not purely technical. They involve social and political choices too. Strong technical infrastructure may make data usable for AI. Community infrastructure can help ensure that its use is legitimate, trusted and shaped by the people it is intended to serve.
That does not mean every decision can, or should, be held up by consultation. Communities are not homogeneous, and different groups will understandably have different priorities, interests and views about acceptable uses of data and AI. Governments still need to make decisions, deliver practical and technical improvements, and demonstrate that those improvements translate into better outcomes for people. The challenge is to build forms of participation that genuinely shape those decisions, while ensuring community engagement enables action rather than inhibits it.
That matters particularly if our starting point is public services. At our launch, Ghana's Government Statistician, Dr Alhassan Iddrisu, challenged us not to stop at infrastructure, but to keep asking what better data ultimately enables, whether that’s better healthcare, greater food security, more effective climate action, or better public services. UN Deputy Secretary-General Amina Mohammed similarly brought the conversation back to people, including young people and girls, and the risk that AI could reinforce existing inequalities.
That is important for how we approach Data Foundations for AI. We shouldn't begin by asking where a government can deploy AI. We should begin with the outcome it is trying to achieve. Where might AI genuinely help? What data would it depend on? What governance and institutional capability would be required? And, importantly, which communities need to be involved in shaping how that data and technology are used?
Our early work is already reinforcing the importance of starting from where countries are, rather than arriving with a predetermined model of what their AI future should look like. Work in Chile and Sierra Leone is helping us understand what implementation looks like in very different contexts and how strengthening foundations needs to respond to the problems governments are actually trying to solve.
There is another reason this broader understanding of infrastructure matters. The technology will keep changing. The models, applications and suppliers governments are considering today may look very different in a few years. By contrast, improving the quality of public data, enabling it to move safely between institutions, establishing effective governance, building institutional capability, and developing trusted relationships with communities takes time, but is also far more durable. Investing in these foundations gives countries the ability to adapt as the technology evolves, rather than designing their AI readiness around the tools available today. That suggests a different way of thinking about AI readiness. The objective shouldn't simply be to prepare countries to adopt today's AI technologies. It should be to strengthen the foundations that give governments and societies agency over how they use AI as the technology evolves.
We came away from New York convinced that this means thinking about data foundations more broadly. Yes, countries need good data, interoperable systems, standards and governance, but they also need capable institutions, trusted relationships and meaningful ways for people to participate in decisions about how data and AI affect their lives.
The AI systems governments use five years from now may be very different from those available today. But the need for strong institutions, trusted data and communities with a meaningful voice in how technology is used will endure. That is a dimension of Data Foundations for AI that we are keen to explore further as the program develops.