Skip to content

Wikimedia Europe

Visual Portfolio, Posts & Image Gallery for WordPress

JohnDarrochNZ, CC BY-SA 4.0, via Wikimedia Commons

Charles J. Sharp, CC BY-SA 4.0, via Wikimedia Commons

Markus Trienke, CC BY-SA 2.0, via Wikimedia Commons

Michael S Adler, CC BY-SA 4.0, via Wikimedia Commons

Benh LIEU SONG (Flickr), CC BY-SA 4.0, via Wikimedia Commons

NASA Goddard Space Flight Center from Greenbelt, MD, USA, Public domain, via Wikimedia Commons

Stefan Krause, Germany, FAL, via Wikimedia Commons

Turning Open Knowledge Into Public-Interest AI-Infrastructure: The Wikidata Embedding Project

As generative AI tools become the default way many people search for information, the question of where that information comes from matters more than ever. Wikimedia Deutschland’s Wikidata Embedding Project offers a concrete answer: it turns Wikidata’s vast, community-verified knowledge base into something small AI systems and humans can use to enable searches by meaning, not just by keyword. It demonstrates what public-interest AI infrastructure can and should look like. 

What it does

The project applies vector-based semantic search to Wikipedia and its sister platforms’ existing data, covering nearly 120 million entries. Until now, Wikidata’s machine-readable data could only be accessed through keyword searches and specialised SPARQL queries — powerful, but difficult for smaller AI applications working with natural language and hard to learn for ordinary users. The new vector database changes that, and it also adds support for the Model Context Protocol (MCP), the emerging standard that lets AI systems communicate directly with external data sources.

In practice, this means a query like “scientist” no longer returns just a keyword match — it surfaces related concepts such as nuclear scientists or Bell Labs researchers, translations, and linked images, giving AI systems the kind of contextual grounding that reduces hallucination and improves traceability.

Built openly, with partners

Wikimedia Deutschland developed the project in Berlin with Jina.AI and DataStax, and released it as a freely accessible public tool. It launched as a fully open and trasnparent alternative to the closed and intransparent datasets that power most commercial AI systems, giving developers a transparent, verifiable foundation to build on instead.

It’s a working demonstration of what public-interest AI infrastructure can look like: multilingual, open, transparent and community-governed — proof that trustworthy AI doesn’t have to mean closed-off AI.

How to get involved

Developers can start building with the public vector database directly and test it against their own use cases, from fact-checking tools to knowledge assistants. 

Anyone working on multilingual NLP or embeddings can help extend coverage beyond the current English, French, and Arabic support. 

Research communities can help, for instance, by flagging feedback or issues back to the Wikidata team. 

Of course, financing further development is always welcome.  

And for anyone simply curious, joining the project’s community discussions or an upcoming webinar is an easy way to follow where it goes next.