European public data to train AI assistants that cite their sources
Tony, creator of Contexte Tech · 1 ottobre 2026
An AI assistant asked about a town or a county has a well-known flaw: when it doesn’t know, it makes things up. A rough population figure, a mayor who doesn’t exist, a statistic out of nowhere.
Yet European public administrations publish reliable, up-to-date figures that anyone can reuse for free. We turned them into five training datasets, published on Contexte Tech under CC BY 4.0.
The idea: answer from the source, or don’t answer
Each dataset holds about forty chat dialogues (JSONL), ready for fine-tuning. They teach a model three things:
- answer only from the official rows given in the message (population, men and women, density, change over time);
- cite the source at the end of every answer;
- politely decline questions outside the data (“Who is the mayor?”, “What’s the weather tomorrow?”) instead of inventing.
They are deliberately small: starter examples to extend with your own data or to duplicate for your region.
Version 2 (3 October): the figure is in the context
A reader on Menéame pointed out the flaw in the first version: forty dialogues teach the format of the answer, not thousands of municipalities, and the statistics office publishes new figures every year. A figure given from memory always ends up wrong.
In version 2, every dialogue carries the official rows in the message: the right row and four others, shuffled. The model learns to read the figure there:
INE data (municipal register, 1 January 2025):
- Torrent (Valencia/València): 90,928 inhabitants, 44,395 men, 46,533 women
- Arucas (Las Palmas): 39,232 inhabitants, 19,421 men, 19,811 women
- …
Question: How many inhabitants does Arucas have?
And when the requested row is not in the data, it declines instead of answering from memory, even for a place seen elsewhere in the dataset. The fine-tune teaches the behaviour; the figures come from the current official table. A script checks that every number in every answer appears in that dialogue’s context: no exceptions. Version 1 stays downloadable in the Versions tab.
The five datasets
| Country | Official source | Dataset |
|---|---|---|
| Germany | Federal Statistical Office (Destatis), via GovData.de: all 400 districts | assistent-landkreise-deutschland |
| Ireland | Central Statistics Office, via data.gov.ie: 26 counties, censuses 1841 to 2022 | assistant-ireland-counties |
| Netherlands | Statistics Netherlands (CBS), via data.overheid.nl: municipalities and provinces | assistent-gemeenten-nederland |
| Italy | Ministry of the Interior, via dati.gov.it: municipalities at the 2021 census | assistente-comuni-italia |
| Spain | National Statistics Institute (INE), municipal register on 1 January 2025 | asistente-municipios-espana |
Each resource shows its exact source in the Source tab, and its status under the GDPR, copyright law and the AI Act in the Compliance tab: no personal data, public data under an open licence.
Use them in one line
pip install contexte
import contexte
ds = contexte.load("contextetech/assistant-ireland-counties")
print(len(ds), ds[0])
To pin a version in your code so an update never changes your results: contexte.load("…", version=1), or contexte get 1ligkia#v1 in the terminal.
A European site for AI practitioners, with more transparency
Contexte Tech aims to be the European site for AI practitioners (developers, companies, researchers, public bodies) where every resource says where it comes from and what you may do with it:
- the source of each resource, with a link to the original data;
- versions: every change is kept, with lines added or removed, and any version stays downloadable;
- compliance for the European Union, the United States and China: personal data, copyright, AI law, licence, distinguishing what the author “declared” from what was “verified”;
- lineage: who started from which resource, and from which version;
- public statistics: views and downloads for each resource.
Why it matters
The AI Act asks model providers to document their training data. Public, sourced data under a clear licence is the simplest case to document, and the easiest to defend.
It is also a concrete way to reuse European open data: these statistics already exist, they are free, and they make assistants more reliable.
Want the same for your country or region? Duplicate a resource, replace the figures with those from your open data portal, and publish your version: it will appear in the Lineage tab of the original.