Skip to main content

Accountability

Data provenance

Where the training data comes from, by source and by legal basis.

Base model
Apertus
Base model provenance
Swiss AI Initiative (ETH Zürich and EPFL)
A walnut half sunk into damp earth, with the split husk beside it.

In short

What was Walnoot trained on?

We trained Apertus further on one open Dutch corpus plus a small source of our own and then on instruction data we put together ourselves. No training data is published, only its provenance.

The best open Dutch model there was disappeared over its data provenance. That is why ours is fixed per source.

Sources

What data is in it?

Per source the legal basis and the processing steps.
NumberSourceBasisProcessing
01GPT-NL Public CorpusBasisCC-BY 4.0, publicly released without registrationProcessingDeduplicated against the rest of our mix. The mix is 70 per cent Dutch, 20 per cent English and 10 per cent code. The English and the code come from this corpus as well. The English comes from the subsets cc_english-pd, cc_loc-pd-books and cc_openalex, the code from cc_github_open_source. They are in there as deliberate replay against knowledge loss.
Registers28 subsets. Of these, 24 hold Dutch text, among them De Rechtspraak, Officiële Bekendmakingen, Tweede Kamer, Openraadsinformatie, Woogle, Eurlex, the Koninklijke Bibliotheek and the Nationaal Archief. Three hold English text and one holds code.
02Instruction data (compiled by WAINUT)BasisLargely synthetic, made with a teacher model under Apache 2.0. Verbatim corpus passages fall under CC BY 4.0. The scientific branch comes from thirteen public datasets, each under the licence of that dataset.ProcessingLicence traced per external source. A gate on personal data for the corpus passages prepared for document tasks in September. No filter on personal data over the data as a whole.

We left newspaper archives out on quality grounds. Text from behind a paywall is not in it either, nor is data whose provenance is not fixed per source.

Method

What did we do with the data?

  1. Selection and filter

    From the corpus we took 28 subsets, at a pinned revision. Every document went through a filter on length and language. Besides the corpus, the continued pretraining contains a small source of WAINUT's own, about 2 per cent of the characters. It consists of synthetic chat pairs, made with the same teacher model as the instruction data, to help the model keep the formatting of text.

  2. Instruction data

    The instruction data is largely synthetic, built from 51 hand-written seed prompts and passages from the corpus. A scientific branch comes from public annotation datasets.

  3. Licence per source

    For the instruction data we looked up the licence of every external source at that source itself. No source in the mix we used is under a non-commercial licence, share-alike, the GPL or a research-only restriction.

  4. Post-training

    In the last step the model learned from answers of its own that a judging step had approved. No new external source came in there.

Public summary of the training contentCopyright policy

Decontamination

Did you train on the test sets?

A model that has already seen the test questions during training scores highly without being able to do anything.

In the measurement after training, two questions from squad-nl shared thirteen consecutive words with the instruction data. At sixteen words they were gone and at twenty, where the line is drawn, nothing was touched. For that data the outcome is clean. The wider check at eight words does not count towards the result and touched 659 items in the instruction data and 33 in the post-training data. The continued-pretraining corpus, in the build the model was trained on, was not measured against the test sets again as a whole. The first pass against the test sets, on 22 July, covered an earlier build.

Result

Corpus
Against a set of our own with which we measure knowledge loss, our corpus comes out at PASS_AMENDED_G2PRIME, which means passed with an amendment that is in the name. That check is not about the test sets. We do not leave that amendment out anywhere.
Held-out
Alongside it stands a G2-FAIL and we do not take that off. Besides the test sets we keep a set of our own aside, to measure whether the model loses knowledge during continued pretraining. On that set the same measurement touched 2,754 of the 6,244 records, 44.11 per cent. That is a failing result and it has not been withdrawn. We pruned that set afterwards, under a rule that was fixed before the outcome of the longer yardstick was known.

How did we search?

We look for thirteen consecutive words in our training data that also appear literally in a test question. For that we first put both texts in lower case and normalise every character, so that a difference in formatting cannot hide a hit. Every hit is then checked against a longer yardstick: sixteen words, then twenty, then twenty-five. Twenty decides. A hit that disappears there was a fixed expression; what remains counts as a leak. The threshold is zero and it applies per dataset.

Do short questions get an exemption?

A test question shorter than that long yardstick gets no exemption. Such a question cannot meet that yardstick by definition. The multiple-choice questions are precisely the short ones. An exemption would therefore cover exactly the questions that matter.

What do we measure against?

For the measurements of July that is a fixed set of 35 files from EuroEval 17.6.0, together 29,273 records, each with its own checksum. Of those, 18,145 come from the test split, 8,674 from the train split and 2,454 from the validation split. Without that freeze a second round measures something else and there is nothing left to compare.

What did the binding measurement show?

On 30 July we measured the two source files of our own source that preserves the formatting of text in the corpus. At thirteen words the outcome there comes to zero. On 12 September, after training, we measured the instruction data and the post-training data the model was actually trained on. There two questions from squad-nl were touched at thirteen words; at sixteen words they were gone and at twenty nothing was touched. That is a fixed phrase and the outcome is clean. The continued-pretraining corpus, in the build the model was trained on, was not measured against the test sets again as a whole. The pass of 22 July covered an earlier build.

Why also measure at eight words?

We also measure at eight words. That measurement does not count towards the result and is there only as a check. On 30 July it touched 723 and 591 of the 29,273 records, per source. On 12 September it touched 659 items in the instruction data and 33 in the post-training data. At such a short fragment ordinary language comes along, such as treaty titles and licence headers. Those figures are here all the same, because a reader who finds them elsewhere would otherwise wonder why they were missing.

Did we find a real leak?

One hit did meet the long yardstick and was therefore a real leak, the Lord's Prayer in test question 1941 of wiki-lingua-nl. That was in the first pass of 22 July, over an earlier build of the corpus with the instruction data of that moment. We name it, because a check that never finds anything says more about how the search was done. How we dealt with it was fixed beforehand and the number of that question is in the pre-registration.

Base model

What about the training data of Apertus?

That data is documented by the Swiss AI Initiative and is not ours to publish.

That it is possible to respect opt-outs at scale and still build a usable model was described by Paul Keller (IViR) in June 2026 on the Kluwer Copyright Blog, in a piece prompted by Apertus. Paul Keller, GPT-NL respects copyright - cui bono? - Part 1, Kluwer Copyright Blog, 25 June 2026