Skip to main content

WAINUT

Press kit

Figures, photos and logo, with the claim exactly as we phrase it.

Description

What is Walnoot 8B?

Walnoot 8B is a Dutch language model from WAINUT. You download the model and run it in your own environment, including for applications where data may not leave that environment. Everything we measure and publish, you can check yourself.

The claim

The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd.

Model
Walnoot 8B Instruct
Version
1.0
Publisher
WAINUT
Licence
Apache 2.0
Size
8.05 billion parameters
Base model
Apertus
Base model provenance
Swiss AI Initiative (ETH Zürich and EPFL)
Figures
17 September 2026

WAINUT

Who is WAINUT?

WAINUT is the Dutch AI company behind Walnoot, founded by Seleman Arefi and Luciano Currie.

Why we built this

Seleman Arefi (left) and Luciano Currie (right), founders of WAINUT, in Rotterdam.Photo: Yasin Celik
Seleman Arefi with his arms folded at a white railing along the Maas in Rotterdam, with the Erasmus Bridge in the background.
Seleman Arefi, co-founder of WAINUT, in Rotterdam.Download the photo (jpg, 3000 × 2000 px, 0.7 MB)
Luciano Currie smiling by the water in Rotterdam, with the Erasmus Bridge and tall towers in the background.
Luciano Currie, co-founder of WAINUT, in Rotterdam.Download the photo (jpg, 3000 × 2000 px, 0.7 MB)
Photo: Yasin Celik

Free to use in coverage of Walnoot and WAINUT, with the photo credit.

50+AI projects in production
1000+professionals trained

Imagery

Logo and mark

Measurement

What exactly does the claim say?

The original goal has been met: match GPT-NL on all six published benchmarks. Five times above it, level on dbrd. WikiLingua-nl falls outside the comparison, because the metric changed there. Walnoot 8B has 8.05 billion parameters against 26.03 billion for GPT-NL.

When does a difference count as a win?

A win only counts as a win when our whole confidence interval lies above their figure. A loss only when it lies entirely below it. If their figure falls inside our interval, the result is a tie. That holds even on a task where we are ahead on points.

What was it measured on?

We measured our column with EuroEval 17.6.0 on the test split, in 10 iterations and with bf16 weights. The GPT-NL column is their own publication of December 2025 on EuroEval 15.16.0, interim scores of the base model, published as point scores without an interval for the figures we take over. The harness is two major versions apart. Across versions the dataset content is not guaranteed to be identical. That is why we state explicitly with every comparison which version was used.

GPT-NL publishes no weights, so nobody outside the consortium can run that model. We have therefore never measured it ourselves. This table sets our measured figure next to their published figure.

To the scoreboardReproduce the figures

Limits

What falls outside the intended use

  • Open-ended reasoning and long, complex analysis
  • Creative work without a brief
  • Code generation
  • Anything where a mistake has immediate serious consequences without a person in between

Why these limits?

Walnoot 8B is built for well-defined tasks. For work that calls for a lot of general knowledge, broad context or long reasoning, larger models are a better fit. You can see that in the EuroEval results too. We publish every score, including the tasks where Walnoot performs less strongly.