Skip to main content

Try it

Run it yourself

Walnoot 8B runs on your own machine or server. No text goes to us and you do not need a key.

Model
Walnoot 8B Instruct 1.0
Format
bf16, about 16 GB
Licence
Apache 2.0

In three steps

What is the shortest way to a first answer?

  1. Download walnoot-8b-instruct-bf16.gguf and the Modelfile from WAINUT/walnoot-8b-instruct. Put them in the same folder.

  2. Build and start the model with Ollama. All the settings and the system prompt are in the Modelfile, so you do not have to set anything else.

    ollama create walnoot-8b-instruct -f Modelfile
    ollama run walnoot-8b-instruct
  3. Type your question and press Enter.

If you would rather work in a window without a command line, use LM Studio. There you enter the settings yourself.

Download on Hugging FaceTo the four programs

What you need for it

What does it fit on?

The file is about 16 GB and that has to fit somewhere.

24 GB of video memory
The model fits on it in full.
Less video memory
The program spreads the layers over card and processor. That works, but it becomes slower.
Laptop with 12 GB of video memory and 32 GB of system memory
A good half of the layers sit on the card.

How fast the model runs depends strongly on your card and your memory. Measure it once on your own machine.

We ship the model only at full size, in bf16. It carries the same bf16 weights the figures on this site were measured on, so you run what we measured.

Why no smaller variant?

Smaller, quantised variants fit on lighter cards, but then our figures no longer hold and the behaviour of the model changes. That is why we do not publish them.

The three that matter

Which settings make a difference?

Repetition penalty at 1.0
Walnoot is trained to quote from a source document and a repetition penalty punishes exactly that.
Context length at 8192
The defaults in LM Studio and Ollama are too tight as soon as a document has to fit as well.
Temperature at 0 for work on a document
The model then picks the most likely word every time and invents less.

Why these values?

Repetition penalty at 1.0

Walnoot is trained to quote from a source document. A repetition penalty punishes exactly that behaviour and makes the model avoid the very words that appear in the source. So set it to 1.0. LM Studio sets it to 1.1 by default.

Context length at 8192

LM Studio starts at 2048 and Ollama at 4096. That is too tight as soon as a document has to fit as well. 8192 is the advice for most machines; 32768 is the serving limit we set and the architecture allows 65536. Leave the KV cache unquantised while you are at it, otherwise reading back a long passage suffers.

Temperature at 0 for work on a document

For summarisation, information extraction and questions about a text, 0 is the right value, because the model then picks the most likely word every time and invents less. For free writing you can take 0.7 with top-p at 0.9.

By program

Which program suits you?

We have tested LM Studio and Ollama on the released model, but not the other two routes. llama.cpp reads the settings from the file we ship; with vLLM your program sends the temperature itself.

  1. LM Studio

    A window to type into, with no command line.Tested

    Load the GGUF file in LM Studio. On this route LM Studio manages the runtime and sampling settings itself, so set them by hand.

    When you load the model:

    • Context Length: 8192
    • Leave the KV cache at full precision; do not use KV cache quantisation such as Q8 or Q4. Offload KV Cache to GPU Memory and Unified KV Cache can stay on.

    Then load the model and go to Settings. Set the following there:

    • Temperature: 0
    • Repeat Penalty: 1.0
    • Top K Sampling: 0
    • Presence Penalty: off
    • Top P Sampling: 1.0
    • Min P Sampling: 0.0

    Finally, paste the full contents of systeemprompt.txt into the System Prompt field.

    NoteDo not skip this. With the default settings you are measuring LM Studio and not Walnoot.

    TestedWe tested this route on the released model on 27 September, with exactly these settings.

  2. Ollama

    Two commands and the settings are right straight away.Tested

    Build the model with the Modelfile we ship and start it. All the settings and the system prompt are in there, so you do not have to set anything else.

    ollama create walnoot-8b-instruct -f Modelfile
    ollama run walnoot-8b-instruct

    NoteIf you call Ollama from a program, use its own route /api/chat. On the OpenAI route /v1/chat/completions Ollama ignores the settings in the Modelfile and silently sets the temperature to 1.0. If you do use that route, send temperature 0 and top-p 1 yourself.

    TestedWe tested this route on the released model on 27 September, with exactly these two commands. What is said above about the two entry points is a property of Ollama itself and not of the model.

  3. llama.cpp

    Your own server at your workplace.Not re-measured

    The settings are in the file and are read from it, so you start the server without sampler flags. Use a release of llama.cpp from October 2025 or later, because older versions do not know the compute step of this model.

    llama-server \
      --model walnoot-8b-instruct-bf16.gguf \
      --ctx-size 8192 \
      --parallel 1 \
      --cache-type-k f16 \
      --cache-type-v f16

    NoteAfter starting, check what the server holds to: temperature 0, top-k 0, top-p 1.0, min-p 0.0 and repetition penalty 1.0. If a program sends its own settings along, those always win over ours.

    TestedWe have not re-measured this route on the released model. The flags above belong with the file we ship.

  4. vLLM, on a server

    Several users at once, in your own data centre or cloud environment.Not run by us

    For a server you do not take the gguf file but the weights themselves, from the ordinary model repository. vLLM serves several requests at once. Plan for a card with at least 24 GB of video memory. How you start vLLM depends on your set-up; these are the settings you give it.

    --model /path/to/walnoot \
      --served-model-name walnoot-8b-instruct \
      --generation-config auto \
      --max-model-len 8192 \
      --dtype bfloat16

    NotevLLM takes no temperature from the files we ship. So have your program send temperature 0 and top-p 1 itself with every request. If you get the message that the architecture is not supported, your vLLM is too old for the base model Walnoot is built on.

    TestedWe have not run this route ourselves and we have no measurement underneath it.

The system prompt

Why does a system prompt belong with it?

Without a system prompt the model invents a role of its own for each conversation and you get answers in a tone you did not ask for. We ship one as systeemprompt.txt, in the same repository as the gguf file. In Ollama it already sits in the Modelfile. In LM Studio you paste its contents into the System Prompt field.