Fine-tune
Everything on this page is optional. Empty fields mean the automatics decide, and the value they would use is shown in the box. Set a number and it applies; put a row back to Auto and the automatics are back.
Fine-tune a model
Every downloaded model has a Fine-tune button on its row on the Offline Models page. The dialog opens on what the model runs at now, and where that came from: Automatic, Measured here, or Set by you. Everything in it is for this computer only.
Measure on this computer
The main event. One button loads the model a few times with different setups and times each one on your own hardware. It takes about three minutes, and chats pause while it runs; you can stop it at any point. Nothing runs without the click.
Afterwards a slider from Faster answers to More room picks between the measured setups. The card under it says what each one means in plain terms: the context in pages in view at once, the speed in words a second, then the detail (reading speed, the expert split, the speed-up file, the cache). The automatic pick is marked on the scale. More room lets the AI hold more of a long chat or document at once; faster answers come from less room. Measure again reruns it.
Two things the measurement teaches the automatics. It times the Compact context memory beside Standard, and Auto uses Compact on this machine only when that came back clean. And when fewer expert layers in main memory ran at least five percent faster than the automatic pick at that context, Auto uses the measured split at the next load. The dialog says what the measurements decided, one line each.
Save Changes keeps the pick and closes the dialog; the outcome shows in the activity card at the bottom right: reloaded now when the model is the one running, otherwise applied at its next load.
Set it yourself
A drawer at the bottom of the dialog, closed until you open it (or already open if you set values by hand before). Each control is one line, with the explanation in its hover text:
- Context size - tokens of reading room. More than the automatics measure fits may load slowly or fail; the model's trained limit still caps it.
- Expert layers in main memory - mixture-of-experts models only. 0 puts everything on the card; fewer than the automatics pick means the rest must fit on the card.
- Context memory - Standard, Compact or Auto. Compact stores the context at half the memory, so a bigger reading room fits. Auto uses Compact only where a measurement here came back clean.
- Leave the speed-up file out - for a model that has one, to measure the difference yourself.
- Back to automatic puts every row back.
Fine-tune this computer
At the bottom of Settings → Engines:
- Worker threads - for the next model load.
- Creativity (temperature), word variety (top-p), rare-word floor (min-p) and the repetition brake (repeat penalty) - the machine-wide generation layer every AI inherits unless it sets its own.
Per AI: advanced generation settings
An AI's form has an Advanced generation settings fold with the same four generation dials for that AI alone. Empty fields mean the model default, shown in the box; a value applies to new replies immediately - two AIs on the same model can answer with different temperaments. Works for offline and online models.
Which value wins
An AI's own setting, then this computer's setting from Settings → Engines, then the model default.