The Archive · Desktop · Guides
Run a model on your computer
Load a model into TALOS Desktop's own engine, or use Ollama or LM Studio, and read what it runs on.
Checked on Desktop 0.1.19
A newer version is out (Desktop 0.1.25): some details may differ.
A local model answers without sending anything over the network. TALOS has its own engine, which it starts and stops, and also works with Ollama and LM Studio if you already have them.
Load a model
Section titled “Load a model”- Open Settings → Model laboratory, then its Hugging Face tab: models to download, and those installed.
- See which models are on disk and which one is loaded in memory, each with its size in GB and whether it fits in the memory available.
- Load or unload a model from there.
No model is recommended or forced: you choose, and the engine’s settings are read from the model file itself. The Download tab follows every download, with its progress, pauses and recovery; the System tab describes this computer. The operations behind it are in Local models.
Read what it runs on
Section titled “Read what it runs on”For its own engine, TALOS says in plain words what the model runs on: the graphics card, by name, or the processor, or a path you chose. Being reachable and having a model loaded are shown on two separate lines.
When the graphics card is not available, the model runs on the processor, and TALOS says so. When the card’s memory is too small for the model, TALOS does not decide by itself: it tells you, and offers the processor — slower — in the menu of the local engine.
If something goes wrong
Section titled “If something goes wrong”- The graphics card is not available: the model runs on the processor — a warning, not an error. It works, more slowly.
- The graphics card’s memory is not enough — choose the processor in the menu of the local engine, or a smaller model.
- The engine did not start — Doctor and the engine’s tab show the last error read.
- Generation ends at once — it happens with local models: a chat model served without its conversation format, a window already full, or a sampling that stops at the first token.