The Archive · Desktop · Concepts
Models and providers
Which models TALOS Desktop can use, where their keys are kept, models that run on your computer, and how usage is counted.
Checked on Desktop 0.1.19
A newer version is out (Desktop 0.1.25): some details may differ.
TALOS Desktop does not ship a model. It works with the model you choose: one from a provider over the network, or one that runs on this computer.
Providers
Section titled “Providers”A provider is who TALOS talks to in order to use a model over the network: OpenRouter, which reaches many models with one access, the direct providers — among them Anthropic, Gemini and OpenAI — and Hugging Face for models to download. Every provider TALOS knows, with its protocol, key and address, is on Providers.
Providers are configured in Settings → Model laboratory, on its Provider tab. For each one it keeps three facts apart, because they are different:
- The credential — whether there is one, and where it comes from: a key you saved, a sign-in, or a key from the environment.
- The address of the service — the default, your own, or still to set.
- The test — whether someone actually tried to reach it, and how it went.
A saved key does not mean a working key: only the test says so. Keys are never kept in the window: they are kept in the operating system’s keyring. Without any access, the server still starts, but read-only: sessions cannot run.
When you run the server from the source, a key from an environment variable wins over the saved one and cannot be removed from the window. The installed app never reads keys from the environment.
Models
Section titled “Models”The model of a session is changed from the bar at the top of the chat. Settings → Model laboratory has four tabs:
- Hugging Face — models that run on this computer: those to download, and those installed;
- Provider — the providers, their keys, and the catalog of their models;
- Download — every download, with its progress, pauses and recovery;
- System — this computer, telling what is measured, estimated or still unknown.
For each model of the catalog you see its provider, how much context it holds, what it takes in and gives out (text, images, audio, files), whether it can use tools, and its prices — a price the provider does not declare reads “not declared”, never zero.
A model in the catalog is a model that exists, not one your access already covers.
Models on this computer
Section titled “Models on this computer”A local model sends nothing over the network. TALOS knows three engines: its own, which it starts and stops itself, and two you may already have — Ollama and LM Studio.
For its own engine, TALOS says what the model runs on: the graphics card (by name), the processor, or a path you chose. Being reachable and having a model loaded are two different lines, on purpose. When the graphics card is not available, or its memory is too small for the model, TALOS says so and offers the processor — slower — without switching by itself. No model is recommended or forced: the engine’s settings come from the model file itself. See Run a model on your computer.
Usage, counted in tokens
Section titled “Usage, counted in tokens”Settings → Costs and usage counts what was really recorded: by day and by model, in tokens — input, output, and those read again from the cache. It shows no amount of money, by choice: the amount is on your provider’s statement, the only place where it is true. A usage that was not measured is a dash, never a zero, and the total of a session is the whole session, not its last message.
The usual suspects of a high usage are a long preamble and heavy attachments, which are paid at every following message.