The Archive · Android · Concepts

Models on the phone and in the cloud

Which models TALOS for Android can use, where their keys are kept, how it tells whether a model fits this phone, and the defaults that pick a model for each job.

Checked on Android 0.1.38

A newer version is out (Android 0.1.40): some details may differ.

TALOS does not ship a model. It works with the one you choose: from a provider, with your own key, or downloaded onto the phone and run there.

Provider How it connects
OpenAI, Anthropic, Google Gemini, DeepSeek Your API key.
OpenRouter Your API key, or sign in with OpenRouter and let it issue one. Hundreds of models through one key.
Ollama The address of an Ollama server you run. No key.
On the phone A GGUF model run by the engine built into the app. No key, no network.

A key is kept in the phone’s Keystore and sent only to its provider; the app shows only whether one is saved. Every provider, with what it needs, is on Providers.

All of this lives in the Model Lab — Models in the menu, or Settings → Models:

  • Providers and access — keys, local addresses, whether each provider is ready, and an optional Hugging Face token;
  • Model catalog — the models your providers list, with what each declares it can do;
  • Local models — the models on this phone, and those you can download;
  • This device — usable memory, storage and the reserve TALOS keeps free, measured on request.

The Model Lab and the chat’s quick picker share one selection: the model you pick in either is the one that answers.

The catalog keeps what a provider observed about a model apart from what is merely declared. A model can be tested from the catalog with a short completion; until then it reads “not tested”, not “working”. A model the provider does not list can be added by hand under Advanced manual models — its capabilities are then marked as yours, not the provider’s.

A local model is a GGUF file run through llama.cpp, inside the app. What you write to it never leaves the phone.

Before anything is downloaded, TALOS measures this phone — free memory, storage, temperature and memory speed — and checks every variant against it:

  • Memory: room to spare — it fits, with RAM left once it is loaded;
  • Memory: little room — it fits, but under load Android may close TALOS;
  • Speed: reads from disk — it barely fits, and the phone would re-read the weights for every word;
  • Not enough memory — it would not run.

The verdict states the context length it was checked at, the RAM left after loading, and a predicted speed. A ledger shows where every byte comes from — weights, KV cache, compute, runtime, and the safety margin — and which numbers are measured, predicted or fixed. Memory is measured against what the model really keeps resident, not against the peak a memory-mapped file appears to use.

The graphics processor is used only after this phone passes a real test, never because of a chipset’s name. See Run a model on the phone.

One limit: a model on the phone cannot drive a Code session yet. Choosing one there is refused with that reason; cloud providers work.

Each model is chosen where it is used:

  • the model that answers — in the composer, while you write;
  • the two models of a deep research — the one that writes the report and the one that checks its citations — in the research itself, next to the plan they govern.

Settings → AI Defaults holds what applies to every conversation:

  • Assistant tone — Balanced, Engineering, Friendly or Concise. The model may suggest a better tone for a conversation; you decide from the notification.
  • Vision routing preference — prefer a model that sees images when you attach one.
  • Let chats use your Library — see Use the Library in chats.

Type to search the guides.