The Archive · Android · Guides

Run a model on the phone

Find a model this phone can hold, check its variants, download it, and let TALOS find the fastest engine — then answer with no network at all.

Checked on Android 0.1.38

A newer version is out (Android 0.1.40): some details may differ.

A model on the phone is a GGUF file run by the engine inside TALOS. It needs a 64-bit ARM phone, and enough memory for the model to stay resident while it writes. How TALOS judges that is explained in Models on the phone and in the cloud.

  1. Open Models from the menu, then Local models, and the Hugging Face tab.
  2. TALOS measures this phone first — free memory, space, temperature and memory speed — and shows the models recommended for this phone.
  3. Narrow the list with the filters — Fits in memory, Chat, Code-oriented, Q4, Declared permissive license — or by publisher and model size, or search by name.

A model marked license required is gated: accept its license on huggingface.co, then search again.

Open a model. Its page has three tabs: Quantizations, Model card and Files. Every quantization is checked against this phone as soon as the page opens:

  • Memory: room to spare, Memory: little room, Speed: reads from disk or Not enough memory;
  • the context length the check assumed — a smaller one needs less memory;
  • a predicted speed, in tokens per second.

Where every byte comes from breaks the estimate down — weights, KV cache, compute, runtime and the safety margin — and Configure the runtime sets the context length and the KV cache type (Automatic, F16 or Q8_0).

TALOS also warns before you download: a file Hugging Face has flagged, a repository missing some of its parts, a file with no published checksum, or a model that does not list your language.

  1. Tap Download on the variant you chose.
  2. Follow it in the Download Center: one download at a time, which you can pause and resume where it left off. The file is verified when it arrives.

Long downloads need TALOS to keep running with the screen off: see Keep long work running. Without a Hugging Face token you share a request limit with everyone else behind your carrier; an optional token, in Providers and access, gives you your own.

Already have the file? Add a model from this phone copies a GGUF file in — for a few minutes it takes twice the space.

Pick the model in the composer, like any other. The first message after starting can take a while: the model has to load. When the phone is too warm, or short on free memory, TALOS waits for your first message instead of loading it in the background.

  • Settings → Privacy and permissions → Which engine is fastest here runs a short, real generation on each engine this phone offers and remembers the fastest. It costs some battery and heat, and runs only when you ask.
  • Doctor → Measure this model’s threads finds the best thread count for the open model on this phone.
  • Doctor → Local model parity checks that a model handles text, tool calls and Stop correctly before you rely on it. See Check what works on this phone.

A model on the phone cannot drive a Code session yet.

Type to search the guides.