The Archive · Desktop · Concepts

What the model reads

The preamble TALOS Desktop sends before your message, your project's instructions, and how a long conversation is kept within the model's window.

Checked on Desktop 0.1.19

A newer version is out (Desktop 0.1.25): some details may differ.

Everything the model knows about your work, it reads in the conversation: a short preamble, your messages and attachments, and what its tools bring back. Knowing what goes in explains both what the agent knows and what a conversation costs.

Before your message, TALOS sends the model a preamble of four blocks, from the one that changes least to the one that changes most:

  1. The engine’s instructions — who the agent is and how it behaves. They cannot be edited.
  2. Your project’s instructions — the content of AGENTS.md or CLAUDE.md in the folder you work on.
  3. The folder map — the shape of the project, not the list of every file. Files are searched when needed.
  4. The working card — the state of version control, the active permission, the model, and how the work is verified.

The preamble is paid at every message, not once. A long instructions file is the most common reason a conversation costs more than expected.

A file of instructions in the project reaches the model with every message: it is where you write the rules the agent must always follow.

  • One file per folder, looked for in this order: AGENTS.md, then CLAUDE.md.
  • A subfolder can have its own: it applies to that part of the project and wins over the one above.
  • The search never climbs above the project’s root: an AGENTS.md in your home folder is not read.
  • There is a cap: 24,000 bytes by default. Beyond it the text is cut at the nearest line break, and the cut is declared, with where to resume; a file left out entirely is named.

Keep them short. See Project files for every file a project can hold.

  • The Library. The model reads a Library document only when it searches for it or opens it. Its content never enters the context by itself.
  • Memory. The model reads memory when it needs it, with its memory search tool. See Where work lives.

What you attach with + in the composer is part of the message, and of every message after it. Each attachment shows an estimate of what it costs in tokens before you send. Up to 10 attachments per message, up to 200,000 characters each — about 50,000 tokens, half a window for one file. See Attach files and images.

A long conversation eventually fills the model’s window. The chat’s context panel keeps it within bounds:

  • The measure — tokens in, the model’s window, what is reserved for the reply, and where the count comes from (the engine, the provider, or an estimate). When the context changed after the count, it says so instead of showing an old number as fresh.
  • Compaction — a summary that frees space, made automatically when the context fills up, or when you press Compact now. The originals stay available: the summary is what the model reads, not what stays on disk.
  • Facts to keep — facts that must survive the summary. They are kept apart from it, and you can add, edit and remove them.
  • Versions — the saved summaries. You can restore an earlier one; later messages stay in the chat.

The model that writes the summary can follow the chat’s, or be chosen apart. See Keep a long conversation in bounds.

Type to search the guides.