Skip to main content

Conversation compaction

A long conversation is a problem for the model before it is a problem for you. Every turn re-sends the whole history, so a thread that has grown through hundreds of tool calls costs more per step and eventually hits the model's input limit, at which point it can no longer answer at all.

Docana handles this the way coding agents do. When a thread's prompt passes a trigger share of the model window, the older part of the conversation is summarized into a structured summary: the goal, what was done, decisions, identifiers, open items, and the user's stated preferences. The next turn starts from that summary plus the most recent messages kept verbatim. Nothing is deleted. The messages stay in the database and in the chat, and the summary is stored on the thread so it is computed once and reused.

The model can still read the transcript​

A summary loses detail by design. So once a thread has been compacted, the model gains a tool, recallConversation, that reads the exact earlier messages of that same thread, by keyword or by message range. The summary itself tells the model the tool exists and when to use it. Short conversations never see the tool, so nothing changes for them.

This applies to every chat surface: the assistant, agents on WhatsApp and other channels, the agent builder, the document builder, and the employee builder, which uses its existing listThreadMessages tool for the same purpose.

Settings​

Super admins tune compaction on the System settings page. Both values are shares of the model's input window, and the page shows what each means in tokens for the platform default model.

SettingMeaningDefault
TriggerCompact once the prompt exceeds this share of the window0.5
Kept verbatimThe most recent span kept as-is after compaction0.2

A lower trigger keeps each step cheaper and the model working with less context. A higher trigger keeps more raw history in front of the model and costs more per step. The switch on the same card turns compaction off platform-wide; clearing the overrides returns to the defaults.

What to expect​

The summary is written by a small model and its usage is billed to the company that owns the thread, on the same turn that triggered it. If the summary cannot be produced, the turn continues with the full history and nothing is stored. A conversation can be compacted more than once; each pass folds the previous summary into the new one.