Building

Compaction provider

Replace how long conversations are summarized so long missions keep the thread.

When a conversation grows toward the model's context limit — automatically at the agent's threshold, or when the user types /compact — AgentParley replaces older history with a summary. A compaction plugin replaces how that summary is written. That is the difference between a long mission that remembers which files it touched and what it decided, and one that starts forgetting.

The platform owns the mechanics: when to compact, what to hand you, and what to keep verbatim. You write one function that turns history into a summary.

compaction.summarize

compaction.summarize(context: {
  messages: { role: string; content: string }[];   // the history to summarize, oldest first
  focus?: string;                                    // present when the user typed /compact <focus>
}): string
  • messages is the historical conversation since, and including, the last checkpoint. On a repeat compaction the first entry is your own previous summary — refine it rather than starting over.
  • The current turn's live request is not in messages; the platform keeps it verbatim after your summary.
  • focus is the user's instruction about what to preserve; honour it.
  • Return one string. Anything else fails the call.

When your summary is refused

The platform runs its built-in summarizer instead, and posts a notice in the chat, when:

  • you return an empty string,
  • your summary is estimated at more than half of the compaction threshold (it wouldn't free enough room), or
  • your handler throws or runs out of time.

The agent is never left without a compaction. (If the history is too large to hand to the sandbox at all, the built-in summarizer is used for that compaction without calling you.)

Use the agent's own model

A summarizer usually needs a model. host.session.llm.complete runs the chat's own agent model — no API key, no model choice, no permission, metered to the account:

// infrastructure/plugins/pi-compactor/index.ts (abridged)
globalThis.compaction.summarize = async (context) => {
  const messages = context.messages || [];
  if (messages.length === 0) {
    return "";
  }

  // The platform re-includes our previous summary as messages[0] on repeated compactions.
  const first = messages[0];
  const hasPriorSummary = first.content.indexOf(SUMMARY_MARKER) === 0;
  const priorSummary = hasPriorSummary ? first.content : null;
  const transcript = serialize(hasPriorSummary ? messages.slice(1) : messages);

  const system =
    "You compact a long agent session into a single structured summary that preserves everything needed to " +
    "continue with NO loss of intent. Output EXACTLY the template below … Your FIRST " +
    'line MUST be "' + SUMMARY_MARKER + '". Output only the summary — no preamble.\n\n' +
    "TEMPLATE:\n" + SUMMARY_MARKER + "\n\n" + TEMPLATE +
    (priorSummary
      ? "\n\nPREVIOUS SUMMARY — refine and extend this; carry its file lists forward:\n" + priorSummary
      : "");

  const result = await host.session!.llm!.complete({
    messages: [
      { role: "system", content: system },
      { role: "user", content: "Conversation to summarize (oldest first):\n\n" + transcript },
    ],
    params: { temperature: 0.2 },
  });

  return result.assistant.text || "";
};

Up to 5 host.session.llm.complete calls per compaction. Alternatives: host.llm (the account's utility model, if set) or your own vendor through host.fetch.

What else it can reach

A compaction handler gets the full host surface your permissions earn — host.session.*, host.agents.*, host.workspace.* (when the chat has an agent), host.fetch, host.media — so it can, for example, also write the summary into a file. It has 120 seconds.

Setup for the user

Install the plugin, then Settings → Engines → Context compaction and choose it. The choice is per account.

See the Pi Compactor walkthrough.