A new method called AutoCompact trains artificial intelligence agents to manage their own context during complex, long-term coding tasks. Researchers introduced AutoCompact, which allows coding agents to learn when to summarize previous work and what information to preserve, thereby improving efficiency and task success.

Current AI coding agents often struggle with long-term projects where they must inspect code, search repositories, edit files, and test changes. As these tasks progress, earlier information can become outdated or irrelevant. Simply keeping all past information can overwhelm the agent's context window, leading to errors or reduced performance. AutoCompact addresses this by enabling the agent to make decisions about context management as part of its core operation.

The AutoCompact method involves training the agent to decide when to compact its context, what parts of the working state to retain, and how to proceed from the compacted state. To gather training data, researchers ran a base agent on various coding tasks. A separate "judge" then reviewed the agent's compaction decisions, the summaries it created, and its subsequent actions. Any flawed outputs identified by the judge were corrected before being executed, ensuring that the agent learned from accurate decisions. This corrected data was then used for supervised fine-tuning, followed by reinforcement learning to further optimize both coding and compaction abilities based on task completion rewards.

Experiments conducted on benchmarks like SWE-bench Verified and SWE-PolyBench Verified demonstrated AutoCompact's effectiveness. The method improved task pass rates by 9.2% and 5.0% respectively, compared to the base model. These gains were observed across different context window sizes, including a 256K window that did not overflow and a 16K window where compaction was triggered by overflow.

Unlike methods that compact context only when a fixed threshold is reached, AutoCompact allows the agent to adaptively decide when compaction is most beneficial for task progress. This proactive approach helps prevent the accumulation of stale information, such as failed attempts or verbose tool outputs, which can distract the agent. By replacing this history with a concise summary of the current working state, AutoCompact allows the agent to focus on the task at hand.

The researchers highlight that AutoCompact moves context compaction from a rigid, rule-based process to a learned behavior. By combining judge-guided supervised fine-tuning with reinforcement learning, coding agents can learn to manage their context effectively, leading to better performance and reduced processing costs.