A new research paper introduces CLIFT, a method designed to improve the training and evaluation of open-source web agents. CLIFT, which stands for Conformal Self-Verification for Web Agent Training and Test-Time Scaling, aims to overcome common challenges in reinforcement learning for agents that interact with web browsers. The primary issues include the difficulty of assigning credit for actions due to sparse success signals and the high cost associated with using large language models (LLMs) for judging every step of an agent's rollout.
CLIFT integrates a process of conformal self-verification during the training phase. The agent is prompted to answer natural-language verification questions about its own actions and observations during a task execution. A "Compositional Conformal Certifier" then processes these self-verification signals. This certifier filters the signals, retaining only those where the URL-conditional evidence aligns with a pre-trained judge.
The method assigns signed trust weights to these verified signals using a technique called polarity-aware lift. The resulting verifier score is then blended into the per-step rewards that the agent receives. This blending provides a more granular and informative reward signal than traditional binary success or failure, which can be too sparse for effective reinforcement learning.
The development of CLIFT comes as open-source web agents demonstrate increasing capability in performing complex browser tasks. However, training these agents with reinforcement learning has been hampered by the limitations of weak supervision. Existing methods often rely on binary task success, which offers limited feedback for learning, or on frontier language model judges, which are computationally expensive and may not be available during deployment.
The concept of self-verification and test-time scaling for LLMs and agents has been an active area of research. For instance, Self-Enhanced Test-Time Scaling (SETS) leverages LLMs' self-verification and self-correction abilities to improve performance on complex reasoning tasks without additional training. Similarly, BrowseConf utilizes verbalized confidence scores to guide test-time scaling for web agents, allowing models to dynamically allocate computational budget based on their self-assessed confidence. Other efforts, such as DeepVerifier, focus on inference-time scaling of verification for deep research agents, using rubric-guided feedback to refine responses without retraining.
The researchers behind CLIFT address the need for more efficient and effective training mechanisms for web agents. By allowing agents to verify their own actions and integrating these verification signals into the reward system, CLIFT aims to provide a more robust and scalable approach to developing agents that can reliably perform tasks on the internet. This method could contribute to wider adoption and improved performance of open-source web agents in real-world applications.
