Published September 23, 2026 | Version 1.0.0

Plan Once, Ground Locally: Plan-and-Execute Web Agents Match ReAct Accuracy at One-Sixth of the Token Cost

Authors/Creators

Description

LLM web agents usually follow the ReAct pattern: the model is called again after every browser action, so each click is paid for with a network round trip and with a prompt that has grown since the last one. We study the alternative in which the model writes the whole plan in one call and the steps are executed locally, with a small on-device decision model (Laya, a ModernBERT-large "System-1" model used zero-shot) mapping each step to a page element, judging when asynchronous content has arrived, and checking actions that left the page unchanged.

We evaluate on a new benchmark of 200 tasks across seven public websites whose ground truth is scraped from the sites themselves or produced by scripted browser runs, and graded by fixed regular expressions. Every task was run three times under four configurations with the same LLM (DeepSeek v4.1 flash): ReAct, ReAct with a reading and turn budget matched to the plan agent, plan-and-execute with Laya, and an ablation that replaces Laya with word-overlap matching, for 2,400 runs, plus 400 further runs with a second model.

Against the budget-matched baseline, planning once is not significantly different on success (88.8% vs 92.0%; paired difference -3.2 points, 95% CI -7.7 to +1.2) while using 83.6% fewer input tokens, 67.3% fewer LLM calls and 80.4% less money, and finishing 4.1x faster. Two negative results accompany that headline. Against the unmatched baseline the plan agent also looks more accurate (+8.2 points, p < 0.001), but the entire gap is an artefact of the baseline's smaller reading window and turn cap. And the neural executor is not significantly better than word-overlap matching (+2.3 points, 95% CI -0.2 to +5.0). A second, cheaper model reproduces the token and call savings and widens the accuracy gap in the plan agent's favour, but not the money saving: in 7.5% of its runs it looped while writing the plan and hit the provider's generation cap.

This record contains the paper and the complete study data: the 200 tasks with their ground truth and sources, one graded row per run for all 2,800 runs, the full per-run traces (every prompt, completion, tool result and model decision), the agent code, and the analysis scripts that recompute every number, table and figure from the recorded runs.

Files

benchmark-and-traces.zip

Files (15.6 MB)

Name Size Download all
md5:06d61995db417a186a006b350abd0110
15.4 MB Preview Download
md5:41c094d83c5cc3cdd606a3be38b37b00
212.2 kB Preview Download

Additional details

Identifiers

ISNI
0009-0001-3183-381X

Software

Repository URL
https://github.com/TheDivyanshShukla/plan-once-ground-locally
Programming language
JavaScript , Python
Development Status
Active

References

  • Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. ICLR 2023.
  • Zhou, S., et al. (2024). WebArena: A realistic web environment for building autonomous agents. ICLR 2024.
  • He, H., et al. (2024). WebVoyager: Building an end-to-end web agent with large multimodal models. ACL 2024.
  • Xu, B., et al. (2023). ReWOO: Decoupling reasoning from observations for efficient augmented language models. arXiv:2305.18323.
  • Kim, S., et al. (2024). An LLM compiler for parallel function calling. ICML 2024.
  • Wang, L., et al. (2023). Plan-and-Solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. ACL 2023.
  • Convai Innovations. Laya: an open System-1 decision model. Hugging Face model convaiinnovations/laya.
  • Warner, B., et al. (2024). Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference. arXiv:2412.13663.
  • Shi, T., Karpathy, A., Fan, L., Hernandez, J., & Liang, P. (2017). World of Bits: An open-domain platform for web-based agents. ICML 2017.
  • Liu, E. Z., Guu, K., Pasupat, P., Shi, T., & Liang, P. (2018). Reinforcement learning on web interfaces using workflow-guided exploration. ICLR 2018.
  • Yao, S., Chen, H., Yang, J., & Narasimhan, K. (2022). WebShop: Towards scalable real-world web interaction with grounded language agents. NeurIPS 2022.
  • Deng, X., et al. (2023). Mind2Web: Towards a generalist agent for the web. NeurIPS 2023 Datasets and Benchmarks.
  • Drouin, A., et al. (2024). WorkArena: How capable are web agents at solving common knowledge work tasks? ICML 2024.
  • Le Sellier De Chezelles, T., et al. (2024). The BrowserGym ecosystem for web agent research. arXiv:2412.05467.
  • Gur, I., et al. (2024). A real-world WebAgent with planning, long context understanding, and program synthesis. ICLR 2024.
  • Abuelsaad, T., et al. (2024). Agent-E: From autonomous web navigation to foundational design principles in agentic systems. arXiv:2407.13032.
  • Lu, X. H., Kasner, Z., & Reddy, S. (2024). WebLINX: Real-world website navigation with multi-turn dialogue. ICML 2024.
  • Zheng, B., Gou, B., Kil, J., Sun, H., & Su, Y. (2024). GPT-4V(ision) is a generalist web agent, if grounded. ICML 2024.
  • Chen, L., Zaharia, M., & Zou, J. (2023). FrugalGPT: How to use large language models while reducing cost and improving performance. arXiv:2305.05176.