CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — agents 20 upvotes

Agent-Editing World Model: Rethinking World Modeling for LLM Agents

QUESTION — How can LLM agents be improved by modeling task progress and editing reasoning states instead of predicting tool responses?

The authors propose the Agent-Editing World Model (AEWM), a framework that models how reasoning and actions shape task progress rather than simulating high-entropy tool responses. The system combines an Action Judge to categorize decisions into Critical, Exploratory, and Noisy with State Revision to edit noisy reasoning-action continuations. Additionally, EditAct integrates these capabilities with real execution to directly change the state underlying subsequent decisions. Trained across Search, Terminal, and Software Engineering, AEWM achieves 70.5% macro-F1 on the Action Judge benchmark and improves average scores across multiple benchmarks.

AEWM achieves 70.5% macro-F1 on our Action Judge benchmark, exceeding the strongest frontier baseline by 10.6 points.

Across six benchmarks and three agent backbones, EditAct improves average scores by 3.2--6.7 points over the strongest baseline.

AEWM-RFT, improves over Self-RFT by 2.2--2.6 points across three domains without online AEWM guidance.

SNHE · 23 Sept 2026 read the original ↗
↑