Trace2Env at a glance
Trace2Env is a learning-free framework for reconstructing simulated environments from historical interaction traces. Offline, it builds an environment worldbook that keeps reusable schemas, grounded evidence, induced behavioral abstractions, and provenance. Online, a dedicated world-model agent actively inspects that knowledge together with persistent episode state and episodic memory, proposes the next observation and lasting state effects, and lets a shared harness validate and commit the transition.
Build environments from tens of traces.
Trace2Env reconstructs a reusable behavioral worldbook from recorded interactions, without requiring the original environment's source code, dependencies, or infrastructure during simulation.
Agentic world modeling.
The world-model agent actively consults the schemas, rules, evidence, state, and memory needed for the current action instead of receiving the entire reconstructed environment as one flattened prompt.
Stateful, long-horizon simulation.
Accepted state effects persist across turns, so later responses reflect the consequences of earlier actions—even when the earlier action itself produced little or no informative output.
Evaluated as a world model.
Measure local fidelity with next-observation prediction on AgentWorldBench, or let a task agent interact successively with the simulated environment and replay its actions in the real environment to measure long-horizon consistency.
Motivation
Consider rm report.txt. The command may print almost nothing, but it changes the environment. Many turns later, cat report.txt should fail. The relevant consequence is therefore not always visible in the observation that immediately follows the action: it has to be inferred when it happens and preserved for later turns.
Conventional prompting places the environment description, interaction history, and current action into a single prediction context. Trace2Env instead separates reusable environment knowledge from what is currently true in this episode, and treats simulation itself as an agentic task.
Method
1. Reconstruct a worldbook
Align action–observation transitions, retain the original observations as grounded evidence, induce action and state schemas, and derive reusable rules, constraints, response contracts, conventions, and provenance. Model weights remain fixed.
2. Maintain an episode workspace
The immutable worldbook describes how the environment behaves. Mutable state records what is currently true in this episode, while episodic memory preserves earlier action–observation turns and details that need not fit into the state schema.
3. Simulate agentically
For each action, the world-model agent selectively inspects relevant knowledge, state, and memory, then proposes state effects plus the next observation. The harness validates represented effects against schemas and constraints, commits accepted state changes, and returns the observation.
Results
We evaluate Trace2Env in two complementary settings: single-turn next-observation prediction across seven environments, and multi-turn interaction consistency on ALFWorld and SciWorld. All worldbooks are reconstructed once with GPT-5.6-Sol and kept fixed across world-model backbones.
Single-turn next-observation fidelity
AgentWorldBench's official five-dimension judge scores format, factuality, consistency, realism, and quality on a 0–100 scale. Trace2Env obtains the best average score with both backbones and outperforms Direct Prompting in every environment–backbone pair.
| Method | Terminal | SWE | Android | Web | Food | Shopping | Benefits | Avg. |
|---|---|---|---|---|---|---|---|---|
| World-model backbone: GPT-5.6-Sol | ||||||||
| Direct Prompting | 55.76 | 65.87 | 62.80 | 54.23 | 81.67 | 77.32 | 88.89 | 69.51 |
| Trace RAG Prompting | 58.42 | 66.50 | 66.50 | 57.12 | 85.57 | 86.07 | 96.11 | 73.76 |
| Worldbook Prompting | 59.07 | 67.82 | 65.53 | 57.70 | 86.74 | 85.89 | 97.68 | 74.35 |
| Harness only | 60.06 | 69.89 | 62.62 | 55.40 | 82.07 | 77.62 | 87.89 | 70.79 |
| Trace2Env | 65.42 | 70.59 | 65.10 | 59.55 | 87.43 | 85.36 | 98.11 | 75.94 |
| World-model backbone: DeepSeek-V4.1-Flash | ||||||||
| Direct Prompting | 60.65 | 63.50 | 57.25 | 52.50 | 80.10 | 77.56 | 89.44 | 68.71 |
| Trace RAG Prompting | 61.90 | 66.18 | 60.78 | 54.65 | 85.97 | 84.40 | 97.56 | 73.06 |
| Worldbook Prompting | 65.24 | 69.31 | 62.65 | 55.64 | 87.53 | 85.54 | 98.10 | 74.86 |
| Harness only | 62.06 | 66.06 | 60.10 | 52.55 | 82.77 | 77.50 | 91.44 | 70.35 |
| Trace2Env | 64.12 | 70.17 | 65.50 | 57.85 | 88.20 | 86.07 | 98.33 | 75.75 |
Trace2Env is best on average with both backbones. The paper also reports trajectory-clustered confidence intervals and analyses on tasks absent from the construction traces.
Long-horizon interaction consistency
A task agent solves each task in the real environment (Real) and inside the world model (WM). The actions generated inside the world model are then replayed in the real environment (W2R). CR = W2R / Real measures how much real competence survives when interaction is mediated by the world model.
A prompted world model can appear successful while continuing from a false state. W2R tests whether the behavior induced by the simulation remains valid in the real environment.
| World model | Real | WM | W2R | CR |
|---|---|---|---|---|
| ALFWorld | ||||
| Direct Prompting | 93% | 97% | 3% | 0.032 |
| Trace2Env | 93% | 88% | 85% | 0.914 |
| SciWorld | ||||
| Direct Prompting | 85% | 92.5% | 45% | 0.529 |
| Trace2Env | 85% | 87.5% | 60% | 0.706 |
Worldbook ablations and trace scaling
On Terminal, grounded evidence provides the strongest standalone gain, while abstraction becomes useful when paired with evidence.
| Variant | Total | Effect |
|---|---|---|
| Direct Prompting | 55.76 | – |
| Schema only | 62.05 | – |
| + Abstraction | 60.69 | −1.12 ± 0.90 |
| + Evidence | 64.18 | +2.45 ± 0.89 |
| Full worldbook | 64.89 | |
| Abstraction given Evidence | +0.69 ± 0.63 | |
| Evidence given Abstraction | +4.26 ± 1.02 |
Case Study
The long-horizon benefit is easiest to see when an action should fail. A single hallucinated success can create a fictitious state that remains internally coherent but no longer matches the real environment.
put command. Direct Prompting hallucinates a successful placement, commits a false state, and reaches a simulated success whose action sequence fails under real-environment replay. Trace2Env predicts the correct Nothing happens., preserves the tissue box in inventory, and keeps later inventory and help responses consistent with that failure, allowing the task agent to recover with the valid move command.Citation
@article{long2026agenticworlds,
title={From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation},
author={Long, Quanyu and Chen, Xiao and Chen, Jianda and Zhang, Haozhen and Hu, Qisheng and Bao, Jianzhu and Wang, Wenya},
journal={arXiv preprint arXiv:2610.06100},
year={2026}
}