From Traces to Agentic Worlds Agentic Language World Models for Interactive Environment Simulation

Quanyu Long1, Xiao Chen2, Jianda Chen1, Haozhen Zhang1, Qisheng Hu1, Jianzhu Bao1, Wenya Wang1

1Nanyang Technological University  ·  2The Hong Kong Polytechnic University

From interaction traces to an agent-operated simulated environment
From interaction traces to an agent-operated simulated environment. Trace2Env reconstructs traces into a reusable environment worldbook containing schemas, grounded evidence, and induced behavioral knowledge. During simulation, a task agent supplies actions, while a world-model agent consults the worldbook together with the current episode state and episodic memory to infer their consequences and predict the resulting observations.
Come and try our live demo.

Trace2Env at a glance

Trace2Env is a learning-free framework for reconstructing simulated environments from historical interaction traces. Offline, it builds an environment worldbook that keeps reusable schemas, grounded evidence, induced behavioral abstractions, and provenance. Online, a dedicated world-model agent actively inspects that knowledge together with persistent episode state and episodic memory, proposes the next observation and lasting state effects, and lets a shared harness validate and commit the transition.

Build environments from tens of traces.

Trace2Env reconstructs a reusable behavioral worldbook from recorded interactions, without requiring the original environment's source code, dependencies, or infrastructure during simulation.

Agentic world modeling.

The world-model agent actively consults the schemas, rules, evidence, state, and memory needed for the current action instead of receiving the entire reconstructed environment as one flattened prompt.

Stateful, long-horizon simulation.

Accepted state effects persist across turns, so later responses reflect the consequences of earlier actions—even when the earlier action itself produced little or no informative output.

Evaluated as a world model.

Measure local fidelity with next-observation prediction on AgentWorldBench, or let a task agent interact successively with the simulated environment and replay its actions in the real environment to measure long-horizon consistency.

Motivation

Consider rm report.txt. The command may print almost nothing, but it changes the environment. Many turns later, cat report.txt should fail. The relevant consequence is therefore not always visible in the observation that immediately follows the action: it has to be inferred when it happens and preserved for later turns.

Conventional prompting places the environment description, interaction history, and current action into a single prediction context. Trace2Env instead separates reusable environment knowledge from what is currently true in this episode, and treats simulation itself as an agentic task.

An agent serves as the environment for another agent. The task agent decides what to do; the world-model agent determines what the environment does next.

Method

Trace2Env framework: environment reconstruction, worldbook, persistent episode workspace, and agentic simulation
Trace2Env framework. Historical traces are reconstructed into a fixed environment worldbook. During a new episode, the world-model agent combines that reusable knowledge with mutable episode state and episodic memory, then proposes a transition that the harness validates and commits.

1. Reconstruct a worldbook

Align action–observation transitions, retain the original observations as grounded evidence, induce action and state schemas, and derive reusable rules, constraints, response contracts, conventions, and provenance. Model weights remain fixed.

2. Maintain an episode workspace

The immutable worldbook describes how the environment behaves. Mutable state records what is currently true in this episode, while episodic memory preserves earlier action–observation turns and details that need not fit into the state schema.

3. Simulate agentically

For each action, the world-model agent selectively inspects relevant knowledge, state, and memory, then proposes state effects plus the next observation. The harness validates represented effects against schemas and constraints, commits accepted state changes, and returns the observation.

Results

We evaluate Trace2Env in two complementary settings: single-turn next-observation prediction across seven environments, and multi-turn interaction consistency on ALFWorld and SciWorld. All worldbooks are reconstructed once with GPT-5.6-Sol and kept fixed across world-model backbones.

Single-turn next-observation fidelity

AgentWorldBench's official five-dimension judge scores format, factuality, consistency, realism, and quality on a 0–100 scale. Trace2Env obtains the best average score with both backbones and outperforms Direct Prompting in every environment–backbone pair.

MethodTerminalSWEAndroidWebFoodShoppingBenefitsAvg.
World-model backbone: GPT-5.6-Sol
Direct Prompting55.7665.8762.8054.2381.6777.3288.8969.51
Trace RAG Prompting58.4266.5066.5057.1285.5786.0796.1173.76
Worldbook Prompting59.0767.8265.5357.7086.7485.8997.6874.35
Harness only60.0669.8962.6255.4082.0777.6287.8970.79
Trace2Env65.4270.5965.1059.5587.4385.3698.1175.94
World-model backbone: DeepSeek-V4.1-Flash
Direct Prompting60.6563.5057.2552.5080.1077.5689.4468.71
Trace RAG Prompting61.9066.1860.7854.6585.9784.4097.5673.06
Worldbook Prompting65.2469.3162.6555.6487.5385.5498.1074.86
Harness only62.0666.0660.1052.5582.7777.5091.4470.35
Trace2Env64.1270.1765.5057.8588.2086.0798.3375.75

Trace2Env is best on average with both backbones. The paper also reports trajectory-clustered confidence intervals and analyses on tasks absent from the construction traces.

Long-horizon interaction consistency

A task agent solves each task in the real environment (Real) and inside the world model (WM). The actions generated inside the world model are then replayed in the real environment (W2R). CR = W2R / Real measures how much real competence survives when interaction is mediated by the world model.

A prompted world model can appear successful while continuing from a false state. W2R tests whether the behavior induced by the simulation remains valid in the real environment.

World modelRealWMW2RCR
ALFWorld
Direct Prompting93%97%3%0.032
Trace2Env93%88%85%0.914
SciWorld
Direct Prompting85%92.5%45%0.529
Trace2Env85%87.5%60%0.706

Worldbook ablations and trace scaling

On Terminal, grounded evidence provides the strongest standalone gain, while abstraction becomes useful when paired with evidence.

VariantTotalEffect
Direct Prompting55.76–
Schema only62.05–
  + Abstraction60.69−1.12 ± 0.90
  + Evidence64.18+2.45 ± 0.89
Full worldbook64.89
  Abstraction given Evidence+0.69 ± 0.63
  Evidence given Abstraction+4.26 ± 1.02
Trace scaling on Terminal
Trace scaling on Terminal. Next-observation score and one-time worldbook construction cost as the number of construction traces increases.

Case Study

The long-horizon benefit is easiest to see when an action should fail. A single hallucinated success can create a fictitious state that remains internally coherent but no longer matches the real environment.

ALFWorld case: Direct Prompting accepts an unsupported action, while Trace2Env preserves the no-op state and supports recovery
One wrong transition derails a whole interaction. The task agent issues an unsupported put command. Direct Prompting hallucinates a successful placement, commits a false state, and reaches a simulated success whose action sequence fails under real-environment replay. Trace2Env predicts the correct Nothing happens., preserves the tissue box in inventory, and keeps later inventory and help responses consistent with that failure, allowing the task agent to recover with the valid move command.

Citation

@article{long2026agenticworlds,
  title={From Traces to Agentic Worlds: Agentic Language World Models for Interactive Environment Simulation},
  author={Long, Quanyu and Chen, Xiao and Chen, Jianda and Zhang, Haozhen and Hu, Qisheng and Bao, Jianzhu and Wang, Wenya},
  journal={arXiv preprint arXiv:2610.06100},
  year={2026}
}