EVOLVINGNAV / PERSISTENT EMBODIED NAVIGATION

Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds

Mingjian Gao1,*Zhaocheng Li1,*Haoyang Huang2,*Wenqiao Zhang1,‡Yingjie NIU3,4,‡,†Hao Zhou1Chao Li3Juncheng Li1Siliang Tang1Yueting Zhuang1

1 Zhejiang University · 2 University of California, San Diego · 3 Deeprobotics · 4 The Chinese University of Hong Kong

* Equal contribution · ‡ Corresponding authors · † Project lead

Predict the world. Inspect the evidence. Keep navigating.
54scenes in
EvoWorld-Bench
803.68Kexecutable
task instances
61.32%first-inspection
success rate
86.18%budgeted search
success rate
01 / THE BIG PICTURE

Navigation in a world
that keeps changing.

From remembered observations to predictive belief — and from belief to evidence-driven action.

Overview of EvolvingNav, predictive 4D belief, EvoWorld-Bench, and persistent navigation tasks
Overview. EvolvingNav converts timestamped 3D histories into a belief over the world at the agent’s future arrival time, then closes the loop with action, visibility-qualified evidence, belief revision, and memory update. EvoWorld-Bench evaluates this loop across predictive navigation, search, replanning, online dynamics, transfer, and embodied question answering.
THE CENTRAL INSIGHT

A remembered world is a hypothesis.

EvolvingNav treats navigation as an evidence loop: predict what will still be true at arrival, inspect only when the observation can change the decision, and revise the memory when the world disagrees.

01

Observe

Build a timestamped 3D history from prior RGB-D observations.

02

Predict

Forecast persistence and relocation at each candidate arrival time.

03

Inspect

Use visibility-aware evidence to test the leading hypothesis.

04

Update

Revise belief, reopen plausible candidates, and replan.

02 / RESEARCH SUMMARY

Abstract

Persistent embodied agents must act from memories that can become stale: objects move while the agent is away, continue evolving during navigation, and may remain hidden even after an inspection. Existing methods do not jointly account for this continued hidden-world evolution and visibility-conditioned belief revision.

We introduce EvolvingNav, which turns timestamped 3D histories into a persistence–relocation belief over the current world. Its event-driven filter forecasts object state at candidate arrival times, invokes a frozen zero-shot vision-language model only when new RGB-D evidence is informative, and replans when evidence invalidates a candidate. We also introduce EvoWorld-Bench, a 54-scene benchmark with 803.68K tasks. EvolvingNav improves first-inspection decisions, budgeted search, and recovery, with the largest gains when environmental change has learnable regularity.

03 / THE METHOD

From memory to action.

EvolvingNav represents each remembered entity with a timestamped 3D history and a factorized belief: a persistence term estimates whether the last observed state remains valid, while a relocation term distributes probability over plausible destinations. Instead of predicting only at query time, the planner forecasts the belief at each candidate’s estimated arrival time.

EvolvingNav pipeline from temporal 3D memory to arrival-time belief, inspection, and evidence-conditioned replanning
Predictive 4D belief and closed-loop navigation. A current-time filter summarizes the latent world, candidate-specific forecasting evaluates future arrival states, and visibility-qualified RGB-D evidence updates only hypotheses that should have been observable.
View full diagram ↗

Execution proceeds in short action chunks. After each informative observation, the posterior is revised, invalidated candidates are suppressed, and previously rejected candidates may be reopened when time or evidence makes them plausible again.

timestamped history → arrival-time belief → short-horizon action → visibility-qualified evidence → replan

04 / EXPLORE THE IDEAHSSD scene · illustrative belief update

Interactive 4D World Memory Explorer

Explore a real HSSD scene from above. Follow the quadruped’s inspection route and see how new evidence changes an illustrative belief over object locations.

Scene images are captured in Habitat-Sim from HSSD. Object markers, probabilities and inspection outcomes illustrate the method; they are not measured model outputs. Scene provenance ↗

05 / THE BENCHMARK

EvoWorld-Bench

EvoWorld-Bench evaluates navigation when the world changes between observations and may continue changing while the agent acts. The latest release contains 54 scenes and 803.68K executable tasks, with persistent temporal histories, causal observability, controlled dynamics, and held-out transfer settings.

Figure 4Figure 4: EvoWorld-Bench construction pipeline from human traces to executable evolving worlds
Benchmark construction. Human traces are aligned into a unified event schema, grounded and normalized, then instantiated as executable static, routine, and random worlds.
Figure 3Figure 3: EvoWorld-Bench task distribution, mobility patterns, and evaluation protocols
Evaluation protocols. N1–N5 and EQA cover predictive navigation, belief-guided search, evidence-aware replanning, online dynamics, held-out transfer, and embodied question answering.
N1Predictive navigationFirst-Inspection SR
N2Belief-guided searchSearch SR / cost
N3Evidence-aware replanningRecovery SR
N4Online dynamicsDynamic SR / recovery
N5 + EQATransfer & reasoningHeld-out SR / accuracy
06 / EXPERIMENTS

Evidence across environments.

EvolvingNav improves both the first destination selected from stale memory and the ability to recover within a search budget. It transfers across FindingDory, GOAT-Bench, EvoWorld-Bench, and physical LYNX M20 trials.

53.22%FindingDory HL-SR
35.43%GOAT-Bench SR
61.32%EvoWorld First-Inspection SR
86.18%EvoWorld Search SR
70.15%EvoWorld SPL
MethodFindingDory
HL-SR ↑
GOAT-Bench
SR ↑
EvoWorld
First-Inspection ↑
EvoWorld
Search SR ↑
EvoWorld
SPL ↑
DynaMem30.3014.5445.3371.9456.83
EvolvingNav (ours)53.22 ± 3.8735.43 ± 2.1661.32 ± 3.2486.18 ± 2.0770.15 ± 1.74
Qualitative EvoWorld-Bench cases showing predictive search and evidence-aware replanning
Qualitative behavior. Arrival-time prediction proposes likely destinations, while newly visible evidence suppresses stale hypotheses and triggers recovery.
Ablation study of EvolvingNav components
Component analysis. Predictive transition modeling, visibility-aware filtering, and online evidence each contribute to robust first inspection and recovery.

Real-world evaluation

Across 64 matched LYNX M20 search episodes, EvolvingNav reaches 34.4% First-Inspection SR, 48.4% Search SR, and 24.3% Recovery SR, with 43.8 m mean travel.

Figure 5a · Indoor caseIndoor physical navigation sequence: the robot follows a corridor and finds a relocated cup
Indoor execution. The robot follows the remembered route, gathers new visibility-aware evidence, and locates the cup at its current position.
Figure 5b · Outdoor caseOutdoor physical navigation sequence: the robot rejects a stale parking location and finds the relocated car
Outdoor recovery. The robot verifies that the last-seen parking spot is stale, then searches the predicted current location.

Resources

CITE THIS WORK

BibTeX

@inproceedings{evolvingnav2027,
  title     = {Beyond the Remembered World: Predictive 4D Belief for Persistent Navigation in Evolving Worlds},
  author    = {Gao, Mingjian and Li, Zhaocheng and Huang, Haoyang and Zhang, Wenqiao and Niu, Yingjie and Zhou, Hao and Li, Chao and Li, Juncheng and Tang, Siliang and Zhuang, Yueting},
  booktitle = {International Conference on Learning Representations},
  year      = {2027}
}