3D maps are fundamental for many robotic tasks, but storing them at a uniform fine resolution is wasteful: most of the scene never matters for the task at hand. Existing adaptive-mapping methods either rely on hand-tuned heuristics or ignore what the map is actually for. We instead frame resolution allocation as a sequential decision problem and train a reinforcement-learning policy that dynamically refines or coarsens the map under an explicit memory budget, optimising directly for downstream navigation rather than reconstruction fidelity. Built on top of PEANUT and evaluated in Habitat, our learned policy retains an 80% point-goal success rate while using 87% less memory on average than a uniformly fine map, and achieves comparable usage to an all-coarse map at far higher success.
A recording of the learned policy navigating an unseen point-goal episode while adapting the 3D map resolution online.
We build our agent on top of PEANUT, and our core contribution is to replace its dense 3D map with a sparse, hierarchical, adaptive-resolution voxel map whose resolution is controlled by a learned policy under an explicit memory budget. The complete perception–mapping–planning pipeline runs as a closed per-frame loop: the simulator emits a new RGB-D frame, the 3D map is updated, a reinforcement-learning policy adapts the map resolution, the 3D map is projected down to a 2D top-down map, the 2D map drives planning, and the resulting action moves the robot, producing the next frame.
End-to-end perception–mapping–planning pipeline. The simulator (Habitat), semantic perception, and prediction–planning stages are reused; the sparse hierarchical 3D map, the adaptive-resolution policy, and the incremental 3D→2D projection are our contribution.
Sparse hierarchical 3D map. PEANUT represents the scene with a dense voxel grid whose memory cost grows with the whole bounded volume regardless of occupancy. We re-implement the map from scratch as a sparse, hierarchical, multi-resolution grid: only observed voxels are stored, and every voxel lives at one of three resolution levels — fine (5cm), medium (10cm), or coarse (20cm).
Per-voxel evidence, carving, and pruning. Every voxel accumulates a hit count, per-class semantic scores, a free-space count, and an RGB color sum over the whole episode. Free-space carving increments the free-space count of every voxel a camera ray traverses before stopping, and a voxel repeatedly seen empty is pruned, keeping the sparse map clean and its footprint low.
Adaptive-resolution policy. Between the map update and the 2D projection, a learned RL policy adapts voxel resolution — refining selected regions (Split), leaving others unchanged (Keep), or collapsing fine voxels into coarser ones (Merge) — to keep the most useful detail while respecting a fixed memory budget.
Incremental 3D→2D projection. Rather than re-projecting the entire grid every frame, we develop an incremental "delta-buffer" projection that tracks only the pixels affected by the latest 3D map changes, keeping the 2D map an exact, always-consistent projection while only paying for what actually changed.
We cast online resolution selection as a Markov decision process and train the policy with Proximal Policy Optimization (PPO). At each step the agent observes a compact 21-dimensional global summary of the map (memory fill, budget, a 16-bin histogram of per-block scores, and the resolution-level fractions) and outputs a pair of thresholds on a per-block score that combines whether the block lies on the route to the goal (weight 0.8) with how much refinement headroom it has (weight 0.2). These thresholds partition all occupied blocks into three groups — Merge, Keep, Split — applied uniformly across the map. The reward sums a budget term that rewards staying just below the memory budget with two route-damage penalties that discourage lengthening or blocking the path to the goal.
Training curves. Top: mean success rate. Middle: median episode reward. Bottom: median memory usage. The policy first over-spends its budget, then learns to compress while restoring a constant 100% training success rate.
On 20 unseen point-goal episodes, the learned policy reaches an 80% success rate while using essentially the same memory as an all-coarse baseline, which only reaches 65%. Compared to an all-fine baseline, our policy cuts memory usage by 87% on average at the cost of a moderate drop in success rate.
| SR ↑ | SPL ↑ | Mem (MB) ↓ | Steps | |
|---|---|---|---|---|
| learned | 0.80 | 0.60 | 3.9 | 78 |
| all_fine | 1.00 | 0.74 | 31.1 | 65 |
| all_coarse | 0.65 | 0.48 | 3.2 | 106 |
learned
all_fine
all_coarse
Although trained on a single episode with a single goal, the policy generalises to unseen episodes and goals: it has effectively learned how to balance resolution and memory given the score distribution, rather than memorising the training episode.
The policy was trained on a single scene and a single memory budget (8MB), so it does not generalise to other budgets out of the box, and randomizing the budget across episodes proved unstable to train without further reward shaping. We also tested generalisation to object-goal navigation, where the policy underperforms an all-coarse baseline — likely because the downstream goal-prediction network goes out-of-distribution when fed non-uniform-resolution maps.
| SR ↑ | SPL ↑ | Mem (MB) ↓ | |
|---|---|---|---|
| all_coarse | 0.60 | 0.38 | 4.0 |
| policy | 0.45 | 0.31 | 7.1 |
| all_fine | 0.65 | 0.42 | 48.8 |
Promising future directions include training on randomized scenes for a more robust policy, and testing the pipeline on agents whose performance is more strongly influenced by the quality of the 3D map.
Adaptive Mapping. To improve efficiency, adaptive mapping techniques adjust voxel resolution to scene content. Geometry-based methods refine regions of high curvature or structural complexity and coarsen flat areas, improving geometric fidelity near surfaces but ignoring semantic relevance. Semantic-based methods add object information to guide subdivision: Panoptic Multi-TSDFs reconstruct separate TSDF volumes per instance, while MAP-ADAPT combines geometric and semantic cues into a single multi-resolution TSDF map by refining user-specified semantic classes. MAP-ADAPT forms the basis of our map representation, but its allocation rests on manually defined class-to-resolution rules that are fixed offline and carry no notion of which regions matter for a given task.
Learning-Based Adaptivity. Reinforcement learning has been applied extensively to active exploration and SLAM, where a policy selects actions or viewpoints that improve coverage or localization. These methods learn how to move to acquire information, but leave the underlying map a passive, fixed-resolution structure. UNRL is, to our knowledge, the first to apply reinforcement learning to the structure of the map itself, learning a budget-aware subdivision policy driven by an entropy signal; its reward, however, targets reconstruction fidelity. We adopt the same budget-aware RL view of map adaptation, but optimize the policy for a downstream navigation objective, instantiated within PEANUT, which relies on a dense, uniformly fine-resolution voxel map.
@misc{baggi2026goaldriven,
author = {Baggi, Lorenzo and De Negri, Marco and De Palma, Vassili and Salvatore, Alessio},
title = {Goal-driven 3D Map Generation with Reinforcement Learning},
howpublished = {3D Vision course project report, ETH Z\"urich},
year = {2026},
}