An Empirical Study on What Matters for Viewpoint-Generalizable Policies in Visual Imitation Learning

VGP teaser: sampling camera poses, generating data in simulation, and zero-shot transfer to a real robot under random cameras.
Top Left: We sample a wide distribution of camera poses in simulation — randomizing cameras costs a standard 2D policy 22.9 points of normalized opening performance that AGP recovers. Middle: We generate demonstrations under these randomized views and train AGP, which grounds visual, proprioception, and action tokens in a shared robot frame. Right: The policy transfers zero-shot to a real robot under random camera configurations.

Abstract

Visual imitation learning is a promising approach to training robot manipulation policies capable of completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations, making deployment in diverse environments a challenge. We present a controlled empirical study of which design choices allow visuomotor policies to generalize across viewpoints. We find that viewpoint generalization improves when dense visual tokens are retained and the action head participates in geometric reasoning. On a suite of simulated tasks that span a wide range of camera poses, we show that these design choices yield a policy that remains performant across viewpoints. As a practical consequence, a policy trained with these design choices also transfers zero-shot from simulation to the real world under random camera configurations.

Keywords: Viewpoint generalization · Visual imitation learning · Sim-to-real

AGP — What Our Study Says Matters

AGP is deliberately not a new architecture. Every component already appears in prior work — our contribution is isolating which of them are responsible for viewpoint robustness, under a common base policy, dataset, and action head.

All variants we study share a base policy: a Diffusion Transformer (DiT) action head trained with a flow-matching objective, which cross-attends from proprioception and noised-action tokens to image tokens and self-attends over the action chunk. We vary this base policy along three axes — whether visual features are compressed before the action head, which visual encoder produces them, and how token positions are encoded — to isolate which choices drive viewpoint generalization. The combination that wins on all three axes — dense tokens, a cross-view finetuned DINOv2 encoder, and robot-frame 4D rotary positional encodings — is what we call AGP.

VGP architecture diagram: multi-view DINOv2 encoder with within-view and cross-view attention feeding a 4D-RoPE DiT action head.
Base policy and design axes, in the AGP configuration. V calibrated RGB views are encoded by a backbone initialized from DINOv2, applying within-view self-attention and cross-view global self-attention to produce dense multi-view image tokens (visual-encoder & token-compression axes). Every image, proprioception, and action token is assigned a spatiotemporal position (x, y, z, h) and attended over with a 4D rotary positional encoding (positional-encoding axis). An N-block DiT, shared across all variants, cross-attends from proprioception/action tokens to image tokens and self-attends over the action chunk.
1

Dense tokens from a cross-view encoder

We initialize from DINOv2 and finetune it with within-view and cross-view global self-attention layers — following the multi-view architecture of Depth Anything 3 — so features can fuse information across cameras without requiring 3D input. Critically, we keep the full set of dense image tokens for the action head to attend over: compressing them into a single global vector before the action head sharply hurts viewpoint robustness, even though it looks harmless under a fixed camera.

2

Robot-frame 4D rotary attention

Each visual, proprioception, and action token gets a 4D coordinate (x, y, z, h) — its 3D position in the robot base frame plus a temporal index — and a 4D RoPE grounds every token in the geometry of the scene. Unlike image-space positional encodings, the geometric relationship between any visual token and the gripper is invariant to the camera pose.

Tasks: Cabinet-Opening and “Do You Know Where Your Camera Is?”

A custom articulated-object task with hard-to-localize over-the-shoulder views, plus two tasks from Jiang et al.’s viewpoint-generalization benchmark.

We evaluate on a custom Cabinet task in PyBullet, inspired by ArticuBot, alongside Pick Place Can and Square Assembly from the “Do You Know Where Your Camera Is?” benchmark (Jiang et al., ICRA 2026). For Cabinet, we train a policy to imitate an expert motion planner opening a cabinet with a small knob handle from PartNet-Mobility. In contrast to front-facing setups that look toward the robot, we randomly sample two over-the-shoulder camera views per trajectory — from which the robot is harder to localize — plus a wrist-mounted camera that keeps the object visible even when the arm occludes it. We generate 989 demonstration trajectories (100k observation-action pairs) and evaluate on 100 held-out trajectories with freshly sampled camera poses; all three tasks randomize cameras during both training and evaluation.

drag to orbit · scroll to zoom · click a camera
Couldn’t start the 3D viewer — WebGL appears to be unavailable in this browser.
Rendered view from the selected camera.
Click a frustum to load its view…

Camera pose distribution (interactive). Each green frustum is one sampled over-the-shoulder camera; drag to orbit the scene and click a frustum to load that camera’s rendered view.

Results

Comparing ways to give a policy camera information: AGP’s 4D RoPE beats Plücker-raymap and canonical-view conditioning on every task but one.

Table 1. Approaches to incorporating camera information for viewpoint generalization on opening an articulated cabinet and on the “Do You Know Where Your Camera Is?” benchmark. Unless marked fixed camera, cameras are randomized during training and evaluation. Mean ± std. over 3 seeds. Best shaded, second-best underlined. *Under the fixed-camera setting, canonical views coincide with the fixed camera views, so Canonical Views is identical to the “w/ Sinusoidal” ablation in Table 2 and we report the same result.
Method Fixed camera Randomized cameras
Cabinet
(Norm. Open)
Cabinet Pick Place Can Square Assembly Avg.
Plücker Raymaps 63.6 ±2.939.5 ±1.468.0 ±0.022.7 ±1.243.4
Canonical Views 59.1 ±3.2*52.8 ±3.262.0 ±0.032.0 ±0.048.9
4D RoPE (AGP) 69.1 ±0.864.4 ±0.896.7 ±3.129.3 ±3.163.5
Best    x Second-best

All methods share the same visual encoder and action-head architecture; only the camera-information mechanism differs. Avg. is computed over the three randomized-camera tasks. Canonical Views wins Square Assembly outright — the one task where a fixed reprojection viewpoint happens to suit the object layout — but AGP leads everywhere else and on average.

What Matters for Viewpoint Generalization?

Ablating each design axis in turn — on Cabinet, Pick Place Can, and Square Assembly — isolates what drives AGP’s robustness, and several results are non-obvious from fixed-camera experiments alone.

Finding 1 Compressing visual information hurts action generation

Many policies squeeze visual observations into a single global vector before the action head. This is fine under a fixed camera, but under randomized cameras the action head loses the spatial information it needs to reason across viewpoints — on Pick Place Can, Diffusion Policy collapses from 56.1 (fixed) to 1.3 (randomized). Keeping all image tokens (no compression) is decisive.

Table 2. Visual token compression, positional encoding, and visual encoder ablations, varied relative to AGP. First column: fixed camera. Remaining columns: cameras randomized during training and evaluation. Mean over 3 seeds.
Method Fixed camera Randomized cameras
CabinetCabinetPick Place CanSquare AssemblyAvg.
Token compression
Diffusion Policy56.133.21.32.712.4
w/ Spatial Softmax55.233.456.022.737.4
w/ Max Pooling60.035.161.322.739.7
AGP — no compression69.164.496.729.363.5

Finding 2 Cross-view attention helps, but matters less than compression or positional encoding

Finetuning DINOv2 helps over a frozen backbone, and adding cross-view attention — letting the encoder fuse information across cameras, as in recent multi-view depth estimation models — adds a further gain. But the spread across encoders is modest next to the swings from removing token compression or fixing the positional encoding: encoder choice is a second-order effect here.

Table 2 (cont.). Visual encoder ablation.
Visual encoder Fixed camera Randomized cameras
CabinetCabinetPick Place CanSquare AssemblyAvg.
w/ ResNet-1867.260.984.726.057.2
w/ DINOv2 (frozen)63.659.888.029.058.9
w/ DINOv2 (finetuned)67.659.393.332.061.5
AGP — DINOv2, finetuned + cross-view69.164.496.729.363.5

Finding 3 Robot-frame positional encoding drives robustness

Image-space positional encodings are fragile under randomized cameras: the same patch index can map to different 3D locations as the camera moves, so the encoding carries no consistent geometric meaning. Grounding every token in the robot frame with 4D RoPE lifts Cabinet (randomized) performance from 39.5 to 64.4.

Table 2 (cont.). Positional encoding ablation.
Method Fixed camera Randomized cameras
CabinetCabinetPick Place CanSquare AssemblyAvg.
ACT58.939.235.320.031.5
w/ Sinusoidal59.139.568.724.044.1
AGP — 4D RoPE69.164.496.729.363.5

Generalization and Robustness

Beyond the in-distribution comparison above: how AGP holds up under cameras it has never seen, against a 3D point-cloud policy, and under noisy calibration and depth at test time.

Out-of-distribution viewpoints. The comparisons above sample test cameras from the same distribution as training, which measures viewpoint robustness. To measure viewpoint generalization, we additionally evaluate Cabinet with a broader camera distribution that observes the scene from unseen elevated and oblique angles.

Camera distribution figure: green frustums for the training distribution, blue for real-world camera poses, orange for out-of-distribution evaluation poses.
Camera distribution. The training/in-distribution camera distribution (green), the poses used in the real world (blue, see Sim-to-Real below), and the poses used for evaluating out-of-distribution generalization (orange).
Table 3. Cabinet with out-of-distribution cameras (unseen elevated/oblique angles).
MethodCabinet (Norm. Open)
Plücker Raymaps29.8 ±5.2
Canonical Views45.6 ±2.3
AGP51.4 ±5.2

All policies drop compared to the in-distribution results in Table 1, but AGP retains its advantage over other camera-conditioned approaches under unseen viewpoints.

Comparison to a 3D policy. Policies that operate directly on point clouds are viewpoint-robust by construction, since their inputs can be expressed in a scene-centric frame. We compare against 3D Diffusion Policy (DP3), which consumes a ground-truth-segmented, fused point cloud of the object plus a 4-point gripper point cloud as proprioception.

Table 4. Comparison against DP3 on the Cabinet task.
MethodFixed cameraRandomized cameras
DP3 (point cloud, GT seg.)50.9 ±2.043.9 ±1.4
AGP69.1 ±0.864.4 ±0.8

DP3 shows only a small gap between fixed and randomized cameras, confirming that point-cloud policies are viewpoint-robust by construction — but AGP achieves a comparably small gap while operating on RGB, and exceeds DP3’s performance under both settings.

Robustness to calibration and depth noise. AGP lifts pixels into the robot frame using camera extrinsics and depth — both imperfect in the real world — so we stress-test a Cabinet policy trained without any noise under increasing test-time perturbations of each.

Line chart: Cabinet normalized opening performance as a function of extrinsic calibration noise, with a dashed line marking expected real-world calibration error.
Fig. 4. Under calibration noise, AGP degrades gracefully and stays robust around the calibration error we expect in the real world (σtrans ≈ 1 cm, σrot ≈ 1°).
Line chart: Cabinet normalized opening performance as a function of additive Gaussian depth noise.
Fig. 5. Performance degrades faster under depth noise — training with depth or calibration noise would likely improve robustness further.

Zero-Shot Sim-to-Real Transfer

A single policy trained entirely in simulation opens real microwaves under camera poses that were never seen during data generation.

We build on ArticuBot, an automated pipeline for generating demonstrations of a Franka robot opening articulated objects. From its set of 322 objects we select five microwaves, generate 2,358 demonstrations and over 100,000 state-action pairs, and render them in IsaacLab with extensive domain randomization over microwave colors and textures, the robot, the table, and the background. We render two over-the-shoulder views and one wrist view per trajectory, with camera intrinsics matched to the ZED2 and ZED-mini cameras used in our real-world setup, and augment depth maps to simulate sensor noise and edge artifacts. At deployment, real-world depth is obtained with Fast-FoundationStereo and cropped to the workstation to match the training distribution; camera poses are unknown during data generation but within the randomized training distribution. We mount three cameras and evaluate all three possible camera pairs, for 30 trials per policy (10 per pair), reporting grasp success (visually confirmed every trial) and normalized opening.

Bar charts: grasp success and normalized opening for Plücker Raymaps, Canonical Views, and VGP.
Sim-to-real transfer on microwaves. Policies trained in simulation and evaluated zero-shot in the real world under randomly sampled camera poses. Averaged over 30 trials across 3 camera configurations (10 trials per configuration).
Grid of simulated microwave-opening rollouts in IsaacLab with domain randomization.
Simulation training data (IsaacLab). Microwave-opening demonstrations rendered with extensive domain randomization over object textures, robot, table, and background. Each row is a distinct scene; columns are timesteps.

Real-World Rollouts — Successes

Zero-shot sim-to-real on the real robot under randomized camera viewpoints.

Failure Cases

Representative failure modes of the zero-shot sim-to-real policy on the real robot.

Takeaways

We study viewpoint generalization, relaxing the common constraint of fixed cameras at training and deployment. Our controlled study identifies what drives it: retaining dense scene representations instead of compressing them before the action head, and grounding every visual, proprioception, and action token in the geometry of the scene. We instantiate these findings as AGP — deliberately not a new architecture, but the combination of existing ingredients responsible for viewpoint robustness — a policy that is performant for all views in the camera distribution, approaches fixed-camera performance, and transfers zero-shot from simulation to a real robot under random cameras.

Limitations. AGP requires calibrated RGB-D input at both training and deployment — depth (sensed or estimated) and camera extrinsics in the robot base frame are needed to lift each pixel into the shared coordinate system the positional encoding relies on. Removing this requirement, e.g. by jointly estimating extrinsics or operating on monocular RGB with learned depth priors, is an important direction. We also demonstrate sim-to-real as a proof of concept on a single object category (microwaves); scaling these findings to large, heterogeneous real-world datasets such as DROID remains an open question.

Appendix

Randomization details: initial-state and camera-pose randomization for demonstration generation, visual domain randomization for sim-to-real, and visualizations of the Cabinet and microwave datasets.

PDF not showing? Open the appendix in a new tab.

BibTeX

Coming soon.