isual imitation learning is a promising approach to training robot manipulation policies capable of
completing a wide variety of tasks. However, policies today remain brittle to viewpoint perturbations,
making deployment in diverse environments a challenge. We present a controlled empirical study of
which design choices allow visuomotor policies to generalize across viewpoints. We find that viewpoint
generalization improves when dense visual tokens are retained and the action head participates in
geometric reasoning. On a suite of simulated tasks that span a wide range of camera poses, we show that
these design choices yield a policy that remains performant across viewpoints. As a practical consequence, a
policy trained with these design choices also transfers zero-shot from simulation to the real world
under random camera configurations.
Keywords: Viewpoint generalization · Visual imitation learning · Sim-to-real
AGP is deliberately not a new architecture. Every component already appears in prior work — our contribution is isolating which of them are responsible for viewpoint robustness, under a common base policy, dataset, and action head.
All variants we study share a base policy: a Diffusion Transformer (DiT) action head trained with a flow-matching objective, which cross-attends from proprioception and noised-action tokens to image tokens and self-attends over the action chunk. We vary this base policy along three axes — whether visual features are compressed before the action head, which visual encoder produces them, and how token positions are encoded — to isolate which choices drive viewpoint generalization. The combination that wins on all three axes — dense tokens, a cross-view finetuned DINOv2 encoder, and robot-frame 4D rotary positional encodings — is what we call AGP.
We initialize from DINOv2 and finetune it with within-view and cross-view global self-attention layers — following the multi-view architecture of Depth Anything 3 — so features can fuse information across cameras without requiring 3D input. Critically, we keep the full set of dense image tokens for the action head to attend over: compressing them into a single global vector before the action head sharply hurts viewpoint robustness, even though it looks harmless under a fixed camera.
Each visual, proprioception, and action token gets a 4D coordinate (x, y, z, h) — its 3D position in the robot base frame plus a temporal index — and a 4D RoPE grounds every token in the geometry of the scene. Unlike image-space positional encodings, the geometric relationship between any visual token and the gripper is invariant to the camera pose.
A custom articulated-object task with hard-to-localize over-the-shoulder views, plus two tasks from Jiang et al.’s viewpoint-generalization benchmark.
We evaluate on a custom Cabinet task in PyBullet, inspired by ArticuBot, alongside Pick Place Can and Square Assembly from the “Do You Know Where Your Camera Is?” benchmark (Jiang et al., ICRA 2026). For Cabinet, we train a policy to imitate an expert motion planner opening a cabinet with a small knob handle from PartNet-Mobility. In contrast to front-facing setups that look toward the robot, we randomly sample two over-the-shoulder camera views per trajectory — from which the robot is harder to localize — plus a wrist-mounted camera that keeps the object visible even when the arm occludes it. We generate 989 demonstration trajectories (100k observation-action pairs) and evaluate on 100 held-out trajectories with freshly sampled camera poses; all three tasks randomize cameras during both training and evaluation.
Camera pose distribution (interactive). Each green frustum is one sampled over-the-shoulder camera; drag to orbit the scene and click a frustum to load that camera’s rendered view.
Comparing ways to give a policy camera information: AGP’s 4D RoPE beats Plücker-raymap and canonical-view conditioning on every task but one.
| Method | Fixed camera | Randomized cameras | |||
|---|---|---|---|---|---|
| Cabinet (Norm. Open) |
Cabinet | Pick Place Can | Square Assembly | Avg. | |
| Plücker Raymaps | 63.6 ±2.9 | 39.5 ±1.4 | 68.0 ±0.0 | 22.7 ±1.2 | 43.4 |
| Canonical Views | 59.1 ±3.2* | 52.8 ±3.2 | 62.0 ±0.0 | 32.0 ±0.0 | 48.9 |
| 4D RoPE (AGP) | 69.1 ±0.8 | 64.4 ±0.8 | 96.7 ±3.1 | 29.3 ±3.1 | 63.5 |
All methods share the same visual encoder and action-head architecture; only the camera-information mechanism differs. Avg. is computed over the three randomized-camera tasks. Canonical Views wins Square Assembly outright — the one task where a fixed reprojection viewpoint happens to suit the object layout — but AGP leads everywhere else and on average.
Ablating each design axis in turn — on Cabinet, Pick Place Can, and Square Assembly — isolates what drives AGP’s robustness, and several results are non-obvious from fixed-camera experiments alone.
Many policies squeeze visual observations into a single global vector before the action head. This is fine under a fixed camera, but under randomized cameras the action head loses the spatial information it needs to reason across viewpoints — on Pick Place Can, Diffusion Policy collapses from 56.1 (fixed) to 1.3 (randomized). Keeping all image tokens (no compression) is decisive.
| Method | Fixed camera | Randomized cameras | |||
|---|---|---|---|---|---|
| Cabinet | Cabinet | Pick Place Can | Square Assembly | Avg. | |
| Token compression | |||||
| Diffusion Policy | 56.1 | 33.2 | 1.3 | 2.7 | 12.4 |
| w/ Spatial Softmax | 55.2 | 33.4 | 56.0 | 22.7 | 37.4 |
| w/ Max Pooling | 60.0 | 35.1 | 61.3 | 22.7 | 39.7 |
| AGP — no compression | 69.1 | 64.4 | 96.7 | 29.3 | 63.5 |
Finetuning DINOv2 helps over a frozen backbone, and adding cross-view attention — letting the encoder fuse information across cameras, as in recent multi-view depth estimation models — adds a further gain. But the spread across encoders is modest next to the swings from removing token compression or fixing the positional encoding: encoder choice is a second-order effect here.
| Visual encoder | Fixed camera | Randomized cameras | |||
|---|---|---|---|---|---|
| Cabinet | Cabinet | Pick Place Can | Square Assembly | Avg. | |
| w/ ResNet-18 | 67.2 | 60.9 | 84.7 | 26.0 | 57.2 |
| w/ DINOv2 (frozen) | 63.6 | 59.8 | 88.0 | 29.0 | 58.9 |
| w/ DINOv2 (finetuned) | 67.6 | 59.3 | 93.3 | 32.0 | 61.5 |
| AGP — DINOv2, finetuned + cross-view | 69.1 | 64.4 | 96.7 | 29.3 | 63.5 |
Image-space positional encodings are fragile under randomized cameras: the same patch index can map to different 3D locations as the camera moves, so the encoding carries no consistent geometric meaning. Grounding every token in the robot frame with 4D RoPE lifts Cabinet (randomized) performance from 39.5 to 64.4.
| Method | Fixed camera | Randomized cameras | |||
|---|---|---|---|---|---|
| Cabinet | Cabinet | Pick Place Can | Square Assembly | Avg. | |
| ACT | 58.9 | 39.2 | 35.3 | 20.0 | 31.5 |
| w/ Sinusoidal | 59.1 | 39.5 | 68.7 | 24.0 | 44.1 |
| AGP — 4D RoPE | 69.1 | 64.4 | 96.7 | 29.3 | 63.5 |
Beyond the in-distribution comparison above: how AGP holds up under cameras it has never seen, against a 3D point-cloud policy, and under noisy calibration and depth at test time.
Out-of-distribution viewpoints. The comparisons above sample test cameras from the same distribution as training, which measures viewpoint robustness. To measure viewpoint generalization, we additionally evaluate Cabinet with a broader camera distribution that observes the scene from unseen elevated and oblique angles.
| Method | Cabinet (Norm. Open) |
|---|---|
| Plücker Raymaps | 29.8 ±5.2 |
| Canonical Views | 45.6 ±2.3 |
| AGP | 51.4 ±5.2 |
All policies drop compared to the in-distribution results in Table 1, but AGP retains its advantage over other camera-conditioned approaches under unseen viewpoints.
Comparison to a 3D policy. Policies that operate directly on point clouds are viewpoint-robust by construction, since their inputs can be expressed in a scene-centric frame. We compare against 3D Diffusion Policy (DP3), which consumes a ground-truth-segmented, fused point cloud of the object plus a 4-point gripper point cloud as proprioception.
| Method | Fixed camera | Randomized cameras |
|---|---|---|
| DP3 (point cloud, GT seg.) | 50.9 ±2.0 | 43.9 ±1.4 |
| AGP | 69.1 ±0.8 | 64.4 ±0.8 |
DP3 shows only a small gap between fixed and randomized cameras, confirming that point-cloud policies are viewpoint-robust by construction — but AGP achieves a comparably small gap while operating on RGB, and exceeds DP3’s performance under both settings.
Robustness to calibration and depth noise. AGP lifts pixels into the robot frame using camera extrinsics and depth — both imperfect in the real world — so we stress-test a Cabinet policy trained without any noise under increasing test-time perturbations of each.
A single policy trained entirely in simulation opens real microwaves under camera poses that were never seen during data generation.
We build on ArticuBot, an automated pipeline for generating demonstrations of a Franka robot opening articulated objects. From its set of 322 objects we select five microwaves, generate 2,358 demonstrations and over 100,000 state-action pairs, and render them in IsaacLab with extensive domain randomization over microwave colors and textures, the robot, the table, and the background. We render two over-the-shoulder views and one wrist view per trajectory, with camera intrinsics matched to the ZED2 and ZED-mini cameras used in our real-world setup, and augment depth maps to simulate sensor noise and edge artifacts. At deployment, real-world depth is obtained with Fast-FoundationStereo and cropped to the workstation to match the training distribution; camera poses are unknown during data generation but within the randomized training distribution. We mount three cameras and evaluate all three possible camera pairs, for 30 trials per policy (10 per pair), reporting grasp success (visually confirmed every trial) and normalized opening.
Zero-shot sim-to-real on the real robot under randomized camera viewpoints.
Use the ‹ / › arrows to see more examples.
Representative failure modes of the zero-shot sim-to-real policy on the real robot.
Use the ‹ / › arrows to see more examples.
We study viewpoint generalization, relaxing the common constraint of fixed cameras at training and deployment. Our controlled study identifies what drives it: retaining dense scene representations instead of compressing them before the action head, and grounding every visual, proprioception, and action token in the geometry of the scene. We instantiate these findings as AGP — deliberately not a new architecture, but the combination of existing ingredients responsible for viewpoint robustness — a policy that is performant for all views in the camera distribution, approaches fixed-camera performance, and transfers zero-shot from simulation to a real robot under random cameras.
Limitations. AGP requires calibrated RGB-D input at both training and deployment — depth (sensed or estimated) and camera extrinsics in the robot base frame are needed to lift each pixel into the shared coordinate system the positional encoding relies on. Removing this requirement, e.g. by jointly estimating extrinsics or operating on monocular RGB with learned depth priors, is an important direction. We also demonstrate sim-to-real as a proof of concept on a single object category (microwaves); scaling these findings to large, heterogeneous real-world datasets such as DROID remains an open question.
Randomization details: initial-state and camera-pose randomization for demonstration generation, visual domain randomization for sim-to-real, and visualizations of the Cabinet and microwave datasets.
PDF not showing? Open the appendix in a new tab.
Coming soon.