Vision-Language-Action Models

JunoTaming Predictive Latents
for Vision-Language-Action Models

Yuchen Zhu1,2·Chenyi Xu2,3·Yulin Zhang4·Gang Xu2·Wentao Zhu2

1 University of Science and Technology of China   2 Eastern Institute of Technology, Ningbo
3 Hangzhou Dianzi University   4 ShanghaiTech University

One action-conditioned JEPA across pretraining, policy learning, and deployment.
Real-robot rollout 1.5× speed
Juno real-robot rollout putting a red cube into a blue bowl
AgileX Cobot Magic · nominal sceneJuno
70–75%success under
real-world shifts
Abstract

Predictive latents are only useful when they stay tied to the control problem.

Joint-embedding predictive architectures (JEPAs) predict masked or future observations in representation space, offering a natural source of predictive latents for vision-language-action (VLA) models. Yet making these latents useful across pretraining, policy learning, and deployment requires addressing three failures: mismatch with embodiment-specific control, interference with action learning, and teacher miscalibration under distribution shifts. We introduce Juno, a unified framework built around one action-conditioned JEPA that serves as a control-aligned representation backbone, a predictive teacher, and an adaptable dynamics model. During pretraining, we train it on embodiment-matched trajectories and use a dynamic CLS loss to transfer motion-weighted patch dynamics to a compact global state. During policy learning, we fuse current-frame JEPA patches into VLA perception and use a decoupled reasoning branch with separate transformation parameters to distill future latent states for action generation. During deployment, we adapt the world model on all observed transitions, including failed rollouts, freeze the adapted teacher, and re-align the policy on verified executions using LoRA adapters and a trainable action head, without expert corrections or task rewards. On SimplerEnv, Juno raises average success from 60.9% to 68.5% over Qwen3GR00T, the strongest baseline, and test-time adaptation further reaches 72.7%; on a real robot, it retains 70%–75% success under background, height, and object shifts where the base policy collapses to 0%.

Control-aligned pretrainingDecoupled predictive reasoningDynamics-first test-time adaptation
The idea

One JEPA, three jobs.

Juno follows a single predictive representation from learning the world to acting in it.

Juno overview: feature pretraining, performance gains, real-world evaluation, and test-time adaptation
01

Control-aligned
pretraining

Match the JEPA to embodiment-specific motion, then make the global state respond to local interaction.

↘
02

Decoupled
reasoning

Fuse current-frame patches and distill future latent states in a branch that still drives actions.

↘
03

Dynamics-first
adaptation

Use every transition to adapt dynamics; trust verified executions when updating the policy.

↘
Architecture

Predict, align, act.

The predictive branch is trained to be action-decodable, not merely visually predictive.

Juno architecture pipeline from JEPA pretraining to predictive reasoning and Juno-TTT
A

Dynamic CLS. Motion-weighted patch changes supervise the compact state token.

B

Mixture-of-transformers branch. Future queries share causal attention while keeping separate transforms.

C

Failure-aware TTT. Failed actions still reveal consequences, so they remain useful for dynamics adaptation.

Evidence

Robustness shows up in the numbers.

Reported paper results across simulation, real-robot transfer, and deployment adaptation.

68.5%SimplerEnv average success
Juno · frozen policy
+7.6 pts vs Qwen3GR00T
72.7%SimplerEnv average success
Juno-TTT · after adaptation
+11.8 pts vs Qwen3GR00T
59.6%RoboCasa-GR1 average
24 tabletop tasks
21 of 24 tasks improved
70–75%Real-robot success under
background, height, object shifts
baseline: 0%
Real robot · frozen policy · 20 trials per condition
ConditionQwen3GR00TJuno
In-domain40%100%
Background shift0%75%
+ Height shift0%70%
+ Object shift0%70%
Real robot evaluation conditions: in-domain, background, height and object shifts
Evaluation conditions on the AgileX Cobot Magic platform.
Real-robot rollouts

See the policy meet the shift.

All embedded rollouts are accelerated to 1.5×, kept at their native 16:9 framing, and shown as looping GIFs.

Nominal scene: put the red cube into the blue bowlJuno · success

Nominal scene

Put the red cube into the blue bowl.

Qwen3GR00T failure example from the real-robot evaluationBaseline · fail

Qwen3GR00T failure example

A representative failure from the real-robot evaluation.

Background shift rollout with a new tablecloth, cube, and bowl75% success

Background shift

New tablecloth, cube, and bowl appearance.

Height and object shift rollout on a raised work plane70% success

Height + object shift

Raised work plane and a new target object.

Juno-TTT

Adaptation recovers control under harder observations.

World-model adaptation uses every transition; policy re-alignment uses verified executions only.

Gaussian observation noise before Juno-TTTFrozen · 40%

Gaussian observation noise

Before Juno-TTT.

Gaussian observation noise after Juno-TTT adaptationJuno-TTT · 65%

Gaussian observation noise

After dynamics adaptation and policy re-alignment.

Dynamic lighting before Juno-TTTFrozen · 55%

Dynamic lighting

Before Juno-TTT.

Dynamic lighting after Juno-TTT adaptationJuno-TTT · 70%

Dynamic lighting

After adaptation to the deployment stream.

BibTeX
@misc{zhu2026junotamingpredictivelatents,
      title={Juno: Taming Predictive Latents for Vision-Language-Action Models},
      author={Yuchen Zhu and Chenyi Xu and Yulin Zhang and Gang Xu and Wentao Zhu},
      year={2026},
      eprint={2610.09940},
      archivePrefix={arXiv},
      primaryClass={cs.RO},
      url={https://arxiv.org/abs/2610.09940},
}