Topics
17 curated papers across all published weeks — sorted by disruptiveness score.
Reliable surgical planning requires models to anticipate not only how instruments will move, but also how the operative visual state will evolve together with…
Read free explainer →
2026-W36
The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets a…
2026-W35
Autonomous laboratories hold great promise for accelerating scientific discovery. To achieve this vision, robots are supposed to dexterously manipulate diverse…
2026-W34
We train a large-scale transformer policy on heterogeneous humanoid datasets that zero-shot adapts to new morphologies and terrains, combining imitation, reinf…
2026-W30
Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interaction, supporting policy…
An autonomous laboratory combines literature-trained language models, Bayesian optimization, and robotic synthesis to discover three previously unreported soli…
We integrate high-resolution tactile sensing into vision-language-action models, enabling contact-rich assembly tasks that pure vision policies fail, with stro…
Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusi…
We demonstrate onboard optical navigation using lunar landmarks and star trackers that maintains kilometer-level accuracy throughout cis-lunar transfer without…
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-ground…
We present PAC-MAN, a perception-aware CBF-RL framework that couples control-barrier safety with deployment-realistic onboard sensing for whole-body humanoid d…
2026-W31
In contact-rich manipulation, action multimodality and reactivity dominate different stages of a single episode. Before contact, multiple trajectories might be…
Pretraining navigation diffusion policies rely on large-scale expert demonstrations. These data are typically generated by a fully-informed oracle planner suit…
Vision-language-action models (VLAs) are moving rapidly towards deployment as general-purpose manipulation policies, but we currently lack basic tools for unde…
2026-W33
Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-camera setups: close…
We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active skill learning algorit…
We present \textbf{DreamX-Phi 1.0}, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and…