[ Research ]
Dyna-2.1: A Physical Agent for End-to-End Workflows
To be useful, a general-purpose robot needs to take over entire workflows without frequent human intervention. Dyna-2.1 introduces our physical agent, a semi-humanoid robot named Taku with a learning system that leverages human experience for enhanced teachability.
Category:
Research
Author:
Dyna Robotics
Date:
September 29, 2026
Read:
19 min
Dyna-2.1: The first physical agent to automate an end-to-end workflow—one hour, uncut.
Listen to our team discussing the breakthroughs.
1. The value of a general-purpose robot lives in workflows, not tasks
Expanding the robot’s workspace, because workflows are done by humans today. The robot has to cover the whole workspace of the role it fills, using base, torso and arms together at every station.
Teachability, because each customer’s workflow is different, changes over time and chains hundreds of subtasks that each need more 9’s. The robot has to learn new skills fast and make them reliable, from human data, simulation and the corrections it collects on the job.
Reasoning, because workflows are non-linear and require conditional decision-making.
2. Dyna-2.1: Introducing the Physical Agent and the Taku Robot
2.1. Meet Taku
Video 2.1: introducing Taku.
Video 2.2: whole-body manipulation, across the workspace. Clockwise from top left: reaching deep into a dryer, picking up a towel, opening a refrigerator, and placing a stack on a low shelf.
2.2. The Physical Agent
[ WORKFLOW RELIABILITY ]
Figure 2.2. Workflow reliability. The chance that a whole cycle finishes without help, for a given per-subtask success rate. Drag the slider; subtask counts are estimates.
Figure 2.3. Laundry Room Workflow. One laundry cycle: thirteen decisions the robot makes, between long runs of loco-dexterous manipulation. The orchestrator is instructed to follow these decisions (Section 3.2).
A whole-body controller, trained with reinforcement learning in simulation, turns task-space target trajectories for the wrists, elbows, chest and footprint (the Unified Robot Representation, URR; Section 3.1.1) into joint targets and wheel velocities at 100 Hz.
The DYNA-2 policy, an improved version of our DYNA-2 world-action model [2], turns the current step into whole-body target trajectories to be tracked by the whole-body controller.
A workflow orchestrator, a vision-language model, tracks the workflow, decides the next step and steers DYNA-2 to carry it out.
[ INGREDIENTS ]
Figure 2.4. Ingredients of a physical agent. Taku and the three model layers, each on its own clock. A person can produce URR target trajectories too, so human data enters where the policy’s commands do.
3. Building the Physical Agent: Contributions Across the Full Stack
3.1. Teachability: Each Skill from the Cheapest Data That Teaches It
Learning a new skill: server servicing.
Learning a new skill: retrieving a drink.
Make human data usable for robots (Section 3.1.1): a shared representation lets human recordings train the robot.
Training controls that maximize teachability (Section 3.1.2): the controller gets better with every demonstration we collect, and a better controller records better data for the policy.
Keep old robot data usable (Section 3.1.3): a hardware revision retrains the controller, not the policy.
3.1.1 Make human data usable for robots
Embodiment-agnostic. The same format describes a human, humanoid, semi-humanoid or tabletop bimanual robot, using more or fewer body components according to what is present or observed.
Captures key task-space constraints. These poses specify hand placement, arm configuration, torso posture and the body’s placement in the workspace. These are the spatial constraints that matter most for coordinated manipulation.
Lossless across equivalent action formats. The represented component poses can be recovered from equivalent formats, such as relative-pose actions, provided the reference frame and starting pose are retained.
[ URR TRACKING ]
Figure 3.1. URR tracking. One teleoperation episode, replayed from the robot's log. The translucent target is the URR command: grippers at the hand targets, the torso at the chest target, the base at the footprint target and elbows at the elbow targets. The solid robot is the measured joint state and odometry as the whole-body controller tracks that command. The robot starts hidden; toggle each layer and drag to orbit.
[ THE DATA LIFECYCLE ]
Figure 3.2. The data lifecycle around URR. Every data source enters as URR for training; at inference, model actions decode to URR target trajectories for the controller.
Learning from human data. Taku reproducing egocentric human demonstrations in simulation, in real time. Left: Taku; right: the person’s head camera. Each recording is converted to URR commands and tracked by the RL whole-body controller.
3.1.2 Training controls that maximize teachability
Video 3.3: the controller, trained at scale. Thousands of simulated Taku robots training in parallel in NVIDIA Isaac Sim.
Figure 3.4. Overdamped versus underdamped. A 15 cm hand step with equal total tracking error: the underdamped response overshoots and rings, the overdamped one never passes the target. The responses are constructed.
3.1.3 Keep old robot data usable
Explicit controller conditioning. We build on DYNA-2’s language steerability [2] to teach the policy each controller’s characteristics. A metadata prompt tells the DYNA-2 policy which controller each robot demonstration came through, so one checkpoint learns how each controller moves and can be steered to the one used at deployment.
Implicit context through asynchronous inference. We also condition the model on future commands during asynchronous inference. This command context provides implicit information about the controller’s dynamics, complementing the explicit controller metadata.
3.1.4 The result: general physical capability and recovery
Video 3.5: physical skills across tasks. Clockwise from top left: pressing a button, placing a stack on a high shelf, inserting a server, and picking a laundry pod. These examples show dexterity and precision across different manipulation tasks.
Video 3.6: robustness when manipulation does not go to plan. Both clips autonomous, at 1× speed. In one, while Taku turns on the washer, a person perturbs the dials and pulls Taku’s gripper off the buttons; Taku returns to the panel and finishes setting the correct cycle. In the other, Taku recovers from a double-towel scenario.
3.2. Reasoning: An Orchestrator that Runs the Workflow
Ling Li, Ming Qin, and Chet Bhateja discuss the design of reasoning and language following, interwoven with footage of the reasoning system guiding the robot through the workflow.
3.2.1 Built on a vision-language model
3.2.2 Reason about the environment and choose the next step
Video 3.7: from folding to dryer unloading. Taku pauses folding to attend to the dryer, adapting its next step as the workflow changes.
[ THE ORCHESTRATOR ]
Figure 3.8. The orchestrator at work. Top: the loop. Bottom: an illustrative stretch of a shift, one decision per row.
3.2.3 Long-term memory
4. Looking Forward
References
Dyna Robotics. Dynamism v1 (DYNA-1) Model: A Breakthrough in Performance and Production-Ready Embodied AI. dyna.co/research/dyna-1, June 2025.
Dyna Robotics. Dyna-2: A 1-Million-Hour Scaling Law for World-Action Models. dyna.co/dyna-2, August 2026.
Dyna Robotics. Not Just a Model, But a Product. dyna.co/research/scaling-customer-deployments, August 2026.
Hietpas. Hotel Laundry Can Maximize Effectiveness with Right Equipment Mix (Part 1 of 2). American Laundry News, September 2011. americanlaundrynews.com.
Dyna Robotics. What 10 Months in Production Taught Us About the Robotics “Bubble”. dyna.co/news/robotics-bubble, May 2026.
Build AI. Egocentric-10K and Egocentric-100K datasets (10,000 and 100,405 hours of egocentric factory-worker video, Apache 2.0). huggingface.co/datasets/builddotai/Egocentric-100K, 2025.
Chi et al. Universal Manipulation Interface: In-The-Wild Robot Teaching Without In-The-Wild Robots. Project page, 2024.
Araujo et al. Retargeting Matters: General Motion Retargeting for Humanoid Motion Tracking. arXiv:2510.02252, 2025.
He et al. OmniH2O: Universal and Dexterous Human-to-Humanoid Whole-Body Teleoperation and Learning. arXiv:2406.08858, 2024.
Yang et al. OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction. arXiv:2509.26633, 2025.
Ha et al. UMI on Legs: Making Manipulation Policies Mobile with a Manipulation-Centric Whole-body Controller. Project page, 2024.
Luo et al. SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control. Science Robotics 11(117), 2026. arXiv:2511.07820.
Liu et al. Visual Whole-Body Control for Legged Loco-Manipulation. arXiv:2403.16967, 2024.
Li et al. AMO: Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control. arXiv:2505.03738, 2025.
Ben et al. HOMIE: Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit. arXiv:2502.13013, 2025.
Xue et al. LeVERB: Humanoid Whole-Body Control with Latent Vision-Language Instruction. arXiv:2506.13751, 2025.
Bronars, Park and Agrawal. Tune to Learn: How Controller Gains Shape Robot Policy Learning. arXiv:2604.02523, 2026.
Kwa et al. Measuring AI Ability to Complete Long Software Tasks. NeurIPS 2025; arXiv:2503.14499, first posted March 2025 as “Measuring AI Ability to Complete Long Tasks”. Summary on the METR blog, March 2025.
[ Stay Updated ]
Our research straight to your inbox.