MobiAgent

MobiAgent: Dual-Loop Recursive Policy Self-Improvement
for Long-Horizon Mobile Manipulation

Chenzhi Liu1,*Zhang Yue1,*Jiehong Lin1,*,†Jianan Wang2Bo Wang1Zhongrui Wang3,‡Xiaojuan Qi1,‡

1 The University of Hong Kong2 Astribot3 Southern University of Science and Technology

* Equal contribution   † Project leader   ‡ Corresponding authors

Pour blue particlesPaper disposalBottle disposalCan disposal

It plans the task, executes it skill by skill,
and improves from its own deployment data.

ABSTRACT

Recursive policy self-improvement.

Long-horizon mobile manipulation presents significant challenges due to compounding execution errors and capacity interference between locomotion and arm control. While recent Vision-Language-Action models excel at short-horizon tasks, they lack the hierarchical reasoning required for multi-stage objectives. Furthermore, existing hierarchical agents suffer from rigid sub-task mapping, inflexible replanning, and a lack of continuous learning. To address these limitations, we introduce MobiAgent, a dual-loop agentic framework that bridges robust deployment execution and recursive policy self-improvement. During deployment, the Inner Loop decouples high-level reasoning from low-level control through highly composable atomic skills. It employs Vision-Language models for receding-horizon planning and visual reflection, dynamically composing skills to ensure robust error recovery. These skills are executed by specialized flow-matching experts that share a unified VLM backbone, maximizing reusability while mitigating capacity interference. Concurrently, the Outer Loop drives automated lifelong learning by autonomously segmenting and verifying deployment rollouts, clustering them to discover atomic skills, and continuously fine-tuning the skill library without human annotations. Evaluations on RoboCasa, BEHAVIOR-1K, and real-world tasks demonstrate the effectiveness of MobiAgent. It outperforms π0.5-TA by 22.5 percentage points on BEHAVIOR-1K and enables robust recovery from execution failures. Through autonomous data recycling, success improves from 7.50% to 27.50% on RoboCasa and from 32.5% to 57.5% on Astribot S1.

THE CHALLENGE

One task.
A long chain of skills.

A mobile manipulator must coordinate navigation and manipulation across an entire environment. Every subtask depends on the last succeeding; one dropped object can derail everything that follows.

MobiAgent closes the loop. The Inner Loop composes atomic skills and checks their outcomes during deployment. The Outer Loop builds a skill library from startup data, then uses verified deployment clips to refine the existing policies.

Dynamic planning and error recovery, from simulated household tasks to real-robot deployment.

SEE IT IN ACTION

From a plan to the real world.

TASK CLIPS & RECOVERY LOGIC
REAL ROBOT · 3×

Pour blue particles

Navigate, grasp, pour, and return the bottle.

REAL ROBOT · 3×

Bottle disposal

Pick up the bottle, carry it, and place it in the bin.

FRAMEWORK WALKTHROUGH

Reflect, retry, replan

How the critic chooses to advance, retry, or replan.

Watch the full story

NARRATED OVERVIEW · 2:59

Choose a chapter Original 4K video

THE FRAMEWORK

Two loops. One shared skill library.

Deployment-time planning and recovery feed an offline loop for recursive policy self-improvement.

INNER LOOP · DEPLOYMENT

Plan. Execute. Reflect.

Task Planner

Compose the next subtask using observations, episode memory, and the atomic skill catalog.

Skill Executor

Generate action chunks with a shared VLM trunk and skill-specific flow-matching experts.

Reflection Critic

Verify the visual outcome and return a verdict, follow-up action, and observable cues.

Visual feedback returns to the planner.

OUTER LOOP · TRAINING

Experience becomes training data.

Data Curator

Segment startup and deployment trajectories, then verify the captioned skill clips.

Skill Generator

Cluster startup instructions into canonical skills, then map new rollout clips to the existing inventory.

Skill Trainer

Bootstrap skill experts, then refine the existing policies using curated deployment data.

Updated skills return to the deployment loop.

A shared VLM backbone supports reusable, skill-specific action experts.

INSIDE THE SKILL EXECUTOR

Shared perception. Specialized actions.

The instruction selects an action expert. Visual and language understanding stay shared across the skill library.

INPUTObservations
+ instruction
What the robot sees and what to do next
SHARED VLMOne backboneCommon visual-language features
ACTIVE EXPERTPick upSkill-specific flow-matching control
OUTPUTAction chunkCoordinated robot motion

Select an atomic skill

EXAMPLE INSTRUCTION

“Pick up the soda can.”

The picking expert handles grasping; the same expert can be reused for different objects and long-horizon tasks.

Choose a BEHAVIOR-style atomic skill to explore its expert and example instruction.

EXPERIMENTS

Better execution.
Learning from deployment.

BEHAVIOR-1K

65.0%

Long-horizon success

Mean over four BEHAVIOR-1K tasks.

VS. TASK-SPECIFIC BASELINE

+22.5pp

Stronger execution

65.0% compared with 42.5% for π0.5-TA.

REAL-ROBOT EVOLUTION

32.557.5%

Learn from deployment

+25.0 percentage points over two training rounds.

ROBOCASA365 · COMPOSITE-SEEN

Learning across six deployment rounds.

16 household tasks, with 2–15 subtasks each. A Franka arm on an Omron mobile base is evaluated with five trials per target scene; success is averaged across tasks and scenes.

7.50%27.50%

Base π0.5 → best MobiAgent round

+20.00 percentage points at R3 and R5.

Early gains are followed by a plateau: R4–R6 range from 25.00% to 27.50%, ending at 26.25% in R6.

About the RoboCasa365 benchmark ↗
MobiAgent roundsCaP-X · 5.00%
RoboCasa success rate by deployment-to-training roundBase π0.5 at R0: 7.50 percent. MobiAgent: R1 18.75, R2 21.25, R3 27.50, R4 25.00, R5 27.50, R6 26.25 percent. CaP-X baseline: 5.00 percent. Peak performance occurs at R3 and R5, with a later plateau.Success rate (%)01020307.50R018.75R121.25R227.50R325.00R427.50R526.25R6
R0: base π0.5 policy. R1–R6: successive MobiAgent deployment-to-training updates.
View all RoboCasa success rates
RoboCasa365 Composite-Seen · success rate (%)
CaP-Xπ0.5 (R0)R1R2R3R4R5R6
5.007.5018.7521.2527.5025.0027.5026.25

Long-horizon performance

Mean success across four BEHAVIOR-1K tasks, with 10 trials per task under matched evaluation conditions.

π0.5-TA
42.5%
MobiAgent
65.0%

+22.5 percentage points over the task-specific baseline.

Autonomous policy evolution

On Astribot S1, 100 demonstrations per task initialize the skill policies. Each update adds approximately 20 automatically curated deployment trajectories per task, without manual annotation.

Iteration 2
57.5%

mean success rate
+25.0 points from bootstrap

32.540.057.5 StartIter. 1Iter. 2
Real-robot taskSuccess rate
Paper disposal50%
Bottle disposal60%
Can disposal70%
Pour blue50%

View success rates by task +
Simulation · success rate (%)
MethodPush radioDispose trashStack storageFetch beerMean
π0.5-TA6040601042.5
MobiAgent7060706065.0

COMPONENT ABLATIONS · BEHAVIOR-1K

What makes the system work?

Task success rate %

Fixed task-specific sub-tasks20.0%
Single shared action head50.0%
Without Reflection Critic10.0%
Without global replanning40.0%
MobiAgent · full system65.0%

Global replanning adds 25.0 percentage points to mean success: 40.0% → 65.0%.

View the complete ablation table
BEHAVIOR-1K · task success rate (%)
MethodPush radioDispose trashStack storageFetch beerMean
Fixed task-specific sub-tasks (w/o Planner)50003020.0
Single shared action head7040504050.0
Without Reflection Critic30100010.0
Without global replanning7040302040.0
MobiAgent (Full)7060706065.0

CITE THIS WORK

BibTeX

arXiv: 2610.03476 ↗

@misc{liu2026mobiagent,
  title = {MobiAgent: Dual-Loop Recursive Policy Self-Improvement for Long-Horizon Mobile Manipulation},
  author = {Chenzhi Liu and Yue Zhang and Jiehong Lin and Jianan Wang and Bo Wang and Zhongrui Wang and Xiaojuan Qi},
  year = {2026},
  archivePrefix = {arXiv},
  eprint = {2610.03476},
  primaryClass = {cs.RO},
  url = {https://arxiv.org/abs/2610.03476}
}