Embodied Intelligence for Long-Horizon Robot Manipulation: A Survey of Challenges, Mechanisms, and Evaluation.
-
Abstract
Long-horizon robot manipulation is a critical bottleneck for embodied intelligence in transitioning from short-horizon skill execution to real-world deployment. Its essence is not merely an extension of action sequences, but rather an embodied decision-making problem jointly constrained by stage dependency, historical state, physical state evolution, and recovery cost. Recognizing that existing surveys are predominantly organized by model type or technical lineage and thus struggle to reveal the relationship between long-horizon failure modes and methodological mechanisms, this paper characterizes long-horizon robot manipulation along three dimensions—stage structure, history dependency, and error propagation—thereby delineating its structural boundary from short-horizon manipulation. A unified analytical framework of structural difficulties and compensation mechanisms is established, mapping core challenges such as combinatorial explosion, state drift, and error cascading onto four functional mechanisms: task organization, future verification, state and semantic support, and execution recovery. The paper further analyzes the conditions for forming and validating long-horizon capability from three perspectives: training supervision, evaluation benchmarks, and real-world deployment boundaries. Long-horizon capability is not a linear extension of single-step control accuracy, but rather a system-level capacity to sustain cross-stage task consistency under partial observability and error accumulation.
-
-