RoboDojo Exposes the Embodied AI Gap: 12.8% vs. 100% in Real-World Manipulation
HKU's RoboDojo unifies 42 simulation and 18 real-world manipulation tasks—and reveals top AI models achieve just 12.8% real-world success against 100% for human experts.
The benchmark gap embodied AI needed
Embodied Artificial Intelligence aims to equip machines with the ability to perceive, reason, and act autonomously in physical spaces. Yet a major bottleneck has been the lack of standardized evaluation metrics—assessments confined to simulation or conducted under disparate hardware configurations and scoring criteria.
On August 12, 2026, The University of Hong Kong announced RoboDojo, a unified benchmarking platform designed to evaluate robotic manipulation across both simulated and physical environments.

Who built it
The Multimedia Laboratory (MMLab) at HKU spearheaded development, co-initiated by Professor Ping Luo, Associate Director (AI Research and Tech Transfer) of the HKU School of Computing and Data Science, and his PhD student Tianxing Chen. The project involved researchers from nearly 20 leading global universities, including UC Berkeley and Tsinghua University.
Professor Luo stated: "To the best of our knowledge, RoboDojo is the first Hong Kong-led benchmark to unify simulation and standardised real-robot evaluation. It moves embodied AI beyond impressive demonstrations towards progress that can be measured, compared and trusted."
What RoboDojo measures
RoboDojo integrates simulation-based evaluation, standardized physical robot testing, and policy benchmarking under a single framework:
- 42 simulation tasks and 18 real-world robotic tasks
- 30 representative robot policies evaluated through XPolicyLab
- Five evaluation dimensions in simulation: generalization, memory, precision, long-horizon execution, and open-vocabulary instruction following
- RoboDojo-RealEval — a reproducible real-world evaluation system with remote cloud access, standardized hardware, scene reset, evaluation protocol, and deployment interface
The platform supports scalable evaluation through heterogeneous parallel simulation in Isaac Sim. Policies integrate once in XPolicyLab and evaluate across simulation and real-world settings with minimal adaptation.

The numbers that matter
The benchmark's initial findings highlight a significant performance gap between current robotic systems and human capabilities:
- Top-performing AI model: 8.80% success in simulation, 12.8% in real-world testing
- Human experts: 76.03% in simulation, 100% in real-world testing
These results underscore the urgent need for more robust AI models capable of executing complex, multi-step tasks in dynamic physical environments.
Community adoption
Since launch, RoboDojo has garnered widespread attention. The project's release generated over 100,000 views on X within its first week, while open-source resources recorded more than 100,000 downloads on Hugging Face.
The public leaderboard and systematic analysis are available at robodojo-benchmark.com.
Why this matters beyond the leaderboard
RoboDojo addresses a credibility problem in embodied AI: impressive isolated demonstrations that fail to generalize. By unifying sim-and-real evaluation under reproducible protocols, it gives the field a shared ruler.
The gap between 12.8% and 100% real-world success is not a rounding error—it is the distance between demo-grade manipulation and deployable robotic reliability.

Looking ahead
RoboDojo is positioned to foster academic-industrial synergy globally, driving embodied AI from proof-of-concept showcases toward sustainable real-world applications. The arXiv preprint (2607.04434) provides full technical details for researchers integrating new policies into the benchmark.
Sources
- The University of Hong Kong — HKU Pioneers "RoboDojo": The First Hong Kong-Led Unified Benchmark for Embodied AI (August 12, 2026)
- arXiv — RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies (July 5, 2026)
- Tech Xplore — Scientists develop 'RoboDojo,' a unified platform to evaluate embodied AI (August 12, 2026)