Contact Before Motion: RynnBrain 1.1 Tests Whether Embodied Brains Can Scale Like LLMs
Alibaba DAMO Academy's RynnBrain 1.1 adds contact-point prediction, native 3D grounding, and a unified cross-embodiment VLA — then deploys it on Unitree G1, Astribot-S1, and Tianji-Wuji with benchmark-leading cognition scores and up to 95% real-robot success.
For years, the robotics industry's quiet complaint has been that vision-language models understand scenes without understanding where to touch them. A policy can narrate a kitchen; it cannot reliably decide which point on a mug handle to close a gripper around, or how that gripper should rotate in the plane before the first millimeter of motion.
On July 16, Alibaba's DAMO Academy released RynnBrain 1.1 — a family of embodied foundation models at 2B, 9B, and 122B-A10B parameter scales — with an explicit bet that the gap is trainable. The technical report, posted to arXiv on July 20, adds two pretraining tasks aimed squarely at manipulation: contact-point prediction (center point plus in-plane grasp orientation) and native 3D grounding (metric 3D bounding boxes in camera coordinates). The team then built RynnBrain-VLA, a vision-language-action policy with a unified cross-embodiment action space, and ran it on three commercially available platforms: Unitree's G1 humanoid, Astribot's S1 bimanual arm, and the Tianji-Wuji dexterous-hand stack.
The headline numbers are not simulation curves. On three long-horizon real-robot tasks — Sprinkle the Petals, Pour the Wine, and Grab the Spatulas — RynnBrain-VLA averaged 86.67% final success across 20 trials per task, versus 60.00% for an otherwise identical VLA initialized from Qwen3.5. A jointly trained multi-task, multi-embodiment variant pushed that average to 91.67%. On the dexterous Grab the Spatulas benchmark, success hit 95%, more than doubling the Qwen baseline's 50%.
From scene description to contact geometry
RynnBrain 1.0, released earlier in 2026, demonstrated that a single embodied multimodal model could handle egocentric understanding, spatial grounding, and planning within one autoregressive framework. RynnBrain 1.1 keeps the Qwen3.5-based decoder-only architecture — vision encoder, projector, LLM backbone — but asks a sharper question: can pretraining outputs be made directly useful to downstream robot policies?
Contact-point prediction is the team's answer for grasping. Rather than predicting axis-aligned grasp rectangles, the model learns to emit a contact center and planar rotation angle — features the authors argue are more precise and more meaningful for policy heads than box corners. On qualitative cases shown in the paper, RynnBrain 1.1-9B selects mug handles over mug centers, adapts when mugs are overturned, and disambiguates referential targets in clutter.
Native 3D grounding extends the same autoregressive token interface into metric space. Given language and camera intrinsics, the 2B and 9B variants predict a nine-dimensional 3D box — center position, dimensions, and orientation in meters and radians — discretized into the shared vocabulary alongside text and 2D points. On SUN RGB-D, RynnBrain 1.1-2B reaches 34.28 AP@15, edging Seed1.5-VL (33.5) and Gemini 2.0 Pro (32.5); scaling to 9B lifts that to 41.12 AP@15. On WildDet3D-Bench, the 9B model scores 23.44 AP3D, surpassing the specialized WildDet3D detector (22.6) that trained with additional in-domain data.
These are not vanity leaderboard entries. They are proxies for whether a single model can bridge "what" and "where in the room" before any motor command fires.
Scaling laws that refuse to behave like LLMs
The paper's scaling analysis is where RynnBrain 1.1 becomes editorially interesting. Comparing RynnBrain 1.1 against raw Qwen3.5 at matched 2B, 9B, and 122B-A10B scales, the authors partition benchmarks into three groups: general embodied cognition, reasoning-intensive cognition, and embodied localization.
General cognition improves monotonically for both model families as parameters grow — familiar territory for multimodal LLM scaling. Reasoning-intensive tasks diverge: RynnBrain 1.1 gains +38.6% from 2B to 122B-A10B, while Qwen3.5 regresses −39.2% on the same bucket. The performance gap widens from 18.2 points to 50.8. The interpretation offered in the report is blunt: multi-view and temporal embodied reasoning is not an emergent property of bigger language priors. Without explicit spatiotemporal supervision, larger Qwen models may override weak visual-spatial signals with more confident but less accurate language.
At the top of the cognition stack, RynnBrain 1.1-122B-A10B leads all evaluated proprietary and open models on VSI-Bench (75.0), MMSI (52.0), and RefSpatial-Bench (79.1) — including Gemini 3 Pro, GPT 5.4, Claude Sonnet 4.6, and Gemini Robotics-ER 1.5 on those shared benchmarks. RefSpatial-Bench scaling is especially clean: 58.5 at 2B, 67.2 at 9B, and 79.1 at 122B-A10B.
One action space, three bodies
Hardware heterogeneity has historically forced labs to train separate policies per morphology — different action dimensionalities, different controllers, little transferable structure. RynnBrain-VLA attacks that fragmentation with a 81-dimensional unified action space organized into semantically aligned body-part groups (arm joints, grippers, hands, head, torso). Embodiment-specific masks activate only the dimensions a given robot actually possesses, so Astribot S1 data and Tianji-Wuji hand data can train a single policy without forcing incompatible action vectors into alignment.
Deployment details matter for readers tracking the gap between paper and warehouse:
- Unitree G1: RynnBrain-VLA predicts 64-dimensional SONIC whole-body motion tokens plus 14-dimensional dexterous-hand commands, decoded through the SONIC controller. On Pull the Chair, it reached 90% success versus 75% for GR00T N1.7 with the same SONIC stack.
- Astribot S1: Gripper-based bimanual tasks with active viewpoint changes — the robot must re-observe while executing.
- Tianji-Wuji: 54 active dimensions across dual arms and Wuji dexterous hands.
All inference runs locally on a workstation with a single NVIDIA GeForce RTX 4090, directly cabled to each robot for real-time execution — no cloud round-trip.
Baselines on the three quantitative long-horizon tasks include GR00T N1.7 (73.33% average success), pi0.5 (65.00%), and the Qwen-initialized VLA (60.00%). RynnBrain-VLA's 91.28% average process score — measuring partial completion of multi-stage tasks, not just final success — suggests fewer catastrophic mid-task failures, not just lucky final frames.
Joint multi-task and multi-embodiment training (RynnBrain-VLAGeneralist) further improved Pour the Wine from 80% to 95% final success without obvious cross-embodiment interference. The authors hypothesize that affordances, task progress, and reach-grasp-manipulate structure are shared across morphologies even when joint commands differ — a claim the robotics field has whispered for years but rarely demonstrated with open weights at three scales.
What ships today
Checkpoints for all three base models — 2B, 9B, and 122B-A10B — are on Hugging Face and ModelScope under the Alibaba-DAMO-Academy organization, with cookbooks and a GitHub repository (alibaba-damo-academy/RynnBrain). That openness matters: embodied foundation models have lagged text LLMs on reproducibility, and a unified training recipe across dense and sparse-MoE scales gives external labs a controlled knob for studying data versus parameters.
The work also arrives amid a crowded July for physical AI. Hugging Face shipped LeRobot v0.6.0 with world-model policies and six new simulation benchmarks; Black Forest Labs and mimic robotics deployed FLUX-mimic video-action models on Audi production lines; Geek+ unveiled its Gravity 4D embodied framework at WAIC 2026. RynnBrain 1.1's differentiator is not a single demo clip — it is the combination of manipulation-aligned pretraining outputs, cross-embodiment VLA training, and published real-robot success rates against matched Qwen and generalist VLA baselines.
The honest ceiling
None of this resolves the industry's throughput problem. Community benchmarks like PhAIL still document warehouse picking rates where the best VLA models land near 64 units per hour against 330 for human teleoperation on the same Franka arm — a gap RynnBrain 1.1 does not claim to close. Long-horizon success in controlled lab tasks is not the coffee test.
What the release does establish is a falsifiable claim: embodied pretraining with explicit contact and 3D supervision produces VLA policies that beat general-purpose VLMs trained on identical downstream data — and that joint training across embodiments can help rather than hurt. For labs deciding whether to fine-tune Qwen or an embodied brain this quarter, that is actionable evidence rather than marketing adjective soup.
The contact point, in other words, is no longer an afterthought bolted onto a captioning model. It is part of the pretraining vocabulary — and the robots on Alibaba's test floor are already speaking it.
Sources
- Alibaba DAMO Academy — RynnBrain 1.1 Project Page (July 16, 2026)
- arXiv — RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model (July 20, 2026)
- GitHub — alibaba-damo-academy/RynnBrain (July 16, 2026)
- Hugging Face — RynnBrain 1.1 Model Collection (July 16, 2026)
- Hugging Face — LeRobot v0.6.0: Imagine, Evaluate, Improve (July 7, 2026)
- Black Forest Labs — FLUX 3 x mimic: The Next Generation of Video-Action Models (2026)
- Hacker News — Show HN: PhAIL – Real-robot benchmark for AI models (2026)