Analysis · 2 min read

Beyond Chatbots: Humanoid Robots Find Their Brains in Multimodal LLMs

Humanoid robots are undergoing a major brain upgrade. By integrating Multimodal Large Language Models (MLLMs), these machines are moving from pre-programmed automation to contextual reasoning, allowing them to adapt to real-world environments in real-time.

By Classy AI News · July 23, 2026

Beyond Chatbots: Humanoid Robots Find Their Brains in Multimodal LLMs

For years, humanoid robots were impressive pieces of hardware running on rigid, pre-programmed code. They could perform specific tasks in controlled environments, but they struggled with unexpected obstacles or unstructured instructions.

In 2026, that limitation is rapidly disappearing. The robotics industry has found its missing link: Multimodal Large Language Models (MLLMs). By using these advanced AI models as "brains," humanoid robots are transitioning from automated machinery to intelligent assistants capable of reasoning.

The Brain Upgrade: From Code to Context

Traditionally, if you wanted a robot to pick up a cup, you had to program the exact coordinates, the grip strength, and the path.

With MLLMs, the robot can process visual inputs and natural language instructions simultaneously. You can simply tell the robot, "Clean up the spill on the counter," and the AI will:

  • Identify the spill using its cameras.
  • Reason that it needs a cloth or paper towel.
  • Locate the cleaning materials.
  • Execute the physical task, adjusting its grip and movement dynamically if the cloth slips or the counter has an obstacle.

Real-World Deployment

This integration is no longer confined to research labs. Major players are deploying these "AI-brained" humanoids into industrial settings:

  • Figure & BMW: Figure’s humanoid robots, powered by custom OpenAI models, are being tested in manufacturing plants to perform complex assembly tasks alongside humans.
  • Tesla Optimus: Tesla is leveraging its self-driving AI stack to train its Optimus humanoid, focusing on warehouse logistics and repetitive factory work.
  • Boston Dynamics: The electric Atlas robot is using visual-language models to navigate complex construction and factory environments autonomously.

The Remaining Hurdles

Despite rapid progress, significant challenges remain before these robots enter our homes:

  • Physical Dexterity: While the AI "brain" can figure out what to do, the mechanical "hands" still struggle with delicate tasks like folding laundry or handling fragile objects.
  • Latency: The time it takes for a robot to process a visual frame, run it through a large model, and send signals to its motors needs to be near-instantaneous for safety.
  • Power Efficiency: Running large AI models locally on a robot requires massive battery power, currently limiting operation times to a few hours.

The Next Frontier

We are moving away from the era of screen-based AI. The next decade of artificial intelligence will be defined by embodied AI—intelligence that can move, touch, and manipulate the physical world.

Newsletter

Get the dispatch

One field. One email when we publish. Privacy.