agentusecasesAll 965 use cases
Security, Science & Hardware

Control robots with generated code and voice commands

Gemini detects objects and writes robot API code, and with 10 in-context demos an ALOHA robot packed boxes and folded a dress; voice commands became robot plans.

Done withGemini

What they did
Gemini 2.5 was given a robot control API (detect object, get grasp pose, move/open/close gripper) and wrote Python plans zero-shot, such as putting a banana in a bowl. For ALOHA box packing and dress folding, 10 demonstrations of interleaved reasoning and actions were added to context. For voice, the same API calls were defined as Live API function-calling tools, with the camera feed streamed in.
How it went
Gemini produced two feasible banana plans, one of which re-planned around the right arm's reach limit. ALOHA packed boxes and folded a dress from the in-context demos. The post gives no success rates.
Worth knowing
The ALOHA setup is in an open-source Colab with example demonstrations. The Live API voice console is on GitHub, and Gemini Robotics-ER access requires a trusted tester waitlist.

Try it yourself with Gemini

Help me build a simulated robot arm controller in [simulator or language]. Use a vision model to detect objects in [image or camera feed], then write code that calls my robot API [API description] to [task, e.g. pack items into a box]. Run it in simulation only, and show me the generated code before anything is run on real hardware. Stop when the task succeeds in simulation three times in a row.

Read the original ↗

Source: developers.googleblog.com · Undated

Five of these in your inbox every morning

The best things people got an AI agent to do, each with the prompt to try it.

More like this