This repository is the public implementation of GPT-Policy, a closed-loop control framework that connects a fixed vision-language model (VLM) to robot tools. At deployment time, the agent can use demonstrations, goal images, interaction history, and execution feedback without gradient updates or task-specific parameter changes.
The implementation currently provides hardware adapters for two robot arm platforms: ARX X5 and I2RT/YAM.
Paper: In-Context Robot Learning with VLM Agents
Authors: Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, and Tong Wu.
Abstract
Enabling robots to adapt to unfamiliar environments as readily as humans remains a moonshot goal of embodied AI. No finite collection of demonstrations can cover every task and situation a robot will encounter, making the ability to learn from context at deployment essential for generalization. Such in-context learning (ICL), however, remains largely beyond the reach of existing robotic policies. The broad agentic capabilities of commercial vision-language models (VLMs), such as GPT-6 Astra, raise a compelling question: can these models learn from demonstrations, examples, and interaction feedback, then translate that information into executable and verifiable robot behavior from a new initial state without gradient updates or persistent changes to task-specific parameters? We introduce GPT-Policy, a general-agent framework for in-context robot learning. GPT-Policy integrates a context compiler that preserves task-relevant visual transitions, a VLM that proposes robot-tool actions, and a constrained controller that verifies and executes each action and reports its outcome. We evaluate its reliability and limitations through task success and efficiency metrics, matched comparisons across models, and controlled context ablations. In real-robot trials, human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. These findings position GPT-Policy as a step toward robot adaptation through in-context learning, providing an empirical foundation for translating the general-purpose capabilities of VLMs into physical behavior and clarifying the challenges that must be overcome for reliable deployment.
- 🚀 [2026/09/16] The paper, project page, and code are now publicly available!
- 🎥 [2026/09/11] The first robot cases and demonstrations are added to the project page!
Four synchronized views of three GPT-6 Astra robot runs. Click any preview to open its MP4.
Plug insertion · top view Plug insertion · right wrist view
Sprite retrieval · search and place the bottle Bottle opening · unscrew and separate the cap
The full experiment gallery includes the other tasks and context comparisons.
GPT-Policy builds one model input from the task, live camera/state observations, task references, and the previous tool result. The VLM emits one structured request; the selected adapter validates and executes it, then returns fresh observations and feedback for the next decision. Adapters support Cartesian targets and waypoint sequences, sequential IK checks, backend-specific timing, gripper control, and append-only run recording.
The available context types are:
- Human Video: a visual procedure that can transfer across embodiments.
- Robot Video / Video + Action: robot interactions, arm roles, and aligned motion references.
- Target Image: the desired object arrangement, position, and spacing.
- Self History: earlier observations, actions, results, and discovered subgoals.
- Human-Robot Interaction: live intent, pointing, corrections, and turn-taking.
Python 3.10 or newer is required.
python -m venv .venv
source .venv/bin/activate
python -m pip install -e .The base install is hardware-free. Install only the backend you need:
source .venv/bin/activate
python scripts/install_drivers.py yam
python scripts/install_drivers.py realsense
# ARX instead: python scripts/install_drivers.py arx
# YAM alternative: python -m pip install -e '.[yam,realsense]'Agent CLIs are external dependencies. Install and authenticate the provider you select; credentials are stored outside this repository.
The default is YAM (yambox) + Codex (gpt-6-astra). Run from the project directory after installation. For different hardware, configure a machine profile first.
source .venv/bin/activate
gpt-policy --checkThe default reports machine: yambox, backend: yam, and hardware_opened: false. Local or explicit profile overrides take precedence.
Replace the path with your reviewed demo.json. This command starts the robot task:
gpt-policy "Use the robot demonstration as a reference. Unscrew and remove the bottle cap, leaving the bottle standing securely on the table." \
--demo /path/to/robot-demonstration/demo.jsonThe default auto mode includes recorded states and actions when available. No extra --demo-mode video+action is needed.
Use the filename printed by Request:, for example:
gpt-policy --input-json request_json/unscrew-remove-bottle-cap.jsonTask requests live in request_json/; generated context, images and recordings live in var/.
| Need | Documentation |
|---|---|
Choose auto, video, or video+action |
Demonstration modes |
| Preview context and see the loaded mode | Context inspection |
| Change machines, cameras, or live-window settings | Runtime configuration |
🎯 Target Image / Self History / Human–Robot Interaction: 100% success on each of the six tasks evaluated under these conditions.
👀 Human Video: Success rate improves from 0% to 67% on both towel and notebook pickup.
🤖 Robot Video + Action: Success rate improves from 0% to 100% on bottle opening and from 0% to 67% on plug reinsertion.
src/gpt_policy/ protocol, input preparation, planning, recording, adapters
configs/default.json default YAM/yambox profile for `gpt-policy "..."`
configs/agents/ provider examples
configs/examples/ local-machine templates
configs/machines/ yambox, arx247 and arx248 deployment profiles
configs/calibration/ measured intrinsics and camera transforms
configs/robots/ portable robot and motion defaults
requirements/ installation entry points for each backend
scripts/ driver installation, evaluation and provider checks
tests/ core offline configuration/runtime regression tests
docs/assets/ figures used in this README
The public tree includes the deployed YAM/ARX machine profiles and their measured calibration. It excludes private task prompts, real credentials, run recordings and evaluation history. The paper's physical demonstration records and complete evaluation environment are not included by implication.
The public repository keeps a focused offline regression suite for configuration, live-window behavior, camera controls and recording layout.
python -m pytest -q
python -m compileall -q src- Release the GPT-Policy pipeline for real-world ARX robots, including the complete harness and format adapters for different context types.
- Release the YAM pipeline with hardware integration and flexible context support.
- Release the RoboDojo simulation pipeline for reproducible evaluation.
- Optimize the agent harness for context construction, feedback, and execution efficiency.
This is a research control loop. The integrator must verify calibration, workspace limits, collision behavior, camera placement, provider configuration, and emergency-stop procedures before energizing a robot. IK acceptance and a model completion message do not establish collision-free motion or physical task success. The project license is intentionally pending; redistribution and commercial use are not granted by this preview.
See THIRD_PARTY.md for third-party notices and optional SDK sources.
@article{cheng2026incontextrobotlearningvlm,
title={In-Context Robot Learning with VLM Agents},
author={Dongzhou Cheng and Taoran Yi and Ye Fang and Xingwu Zhang and Fan Feng and Yixuan Li and Gengxiong Zhuang and Rongze Wang and Shuai Yang and Wei Song and Weizhi Xue and Minyan Wu and Jie Gui and Jiaqi Wang and Tong Wu},
journal={arxiv:2609.19138},
year={2026}
}GPT-Policy integrates optional ARX, I2RT/YAM, RealSense, and provider CLI interfaces, with reference to RoboCurve's inspect-robots project. Please see THIRD_PARTY.md before redistributing a deployment that includes external SDKs.


