Robotics harnesses & evaluation

Build the harness.
Measure what happens.

We connect AI agents to robots, then test what they can do. Real hardware. Physical feedback. Results we can inspect and repeat.

SO101 robot arm holding a pen over a drawing sheet
SO101 drawing harnessFrom code to physical feedback
5 peopleBuilding Green Robin Lab
Real robot prototypeCode and experiments published
Capability + safetyOur research direction

What we build

The model is one part.
The whole loop matters.

A robot needs tools, sensors, calibration, and a way to check its work. We build that harness and evaluate the system together. What changed? Did the task succeed? Can it do it again?

01 / HARNESS

Give agents a body.

Connect robot control, camera observations, calibration, and reusable tools so agents can act and inspect the result.

02 / CAPABILITY

Measure the outcome.

Our evaluation work asks how models and harness choices affect task completion, error recovery, and consistency on physical tasks.

03 / SAFETY

Test the boundaries.

Our next milestone is evaluating AI safety and harm in robotics, including how agents respond to hazardous requests.

Public prototype / Drawing

A robot drew it.
We kept the rough attempts too.

Our SO101 harness turns a robot arm into a pen plotter. An AI coding agent wrote the control code from scratch and used camera images to close the feedback loop.

The first ArtScience Museum drawing was rough. After changes to the pen mount, calibration, and stroke design, the robot produced a Petronas Twin Towers sketch.

88Strokes in the final sketch
6.5 minApproximate drawing time
3Camera views
Read the experiment ↗
Petronas Twin Towers sketch drawn by the SO101 arm on paper
The physical result. Code, calibration, and lessons are in the public repository.

Next research milestone

Capability brings
responsibility.

Planned evaluation work

We are developing robotics evaluations for AI safety and harm. Our focus includes hazardous material and weapon related scenarios, alongside evaluation of the harness and the robot's capabilities.

We want evidence about where an agent's boundaries hold and where they fail. These are research questions for the next phase.

  • Does the agent recognize a hazardous request?
  • Does it maintain safety boundaries when it can act?
  • How does changing the harness affect the outcome?

About the lab

Small team.
Open questions.
Let's build.

Green Robin Lab is a team of five working on robotics harnesses, capability evaluation, and AI safety. Dylan Ler shares code and research notes on agent tools, model behavior, and evaluation.

Get in touch

Working on the same questions?

We'd like to hear from people building robot systems, agent harnesses, and evaluations.

admin@greenrobin.us ↗