ENCORE: Few-Shot Agentic Discovery
of Manipulation Strategies

Yifan Kang1,2  ·  Zihan Wang2  ·  Zhiwen Fan3  ·  Bangya Liu2,4

1MIT   22077AI   3Texas A&M University   4University of Wisconsin–Madison

WRL @ NeurIPS 2026  ·  Oral

MIT 2077AI Texas A&M University University of Wisconsin–Madison

Abstract

Coding agents can now write, run, and debug programs with little human help. Robot tasks, however, are usually specified by a sentence that leaves out how to grasp, in what order to make contact, and what the result should look like, and an agent given only the sentence must find these details by trial and error. We introduce ENCORE, which gives the agent a few demonstrations as evidence to read rather than as training data. A deterministic builder distills each demonstration into a pack of multi-view keyframes, gripper events, frame strips, and the full trajectory. A coding agent studies the pack, writes a policy program against a fixed perception and action API, refines it iteratively over a few development rollouts, and freezes it before a sealed evaluation that never reveals the success signal. On LIBERO-PRO, the agent's first program already succeeds in half of the perturbed tasks with demonstrations and in one task without them, and the frozen programs outperform the strongest prior agentic system run with the same language model (96.3% against 89.3%). On RoboDojo tasks whose instructions leave the goal unstated, no program succeeds without demonstrations. ENCORE also runs on a real bimanual robot, learning cube handover and cup inversion from five demonstrations each.

Method

ENCORE overview: demonstration pack, agentic coding loop, robot execution
The agent queries the pack, writes a program, refines it iteratively, and freezes it.

The first program carries the strategy

Development success by program version, K=3 vs K=0
First program: 41% development success with demonstrations, 1% without.

A strategy the agent could not find alone

The demonstrations open a different drawer, but show the pull that works: 50/50 with them, 0/50 without.

Composing demonstrated skills

Two demonstrated skills recombined into a new task
Two packs, neither showing the goal; the sentence picks which half of each to reuse.

LIBERO-PRO

LIBERO-PRO success per suite and axis
Pos moves objects; Task rewrites the sentence.

RoboDojo

RoboDojo successes out of 50 per task for K=3, images only, and K=0
When the sentence leaves the goal unstated, no program succeeds without demonstrations.

Real robot

Cup inversion: demonstration, agent reading, frozen program
Goal states of cube handover and cup inversion
Five demonstrations each. Cup inversion 9/10, cube handover 9/10.

BibTeX

@article{kang2026encore,
  title     = {ENCORE: Few-Shot Agentic Discovery of Manipulation Strategies},
  author    = {Kang, Yifan and Wang, Zihan and Fan, Zhiwen and Liu, Bangya},
  journal   = {arXiv preprint},
  year      = {2026}
}