We turn an imitation policy into a dense reward signal instead of a controller. On sparse-reward manipulation tasks, the resulting agent reaches 50% success in as little as half the interaction — and on real hardware, without the early safety stops that make robot reinforcement learning expensive.
- 20-57% fewer steps to 50% success on every task tested
- 0 safety resets in the first 1k real-robot steps (baseline 35-38 safety resets)
- 90% insertion into a rotating socket where demonstrations only saw stationary insertions
Showing a robot what to do is not the same as making it good at the job
Teleoperate a robot through a task fifty times and you can train a policy that reproduces what you did. That is a remarkable capability, and it is where most of modern manipulation starts. But it has a ceiling that is easy to miss in a benchmark and impossible to miss on a production line.
Imitation learning optimizes agreement with the demonstrator. It does not optimize task success, robustness to variation, or cycle time — the three things a deployed robot is actually judged on. If the demonstrator was slow, the policy is slow. If the demonstrator never encountered a particular disturbance, the policy has no opinion about it. Getting past that ceiling means collecting more and better data, forever.
Reinforcement learning is the principled alternative: let the robot optimize task success directly, discover behaviors the demonstrator never showed, and — because discounted returns reward finishing sooner — get faster as it improves. The catch is what happens in the first hour. Real manipulation rewards are sparse: the robot gets a single bit of feedback at the end of an episode and nothing at all before that. Until it stumbles into a success, it is performing an unstructured search with a heavy arm in a real workspace. That search is slow, it drops things, and it drives the arm into places it should not go — every one of which is a human walking over to reset the cell.
The standard fix is to use the demonstrations as an action prior: propose actions from the imitation policy, or mix demonstrations into the replay buffer. That helps, but it leaves the fundamental problem in place. The agent still has no feedback about whether the last ten steps made things better or worse. It only finds out at the end.
The idea: read the prediction as a direction, not a command
Modern imitation policies do not predict one action. They predict an action chunk — in our case, twenty actions covering roughly the next second of motion. Everyone treats that chunk as something to execute.
We treat it as something to measure against.
If you know how the robot's controller turns commands into motion, that chunk tells you where the hand/EE is going to be a moment from now, if the demonstrated behavior is followed. That is not a command. It is a local estimate of what progress looks like from right here — and unlike a sparse success bit, it is available at every single control step.
Task-space Imitation Guidance for Efficient Reinforcement learning (TIGER) is built on that reframing. The imitation policy is never executed. It is queried, converted into a short-horizon end-effector reference, and used to hand the reinforcement learner a dense reward for closing the gap. The sparse task reward stays in charge of what "done" means.

Three pieces that make it work on a real machine
Predicting where the hand/EE will actually go, not where the policy asked it to go. A policy's action space and the robot's realized motion are not the same thing. Between them sit controller scaling, frame conventions, command attenuation, latency and finite tracking bandwidth. Rolling out predicted actions naively gives a reference trajectory the robot could never follow, and rewarding against it teaches the wrong thing.
So TIGER maps each chunk through the controller before using it: the known analytic transforms first, then a small learned residual fitted from the same demonstrations, then integration into an end-effector trajectory anchored to the robot's current measured pose. The residual is deliberately low-order — it models systematic execution effects like latency and attenuation, not contact physics. It is fitted once per robot and controller in under a second, saturates after about five demonstrations, and costs at most 50 microseconds per chunk at run time. It also transfers: a mapping fitted on one task and evaluated on two hundred unseen demonstrations from another shifted its contact-phase prediction error by 0.05 mm, while the no-mapping baseline degraded from 32.8 mm to 42.0 mm.
Rewarding progress, not proximity. The dense reward is the change in pose alignment between consecutive steps, not the alignment itself. This matters more than it sounds. A reward for being well-aligned pays a robot to find a comfortable pose near the reference and stop. A reward for becoming better aligned pays only for movement in the right direction. We also scale the sparse success reward so that the entire dense return from a failed episode cannot exceed the reward for one success — the agent is never better off imitating well and failing.
A warm start that knows which directions are safe to try. Before any robot moves, TIGER pretrains the actor and critic offline on the demonstrations. Offline RL normally handles unfamiliar actions by being pessimistic about all of them uniformly, which produces a policy too conservative to improve. TIGER instead relaxes that pessimism selectively: for candidate actions the controller-aware mapping predicts will make task-space progress, the conservative penalty is softened; everywhere else it stays fully in force. The agent arrives online already achieving some success, and — critically — already biased toward the region of behavior the demonstrations covered.
Results
We evaluated TIGER on seven sparse-reward simulated tasks across MetaWorld, Robomimic and a humanoid shape-sorting insertion task, on all ten LIBERO-Spatial tasks, and on three contact-rich tasks on a real Franka Emika Research 3. Every simulation number below is averaged over five random seeds. Because guidance enters only through the reward and the pretraining objective, TIGER is not tied to one algorithm — we ran it on top of two different off-policy backends and report both in the paper.
It gets to competence sooner on every task we tried. Against the matched RL backend, TIGER reached 50% success earlier on all seven tasks, cutting the required interaction by 20% to 57%.
The gains are largest exactly where they should be. On tasks where the bottleneck is sparse exploration rather than fine control, the difference is not a speed-up but the difference between learning and not learning. On Robomimic ToolHang — a two-stage assembly task requiring up to 600 control steps — neither RL baseline ever succeeds, and the behavior-cloning policy tops out at 50%; TIGER reaches 79%. On the humanoid shape-sorting task, where tight insertion tolerances make standalone imitation brittle, behavior cloning reaches 10% and TIGER reaches 93%.

Two further comparisons are worth noting. Against learned reward models — ReWiND and RoboMeter, trained on the same demonstrations with the same RL backend — TIGER tied the best result on MetaWorld Assembly and came out ahead on Box Close, Stick Pull and Robomimic Square, without needing a learned visual or language reward function. And in the multi-task setting, one shared policy across all ten LIBERO-Spatial tasks with only ten demonstrations each, TIGER reached both the 50% and 90% thresholds earlier than the baseline and finished at 95% success.
On real hardware, the early-training difference is the operational one. We trained three tasks on a Franka FR3: picking a randomly placed block, opening one specific small-handled drawer in a stack, and inserting an object into a socket on a rotating disk. Block picking converged in roughly 8k environment steps against roughly 12k for the baselines; drawer opening in roughly 7k against 12k.

More telling than the step count is what happened to the workspace. We logged every safety violation — every time the arm left its task-specific safe region or triggered a controller fault, each of which terminates the episode and requires a reset. In the first thousand steps of block picking, the baselines each incurred on the order of 35 to 38 resets. TIGER incurred 0.
The rotary insertion task is where the reframing pays off most clearly, because it is the case where imitation alone cannot work. All fifty demonstrations were collected on a stationary disk — the demonstrator never once tracked a moving target. Asked to insert into a socket that rotates during the attempt, the behavior-cloning policy succeeds 10% of the time. TIGER, guided by that same mismatched prior, reaches 90%, and was the only method to complete the full curriculum on both seeds within the fixed 26k-step budget — about three hours of robot time per run.
That is the point of using demonstrations as guidance rather than as a policy. Imperfect guidance still tells you which way is forward. It does not cap what you can learn.
What this changes, and what it does not
For deployment, the interesting numbers here are not the success rates — they are the interaction budget and the reset count. A method that reaches competence in three hours instead of five, without a human intervening every few minutes, is a method you can actually run on a fleet. And because TIGER's guidance enters only through the reward and the offline objective, it drops on top of an existing behavior-cloning policy and an existing off-policy learner rather than replacing either.
The central assumption is also the central limitation: TIGER assumes short-horizon end-effector motion is an informative proxy for task progress. That holds well for reaching, grasping, insertion and articulated-object manipulation. It holds less well when success depends primarily on force regulation, tactile feedback, deformable-object dynamics, in-hand dexterity, or hidden object state — cases where the hand/EE can be in exactly the right place and the task can still be going wrong. Guidance is also only locally consistent: when demonstrations are multimodal, the imitation policy commits to one short-horizon reference per replanning step rather than enforcing global trajectory agreement, which is usually what you want, but means systematically misleading predictions can bias the shaped reward. And we still see the transient dip in performance during the offline-to-online handoff that is familiar from the offline RL literature. Improving that handoff is where we are looking next.
About Sanctuary AI
Sanctuary AI helps industrial leaders automate complex tasks by deploying Physical AI with production-ready performance across existing and future robotic hardware. The company is expanding what automation can achieve today while preparing customers for the next generation of industrial dexterous technologies. Sanctuary AI’s commanding IP portfolio, proprietary hydraulic hands, and advanced AI systems uniquely position the company as a leader in Physical AI.