# Sim-to-Real Transfer in Robotic Manipulation Agents: 7 Proven Strategies from the AI Agent Book

> Unlock sim-to-real transfer for robotic manipulation. Discover 7 AI agent book strategies including domain randomization, LLM rewards, and hardware fine-tuning. Improve your robot's real-world performance today.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-17

---

**Effective sim-to-real transfer requires high-fidelity physics simulation, domain randomization, curriculum learning, hybrid data mixing, LLM-as-Judge reward shaping, and targeted fine-tuning on physical hardware.**

The bojieli/ai-agent-book repository documents a production-grade pipeline for training robotic manipulation agents in simulation and deploying them on real robots. According to the source code and experimental chapters, closing the reality gap depends on specific engineering strategies implemented in the SAPIEN-based SimpleVLA-RL framework rather than algorithmic modifications alone.

## High-Fidelity Physics Simulation with SAPIEN

The foundation of successful transfer rests on accurate dynamics modeling. In [`chapter8/SimpleVLA-RL/robotwin2_tasks_description.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/SimpleVLA-RL/robotwin2_tasks_description.md), the RoboTwin 2 benchmark implements dual-arm manipulation tasks using the SAPIEN engine with precise contact dynamics, joint limits, and sensor noise models.

Realistic physics reduce the "reality gap" that causes policies to diverge when transferred to physical robots. The repository emphasizes accurate contact modeling between grippers and objects, ensuring that simulation captures the subtle mechanics of grasping and manipulation that dominate real-world performance.

## Domain Randomization for Visual and Physical Robustness

To prevent overfitting to simulation-specific artifacts, the codebase applies extensive domain randomization. As documented in [`chapter8/SimpleVLA-RL/vla-rollout-analysis.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/SimpleVLA-RL/vla-rollout-analysis.md), each rollout randomizes object positions, orientations, appearances, lighting conditions, and backgrounds.

This strategy exposes agents to a wide distribution of visual and physical conditions during training, encouraging robust perception and control that generalizes to unseen real-world variations. The following snippet demonstrates the randomization implementation:

```python
from verl.envs import RoboTwin2Env
import numpy as np

def make_randomized_env(seed: int = 0):
    env = RoboTwin2Env()
    env.seed(seed)

    # Randomize object pose and appearance

    for obj in env.scene.objects:
        obj.set_pose(
            position=np.random.uniform(-0.05, 0.05, size=3),
            orientation=np.random.uniform(-np.pi, np.pi, size=3),
        )
        obj.set_color(np.random.rand(3))          # random RGB

    
    # Randomize lighting

    env.scene.set_lighting(
        intensity=np.random.uniform(0.8, 1.2),
        direction=np.random.normal(size=3)
    )
    return env

```

## Curriculum Learning for Progressive Difficulty

The training pipeline employs curriculum learning to stabilize skill acquisition before confronting full environmental complexity. As described in [`book-en/chapter8.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter8.md), the system starts with easy, deterministic scenarios and progressively increases stochasticity by adding distractors and reducing simulation step sizes.

This staged approach allows policies to master basic manipulation skills early, preventing catastrophic failure when later exposed to the full noise and uncertainty of real hardware. The curriculum transitions through three distinct stages:

```python
def curriculum_step(env, stage):
    if stage == 0:          # Easy: fixed object pose, no noise

        env.set_randomization(enabled=False)
    elif stage == 1:        # Medium: pose randomization only

        env.enable_pose_randomization(True)
        env.enable_visual_randomization(False)
    else:                   # Hard: full randomization

        env.enable_pose_randomization(True)
        env.enable_visual_randomization(True)

```

## Hybrid Simulation and Real-World Data

To anchor policies to genuine dynamics while maintaining sample efficiency, the repository mixes simulated rollouts with tele-operated real robot trajectories. According to [`book-en/chapter6.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter6.md), the XLeRobot dataset provides "real-world ceiling" trajectories recorded from physical hardware.

This hybrid approach leverages large-scale simulation data for sample-efficient reinforcement learning while using scarce real-world data to prevent divergence from actual robot dynamics. The training loop alternates between simulation batches and occasional real-world collection:

```python
sim_env = make_randomized_env()
real_env = RealXLeRobotEnv()          # defined in chapter6/XLeRobot

for epoch in range(num_epochs):
    # Simulated rollouts

    for _ in range(sim_batch):
        traj = run_episode(sim_env, policy)

    # Occasionally collect real rollouts

    if epoch % real_interval == 0:
        real_traj = run_episode(real_env, policy, record=True)
        replay_buffer.add(real_traj)

    # Update policy using combined buffer

    policy.update(replay_buffer.sample(batch_size))

```

## Cross-Domain Reward Shaping with LLM-as-Judge

Consistent reward signals across simulation and reality prove critical for stable transfer. The evaluation system in [`book-en/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter7.md) implements an LLM-as-Judge model that provides dense feedback during simulation training and refined scores during real-world evaluation.

This learned reward model prevents reward hacking on simulator-specific quirks while providing the dense supervision necessary for sample-efficient learning. The implementation uses a lightweight language model to evaluate trajectory success:

```python
from llm_judge import LLMJudge

judge = LLMJudge(model_name="gpt-4o-mini")
def compute_reward(state, action, next_state):
    prompt = f"State: {state}\nAction: {action}\nNext state: {next_state}\nIs the action successful? Yes/No."
    verdict = judge.ask(prompt)
    return 1.0 if "Yes" in verdict else 0.0

```

## Fine-Tuning on Real Robot Hardware

Despite robust simulation training, residual gaps in sensor noise, latency, and actuation characteristics require final adaptation. The SFTvsRL experiments in [`chapter8/SFTvsRL/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/SFTvsRL/README.md) demonstrate that a short fine-tuning phase on real XLeRobot episodes closes these remaining discrepancies.

This phase typically involves either supervised fine-tuning (SFT) on successful real trajectories or reinforcement learning with real-world feedback, adjusting network weights to match the specific characteristics of the target hardware platform.

## Simulation-Aware Architecture Design

The repository advocates for agents that treat the simulator as a first-class tool during training. As detailed in [`book-en/chapter8.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter8.md), this design principle exposes simulator-specific APIs including `reset`, `step`, and `seed`, and allows policies to query perfect simulation state for planning during training.

This architecture enables policies to exploit simulation affordances—such as ground-truth object poses—during learning while learning to operate with only partial observations when deployed on physical hardware.

## Summary

Effective sim-to-real transfer in robotic manipulation requires a systematic pipeline rather than isolated tricks:

- **High-fidelity simulation** using SAPIEN with accurate contact dynamics forms the necessary foundation
- **Domain randomization** of visual and physical parameters prevents overfitting to simulation artifacts
- **Curriculum learning** gradually increases difficulty from deterministic to fully stochastic scenarios
- **Hybrid data pipelines** mix abundant simulation data with scarce but crucial real-world trajectories
- **LLM-as-Judge** provides consistent reward signals across both simulated and real environments
- **Real-world fine-tuning** closes residual gaps in sensor characteristics and actuation dynamics
- **Simulation-aware architectures** exploit perfect state information during training while learning robust inference

## Frequently Asked Questions

### What is domain randomization in sim-to-real transfer?

Domain randomization is a training technique where simulation parameters—including object colors, lighting intensity, camera poses, and friction coefficients—are randomized across episodes. According to the AI Agent Book source code in [`chapter8/SimpleVLA-RL/vla-rollout-analysis.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/SimpleVLA-RL/vla-rollout-analysis.md), this exposes policies to a broad distribution of environmental conditions, forcing them to learn robust features that generalize to the specific (but unknown) parameters of the real world.

### Why is the RoboTwin 2 benchmark important for manipulation research?

RoboTwin 2, described in [`chapter8/SimpleVLA-RL/robotwin2_tasks_description.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter8/SimpleVLA-RL/robotwin2_tasks_description.md), provides standardized dual-arm manipulation tasks built on the SAPIEN physics engine. It serves as the primary testbed for validating sim-to-real strategies because it accurately models contact dynamics, joint limits, and sensor noise that dominate real-world manipulation performance, allowing researchers to verify that simulation success correlates with hardware deployment success.

### How does LLM-as-Judge prevent reward hacking during sim-to-real transfer?

Traditional hand-crafted rewards often exploit simulator-specific physics quirks that fail to transfer. The LLM-as-Judge system documented in [`book-en/chapter7.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter7.md) evaluates trajectories using language models trained on human preferences, providing reward signals based on task semantics rather than simulator physics. This approach yields consistent evaluation criteria across both simulated and real environments, preventing policies from optimizing for simulation artifacts that do not exist on physical hardware.

### What is the difference between curriculum learning and fine-tuning in this pipeline?

Curriculum learning occurs during the initial simulation training phase, gradually increasing environmental difficulty by adding randomization and noise, as implemented in the training loop referenced in [`book-en/chapter8.md`](https://github.com/bojieli/ai-agent-book/blob/main/book-en/chapter8.md). Fine-tuning occurs after simulation training is complete, using a small dataset of real robot trajectories to adapt the policy to specific hardware characteristics. The former stabilizes learning in simulation, while the latter bridges the final gap to physical deployment.