Expert Data Expansion#
EmbodiChain supports two expert-authoring paradigms for synthetic demonstration data:
Handwritten trajectories implement the expert in a registered Python task, using
MotionGeneratordirectly or composing reusable Atomic Skills.Task Program declares the task flow in typed YAML and lets the configured environment integration lower Semantic Calls to the same Atomic Skills.
The two paradigms differ only in how the expert plan is authored. Both produce
lazy DemoSegment objects, and both use the
same Gym executor, env.step() path, validation rules, dataset manager, and
transactional commit boundary.
Both paradigms retain an overall task instruction and a finer instruction for each semantic segment. See Task and segment language for expert data for task/program configuration, per-episode overrides, LeRobot readback, online training, and migration from recorder-owned language.
Paradigm 1: Handwritten Trajectories#
A handwritten expert lives beside its registered task entry point, for example:
embodichain_tasks/embodichain_tasks/manipulation/tableware/stack_blocks_two.py;embodichain_tasks/embodichain_tasks/manipulation/tableware/blocks_ranking_rgb.py.
Both examples build a MotionGenerator backed by TOPPRA and pass it to an
AtomicActionEngine. The engine then plans reusable PickUp and Place
invocations while the task retains control over scene queries, command timing,
segment metadata, and success checks.
Direct Motion Generator or Atomic Skills#
These are two abstraction levels, not two unrelated planning engines:
Use
MotionGenerator.generate()directly when the task already owns the joint or Cartesian waypoints and only needs a planned joint trajectory. WrapPlanResult.positionsas aDemoSegmentaction iterable and combinePlanResult.successwith a physical validator. See Motion Generator for the complete planning API.Use
AtomicActionEnginewhen the behavior matches reusable skills such as PickUp, Place, MoveEndEffector, Pour, or HandOver. The engine owns the shared motion generator, command profiles, typed goals, and projected semantic context. See Atomic actions for the action contracts.
When using MotionGenerator directly, its planned positions are ordered for
the selected control_part. The actions yielded to the environment must
match its active action space. If the environment controls more joints than
the planned arm, merge the arm plan with hold or gripper commands before
yielding it; do not silently rely on incompatible dimensions.
Build the planning services#
The two-block stacking task demonstrates the recommended Atomic Skills setup:
Motion Generator and AtomicActionEngine setup
1 def _initialize_atomic_actions(self) -> None:
2 """Create the right-arm atomic-action engine and object semantics."""
3 from embodichain.lab.sim.atomic_actions import (
4 Affordance,
5 AtomicActionEngine,
6 ControlPartCommandProfile,
7 ObjectSemantics,
8 )
9 from embodichain.lab.sim.motion.motion_generator import (
10 MotionGenCfg,
11 MotionGenerator,
12 )
13 from embodichain.lab.sim.motion.planners import ToppraPlannerCfg
14
15 hand_dof = len(self.robot.get_joint_ids(name=HAND_CONTROL_PART))
16 hand_open_qpos = torch.full(
17 (hand_dof,), HAND_OPEN_QPOS, dtype=torch.float32, device=self.device
18 )
19 hand_close_qpos = torch.full(
20 (hand_dof,), HAND_CLOSE_QPOS, dtype=torch.float32, device=self.device
21 )
22 motion_generator = MotionGenerator(
23 cfg=MotionGenCfg(planner_cfg=ToppraPlannerCfg(robot_uid=self.robot.uid))
24 )
25 self._action_engine: AtomicActionEngine = AtomicActionEngine(
26 motion_generator,
27 control_profiles={
28 HAND_CONTROL_PART: ControlPartCommandProfile.joint_positions(
29 open=hand_open_qpos,
30 grasp=hand_close_qpos,
31 )
32 },
33 )
34 self._stack_block_semantics: ObjectSemantics = ObjectSemantics(
35 affordance=Affordance(),
36 geometry={},
37 label=STACK_BLOCK_UID,
38 entity_id=STACK_BLOCK_UID,
39 )
40
The task constructs the motion generator once, declares command profiles for the gripper, and reuses the engine for every episode. Object semantics are explicit inputs to the skills rather than being inferred from simulator names.
Return semantic demonstration segments#
create_demo_segments() is the preferred handwritten expert API:
1 def create_demo_segments(self, **kwargs: Any) -> Iterable[DemoSegment]:
2 """Yield grasp/lift and placement subgoals from one continuous plan.
3
4 Args:
5 **kwargs: Reserved for expert-planning options.
6
7 Yields:
8 A measured lift followed by a settled stacking placement.
9 """
10 del kwargs
11 (
12 pick_success,
13 place_success,
14 pick_trajectory,
15 place_trajectory,
16 source_pose,
17 target_pose,
18 ) = self._plan_stack()
19 yield DemoSegment(
20 actions=self._iter_segment_actions(pick_trajectory, settle=False),
21 name="pick_block_2",
22 target_uid=STACK_BLOCK_UID,
23 instruction="Grasp and lift block 2 clear of the table.",
24 progress_total_steps=int(pick_trajectory.shape[1]),
25 metadata={
26 "segment_index": 0,
27 "segment_count": 2,
28 "planning_success": pick_success.detach().cpu().tolist(),
29 "source_pose": source_pose.detach().cpu().tolist(),
30 "atomic_actions": ["pick_up"],
31 },
32 validator=partial(
33 self._validate_pick,
34 pick_success.detach().clone(),
35 source_pose.detach().clone(),
36 ),
37 )
38 yield DemoSegment(
39 actions=self._iter_segment_actions(
40 place_trajectory,
41 clear_grasp_dynamics=False,
42 ),
43 name="place_block_2_on_block_1",
44 target_uid=STACK_BLOCK_UID,
45 instruction="Place block 2 on top of block 1, release it, and let the stack settle.",
46 progress_total_steps=int(place_trajectory.shape[1]) + SETTLE_STEPS,
47 metadata={
48 "segment_index": 1,
49 "segment_count": 2,
50 "planning_success": place_success.detach().cpu().tolist(),
51 "target_pose": target_pose.detach().cpu().tolist(),
52 "atomic_actions": ["place"],
53 },
54 validator=partial(
55 self._validate_stack,
56 (pick_success & place_success).detach().clone(),
57 ),
58 )
59
The two segments keep three outcomes separate:
actionsis the controller-command stream consumed by the common runner;planning_successrecords whether PickUp and Place were planned; andvalidatorchecks measured lift after the grasp, then planning success and the physical stack after placement and settling.
The task creates typed PickUp and Place requests and threads the projected held-object state from the first plan into the second:
Compile the PickUp and Place trajectory
1 pick_compiled = self._action_engine.compile(
2 (
3 ActionInvocation(
4 skill_id="pick_up",
5 goal=GraspGoal(
6 self._stack_block_semantics,
7 grasp_xpos=grasp_pose,
8 ),
9 binding=pick_binding,
10 motion_policy=MotionPolicy(sample_count=PICK_SAMPLE_INTERVAL),
11 skill_options=PickUpOptions(
12 pre_grasp_distance=0.12,
13 lift_height=0.15,
14 hand_interp_steps=HAND_INTERP_STEPS,
15 ),
16 ),
17 ),
18 self._action_engine.initial_context(
19 scene=SceneSnapshot(
20 timestamp=0.0,
21 version=0,
22 entities={STACK_BLOCK_UID: EntityState(source_pose)},
23 ),
24 control_dt=self.step_dt,
25 ),
26 )
27 pick_success = pick_compiled.plan_success
28 pick_trajectory = pick_compiled.trajectory.positions
29 picked_context = pick_compiled.projected_context
30 pick_trajectory = self._insert_grasp_hold(pick_trajectory)
31
32 target_pose = source_pose.clone()
33 target_pose[:, :3, 3] = base_pose[:, :3, 3]
34 target_pose[:, 2, 3] += BLOCK_HEIGHT
35 held = picked_context.get_held_object(CONTROL_PART)
36 if held is None or not bool(pick_success.all().item()):
37 if held is None:
38 pick_success = torch.zeros_like(pick_success, dtype=torch.bool)
39 return (
40 pick_success,
41 torch.zeros_like(pick_success, dtype=torch.bool),
42 self._ensure_nonempty_trajectory(pick_trajectory),
43 pick_trajectory[:, :0],
44 source_pose,
45 target_pose,
46 )
47
48 place_eef_pose = torch.bmm(target_pose, held.object_to_eef)
49 place_compiled = self._action_engine.compile(
50 (
51 ActionInvocation(
52 skill_id="place",
53 goal=PlaceGoal(place_eef_pose),
54 binding=place_binding,
55 motion_policy=MotionPolicy(sample_count=PLACE_SAMPLE_INTERVAL),
56 skill_options=PlaceOptions(
57 lift_height=0.10,
58 hand_interp_steps=HAND_INTERP_STEPS,
59 ),
60 ),
61 ),
62 picked_context,
63 )
64 place_success = place_compiled.plan_success
65 place_trajectory = place_compiled.trajectory.positions
66 return (
67 pick_success,
68 place_success,
69 self._ensure_nonempty_trajectory(pick_trajectory),
70 self._ensure_nonempty_trajectory(place_trajectory),
71 source_pose,
72 target_pose,
73 )
74
For a multi-object episode, yield segments lazily. The
BlocksRankingRGBEnv.create_demo_segments() example yields the red-block
grasp and placement first, then queries the updated reference-block pose before
planning the blue-block grasp and placement. This avoids planning later
subtasks against stale scene state.
Legacy tasks implementing create_demo_action_list() remain supported and
are wrapped as one segment named legacy. New tasks should implement
create_demo_segments() so subtask boundaries, language, metadata, and
validators remain explicit.
Try the handwritten expert#
The shipped stacking config includes a dataset recorder. First smoke-test the expert while suppressing dataset writes:
embodichain run-env \
--gym_config embodichain_tasks/configs/tasks/manipulation/tableware/stack_blocks_two/env.json \
--headless \
--filter_dataset_saving \
--max_episodes 1
To collect it, run the same config without --filter_dataset_saving. Customize
the recorder using Configure Dataset Recording:
embodichain run-env \
--gym_config embodichain_tasks/configs/tasks/manipulation/tableware/stack_blocks_two/env.json \
--headless \
--max_episodes 5
Paradigm 2: Task Program#
Task Program replaces the task-specific expert method with declarative task intent. It does not introduce a second rollout or recording API: the configured adapter compiles the program, creates lazy demonstration segments, and returns them to the same executor used by handwritten tasks.
The current Pour Water example is a complete configuration-defined task:
env.yamlowns the physical scene, environment values, and dataset recorder, without any Task Program fields;task.cobotmagic.yamlis the runnable deployment that selects the environment, Task Program, embodiment, and execution policy;program.yamlowns embodiment-independent intent and targets;integration.yamlowns task-specific contracts, scene binding, semantic defaults, action options, and allowlisted runtime services;the embodiment component owns the robot, sensors, endpoints, and compatible skill profile.
Compose the deployment#
The runnable deployment is intentionally thin:
1id: PourWater-v1
2
3environment:
4 component: env.yaml
5
6task_program:
7 program: task_program/program.yaml
8 integration: task_program/integration.yaml
9 execution_policy: ../../../../components/execution_policies/trajectory_open_loop_slow.yaml
10
11embodiment:
12 component: ../../../../components/embodiments/cobotmagic.yaml
Its selected env.yaml is reusable outside Task Program because it contains
only physical simulation and ordinary Gym values:
Pour Water reusable environment
1environment_id: pour_water
2physics: default
3max_episodes: 5
4max_episode_steps: 2400
5num_envs: 1
6arena_space: 3.0
7
8render_cfg:
9 tone_mapping_enabled: true
10 tone_mapping_exposure: 1.0
11
12simulation:
13 light:
14 direct:
15 - uid: light_1
16 light_type: rect
17 color: [1.0, 0.95, 0.88]
18 intensity: 5.0
19 init_pos: [0.35, -0.5, 2.1]
20 direction: [0.35, 0.4, -1.0]
21 rect_width: 1.0
22 rect_height: 0.8
23 - uid: fill_light
24 light_type: rect
25 color: [0.88, 0.94, 1.0]
26 intensity: 1.5
27 init_pos: [1.25, 0.6, 1.7]
28 direction: [-0.5, -0.5, -1.0]
29 rect_width: 0.8
30 rect_height: 0.8
31 background:
32 - uid: table
33 shape:
34 shape_type: Mesh
35 fpath: PourWaterAssets/table.glb
36 attrs:
37 mass_props:
38 mass: 10.0
39 material_props:
40 static_friction: 0.95
41 dynamic_friction: 0.9
42 restitution: 0.01
43 body_scale: [1, 1, 1]
44 body_type: kinematic
45 init_pos: [0.725, 0.0, 0.825]
46 rigid_object:
47 - uid: cup
48 shape:
49 shape_type: Mesh
50 fpath: PourWaterAssets/cup.glb
51 collision:
52 approximation: convex_decomposition
53 max_hulls: 16
54 attrs:
55 mass_props:
56 mass: 0.03
57 rigid_props:
58 max_depenetration_velocity: 10.0
59 min_position_iters: 32
60 min_velocity_iters: 8
61 collision_props:
62 contact_offset: 0.003
63 rest_offset: 0.001
64 material_props:
65 static_friction: 0.9
66 dynamic_friction: 0.8
67 restitution: 0.01
68 init_pos: [0.75, 0.1, 0.921]
69 body_scale: [1, 1, 1]
70 - uid: bottle
71 shape:
72 shape_type: Mesh
73 fpath: PourWaterAssets/bottle.glb
74 collision:
75 approximation: convex_decomposition
76 max_hulls: 8
77 attrs:
78 mass_props:
79 mass: 0.01
80 rigid_props:
81 max_depenetration_velocity: 10.0
82 min_position_iters: 32
83 min_velocity_iters: 8
84 collision_props:
85 contact_offset: 0.003
86 rest_offset: 0.001
87 material_props:
88 restitution: 0.01
89 init_pos: [0.75, -0.1, 0.962]
90 body_scale: [1, 1, 1]
91 rigid_object_group: []
92 articulation: []
93
94env:
95 # Update position targets at the physics cadence instead of holding each
96 # target for four substeps, which produces in-hand acceleration impulses.
97 target_control_frequency: 100
98 events:
99 settle_pour_objects_on_reset:
100 func: wait_for_dynamic_objects_to_settle
101 mode: reset
102 params:
103 entity_cfgs:
104 - uid: bottle
105 - uid: cup
106 min_steps: 10
107 max_steps: 180
108 check_interval_steps: 2
109 required_stable_checks: 3
110 timeout_behavior: raise
111 dataset:
112 lerobot:
113 func: LeRobotRecorder
114 mode: save
115 params:
116 robot_meta:
117 robot_type: CobotMagic
118 extra:
119 scene_type: Commercial
120 task_description: Pour water
121 data_type: sim
122 use_videos: true
123 control_parts: [left_arm, left_eef, right_arm, right_eef]
During config_to_cfg(), component paths are resolved relative to
task.cobotmagic.yaml. The resolver expands the physical environment and
embodiment, checks scene-binding UIDs and semantic contracts, and binds trusted
provider identities into the otherwise embodiment-independent program.
Declare the semantic workflow#
The program describes Pick, transport, Pour, and Place without importing Python callables or naming simulator joints:
Pour Water Task Program
1program_id: pour_water_with_right_arm
2instruction: Pour water from the bottle into the cup and return the bottle to the table.
3targets:
4 bottle_return_pose:
5 kind: cyclic_pose
6 values:
7 - position: [0.75, -0.1, 0.962]
8 quaternion_xyzw: [0.0, 0.0, 0.0, 1.0]
9 bottle_release_pose:
10 kind: cyclic_pose
11 values:
12 # Open above the support instead of driving a held bottle into the table.
13 - position: [0.75, -0.1, 0.969]
14 quaternion_xyzw: [0.0, 0.0, 0.0, 1.0]
15 bottle_pour_pose:
16 kind: cyclic_pose
17 values:
18 - position: [0.7776451358490485, 0.04643509710227371, 1.101]
19 quaternion_xyzw: [0.0, 0.0, 0.0, 1.0]
20 cup_rest_pose:
21 kind: cyclic_pose
22 values:
23 - position: [0.75, 0.1, 0.921]
24 quaternion_xyzw: [0.0, 0.0, 0.0, 1.0]
25program:
26 kind: sequence
27 items:
28 - kind: segment
29 name: pick_bottle
30 instruction: Grasp and lift the bottle.
31 steps:
32 kind: invoke
33 call:
34 kind: pick
35 object: bottle
36 grasp: bottle_grasp
37 - kind: segment
38 name: move_bottle_above_cup
39 instruction: Move the held bottle above the cup.
40 steps:
41 kind: invoke
42 call:
43 kind: registered
44 call_id: simulation.move_held_object
45 arguments:
46 target: cup_pour_pose
47 # Read the physical bottle pose before allowing the pour segment.
48 validators:
49 - kind: object_near_target
50 object: bottle
51 target: bottle_pour_pose
52 position_tolerance: 0.01
53 - kind: segment
54 name: pour_water
55 instruction: Tilt the bottle to pour water into the cup and return it upright.
56 steps:
57 kind: invoke
58 call:
59 kind: registered
60 call_id: simulation.pour
61 arguments:
62 object: bottle
63 - kind: segment
64 name: return_bottle
65 instruction: Place the bottle back on the table and release it.
66 steps:
67 kind: invoke
68 call:
69 kind: place
70 object: bottle
71 at:
72 kind: target_ref
73 target: bottle_release_pose
74 post:
75 - kind: wait_stable
76 entity: bottle
77 - kind: wait_stable
78 entity: cup
79 validators:
80 - kind: object_near_target
81 object: bottle
82 target: bottle_return_pose
83 position_tolerance: 0.05
84 - kind: object_near_target
85 object: cup
86 target: cup_rest_pose
87 position_tolerance: 0.002
The trusted integration maps bottle and cup to physical scene objects,
selects the primary manipulator, configures action options, and allowlists the
registered transport and pour lowerers:
Pour Water integration
1integration_id: pour_water_v1
2program_id: pour_water_with_right_arm
3
4requires:
5 scene_contract: pour_water_scene_v1
6 embodiment_contract: single_arm_parallel_gripper
7
8scene_binding:
9 contract_id: pour_water_scene_v1
10 registry_id: task_program_pour_water
11 rigid_objects:
12 - entity_id: cup
13 simulation_uid: cup
14 dynamics: dynamic
15 semantic_type: cup
16 - entity_id: bottle
17 simulation_uid: bottle
18 dynamics: dynamic
19 semantic_type: bottle
20 affordances:
21 - entity_id: bottle_grasp
22 kind: antipodal_grasp
23 # Wrist roll axis expressed in the bottle's grasp frame.
24 internal_axis: [0.8665034216893454, 0.49392778336146054, 0.07216068891225036]
25
26profile:
27 defaults:
28 pick_up:
29 primary: primary_manipulator
30 move_held_object:
31 primary: primary_manipulator
32 pour:
33 primary: primary_manipulator
34 place:
35 primary: primary_manipulator
36 action_options:
37 pick:
38 kind: pick_up
39 hand_interp_steps: 44
40 grasp_settle_steps: 100
41 fixed_object_to_eef:
42 - -0.0530918874
43 - 0.4963395894
44 - 0.8665033579
45 - 0.035870254
46 - -0.0525476672
47 - -0.8679134846
48 - 0.493927747
49 - 0.0204655528
50 - 0.9972059727
51 - -0.0193091929
52 - 0.0721606836
53 - 0.0321167707
54 - 0.0
55 - 0.0
56 - 0.0
57 - 1.0
58 simulation.move_held_object:
59 kind: move_held_object
60 simulation.pour:
61 kind: pour
62 rotate_angle: -1.0471975511965976
63 place:
64 kind: place
65 hand_interp_steps: 44
66 release_settle_steps: 60
67 preserve_current_object_orientation: true
68 effect_monitors: {}
69
70runtime_services:
71 registered_semantic_lowerers:
72 - kind: move_held_object
73 target_id: cup_pour_pose
74 reference_entity_id: cup
75 relative_pose:
76 - 1.0
77 - 0.0
78 - 0.0
79 - 0.027645135849048517
80 - 0.0
81 - 1.0
82 - 0.0
83 - -0.053564902897726294
84 - 0.0
85 - 0.0
86 - 1.0
87 - 0.18
88 - 0.0
89 - 0.0
90 - 0.0
91 - 1.0
92 - kind: pour
93 object_id: bottle
The program remains provider-independent. Simulation UIDs, grasp affordances, fixed object-to-end-effector transforms, and live target resolution stay in the trusted integration. For a detailed deployment walkthrough, see Configure and Run an Embodied Task Program; for constructing and compiling the language directly in Python, see Authoring a Task Program in Python.
Run the configured Task Program#
Pour Water already configures LeRobotRecorder in env.yaml, so its
runnable deployment can record directly:
embodichain run-env \
--gym_config embodichain_tasks/configs/tasks/manipulation/tableware/pour_water/task.cobotmagic.yaml \
--headless \
--max_episodes 1
The Task Program bridge plans and yields actions lazily. It never calls
env.step(); stepping, annotations, validation, commit, retry, and discard
remain owned by the shared environment rollout.
Configure Dataset Recording#
The expert paradigm and recorder configuration are independent. Add a dataset
functor under env.dataset in an inline Gym config or in a reusable
env.yaml environment component. The example below also declares a
handwritten task’s overall goal; a Task Program declares that goal on its
program root instead:
max_episodes: 5
max_episode_steps: 600
env:
task_instruction: Pick and place the object.
dataset:
lerobot:
func: LeRobotRecorder
mode: save
save_failed_episodes: false
params:
save_path: outputs/lerobot/expert_demos
robot_meta:
robot_type: CobotMagic
extra:
scene_type: tabletop
task_description: pick_and_place
data_type: sim
use_videos: true
Important fields are:
env.task_instructionowns a handwritten task’s overall goal. Task Programs use rootinstructioninprogram.yaml. Segment instructions belong to the task’sDemoSegmentvalues or program segments. Recorderinstructionis a legacy fallback; see Task and segment language for expert data.collection.target_episodesis the exact number of persisted environment rows, not the number of vector batches. The legacymax_episodessetting and--max_episodesoption map to this target.max_episode_stepsmust exceed the longest valid expert execution, including gripper holds and settling actions.save_failed_episodesbelongs besidefuncandmode. It defaults tofalse; when enabled, a failed or truncated attempt with recorded frames is committed withsuccess=falsemetadata and counts towardcollection.target_episodes.params.save_pathis the parent directory for auto-numbered datasets. If omitted, the default is~/.cache/embodichain_datasetsor the value ofEMBODICHAIN_DATASET_ROOT.params.use_videoscontrols RGB dataset videos. It has an effect only when image observations from configured sensors are present.Dataset frequency is derived from
env.step_dtand must be an integer number of frames per second for LeRobot.
The real Pour Water recorder block can be inspected directly:
1 dataset:
2 lerobot:
3 func: LeRobotRecorder
4 mode: save
5 params:
6 robot_meta:
7 robot_type: CobotMagic
8 extra:
9 scene_type: Commercial
10 task_description: Pour water
11 data_type: sim
12 use_videos: true
13 control_parts: [left_arm, left_eef, right_arm, right_eef]
Execution, Validation, and Persistence#
Without --preview or --replay, embodichain run-env performs offline
data expansion:
Resolve
create_demo_segments()from the handwritten task or configured Task Program bridge.Execute every yielded action through
env.step(action)and record the resulting transition.Run the segment validator after its action iterable is exhausted.
Check episode termination and final task success.
Commit selected environment rows with an explicit reset, or discard the attempt with
reset(options={"save_data": False}).
Failed attempts are discarded and retried by default, up to
collection.max_attempts (with demo_max_attempts as a legacy fallback).
Empty plans and exceptions are always
discarded because they do not form a complete dataset transaction. With
save_failed_episodes: true, a failed or truncated attempt is retained only
when every selected row contains recorded frames.
collection.target_episodes controls the final dataset size and num_envs
controls only the parallel width. If target_episodes=10 and num_envs=4,
the runner uses three vector batches and commits only two rows from the final
batch. The legacy --max_episodes option maps to the same target field.
For a configured Expansion task, keep the collection policy beside the task-facing expansion overrides:
runtime:
num_envs: 16
max_episode_steps: 1200
collection:
target_episodes: 64
max_attempts: 3
selection:
mode: sequential
start_recipe_index: 0
Sequential selection consumes logical recipe IDs from the start index and automatically groups them into batches. Explicit selection is useful when a small, known set of recipes is required:
collection:
target_episodes: 4
selection:
mode: explicit
recipe_indices: [0, 16, 32, 48]
The explicit recipe list must contain exactly target_episodes entries.
Recipe IDs identify logical candidates; they do not represent batch offsets.
Collection terms have one meaning across handwritten and Expansion tasks:
Term |
Meaning |
|---|---|
|
One environment row executed to completion and submitted as data. |
|
One execution try. A failed attempt can be discarded and retried. |
|
The selected environment rows executed in one vectorized pass. |
|
The maximum number of rows available in one batch. |
|
The stable logical ID of an Expansion candidate. |
Changing num_envs changes the number of batches and the parallel capacity;
it does not change target_episodes. A failed attempt and any prepare,
commit, or discard reset also leave the target episode count unchanged.
Collection accounting distinguishes successful data rows from execution attempts. The expansion manifest records the target, planned, committed, and rejected episode counts, total attempts, batch count, and prepare, commit, and discard reset counts. Prepare reset runs before a batch, commit reset finalizes selected rows, and discard reset clears a failed attempt. Reset boundaries do not add episodes.
An episode is the complete task; a segment is one semantic subtask. Do not use
generate_function(num_traj=...) to repeat subtasks: direct callers may pass
only None or 1. Yield multiple DemoSegment objects instead so the
task owns their order, live-state dependencies, and validation.
Useful modes and options are:
--headlessdisables the GUI for collection throughput.--previewopens interactive inspection and does not save a dataset.--filter_dataset_savingexecutes the expert while suppressing structured dataset writes.--num_envsoverrides collection parallelism.--max_episodesoverrides the configuredcollection.target_episodes.collection.max_attemptscontrols retries for one selected batch.
See Running Environments with run-env for preview, dataset recording, debug video, trajectory recording, and replay modes, and CLI Reference for the complete argument list.
Recorded Data#
LeRobotRecorder creates an auto-numbered dataset directory containing:
data/for Parquet action, state, and annotation features;videos/for RGB observations whenuse_videosis enabled;meta/for LeRobot metadata, task/subtask mappings, and EmbodiChain episode metadata.
The primary fields include observation.state, action, and
observation.images.{sensor_name}. Segment-aware episodes additionally
record subtask_index and annotation.segment_* boundaries, plus terminal
and truncation annotations. The overall goal resolves as sample["task"];
the current segment instruction resolves as sample["subtask"] through
meta/subtasks.parquet. Both levels remain present in segment fragments.
Depth and segmentation observations have their own
numeric or configured sidecar representation; see
Dataset Functors for the complete schema.
Inspect Recorded LeRobot Data#
Use EmbodiChain’s structural preview on the parent directory of auto-numbered datasets:
embodichain preview_lerobot_data \
outputs/lerobot/expert_demos \
--latest \
--episode 0
For Pour Water without an explicit save_path, use the default parent:
embodichain preview_lerobot_data \
~/.cache/embodichain_datasets \
--latest \
--episode 0 \
--expect-segments 4
The command validates frame and timestamp continuity, one episode-level task,
subtask mappings, contiguous segment ranges, terminal annotations, and the
EmbodiChain metadata sidecar. --expect-segments is an optional assertion;
omit it when the expert’s segment count is data-dependent.
Use LeRobot’s lerobot-dataset-viz with the exact auto-numbered dataset
directory when interactive Rerun plots and camera playback are needed. The
EmbodiChain preview focuses on structure and annotations; it does not render
images or time-series plots.
Best Practices#
Keep planning and stepping separate: experts yield actions; the runner calls
env.step().Validate physical outcomes from live simulator state, not only planner success or projected semantic state.
Yield later subtasks lazily when they depend on object poses changed by earlier segments.
Smoke-test with
--filter_dataset_savingbefore a long collection run, then inspect one committed episode before scalingnum_envs.Keep each task’s reusable environment component (or its
envs/default.yamlandenvs/newton.yamlvariants),task.<embodiment>.yaml, andtask_program/{program,integration}.yamltogether so physical UIDs, contracts, and canonical IDs stay aligned. Keep task-facing expansion overrides underexpansion/and reusable embodiment and execution-policy components underconfigs/components/.Move reusable motion behavior into Atomic Skills instead of copying task-local trajectory logic across Python tasks or registered Semantic Calls.