Dataset Functors#
This page lists all available dataset functors that can be used with the Dataset Manager. Dataset functors are configured using DatasetFunctorCfg and are responsible for collecting and saving episode data during environment interaction.
Note
This page covers structured dataset export. If you only need human-viewable debug or demo videos from a fixed camera, use record_camera_data on Event Functors.
Tip
Using an AI coding agent? Use the /add-functor skill to scaffold a new dataset functor with the correct signature, DatasetFunctorCfg registration, and module placement in datasets.py.
Recording Functors#
Functor Name |
Description |
|---|---|
Records episodes in LeRobot dataset format. Handles observation-action pair recording, format conversion, and episode saving. Requires LeRobot package to be installed. {"func": "LeRobotRecorder", "mode": "save",
"save_failed_episodes": true,
"params": {"robot_meta": {"robot_type": "CobotMagic"},
"instruction": {"lang": "Pour water from bottle to cup"},
"extra": {"scene_type": "Commercial",
"task_description": "Pour water",
"data_type": "sim"},
"use_videos": true}}
|
|
Drop-in async variant of {"func": "AsyncLeRobotRecorder", "mode": "save",
"save_failed_episodes": true,
"params": {"robot_meta": {"robot_type": "CobotMagic"},
"instruction": {"lang": "Pour water from bottle to cup"},
"extra": {"scene_type": "Commercial",
"task_description": "Pour water",
"data_type": "sim"},
"use_videos": false,
"image_writer_threads": 4}}
|
LeRobotRecorder#
The LeRobotRecorder functor enables recording robot learning episodes in the LeRobot dataset format, which can be used for training with LeRobot’s imitation learning algorithms.
Features#
Records observation-action pairs during episodes
Converts data to LeRobot format automatically
Saves episodes when they complete
Supports RGB, depth, and segmentation-mask camera observations
Supports robot state (qpos, qvel, qf)
Supports custom observation features
Auto-incrementing dataset naming
Parameters#
save_failed_episodes is a DatasetFunctorCfg lifecycle option,
not a recorder constructor parameter. Place it beside func and mode in
JSON/YAML, or pass it directly to DatasetFunctorCfg in Python. Its default
is false. During run-env expert generation:
falsediscards and retries failed or truncated attempts;truecommits a failed/truncated attempt when every selected environment row contains at least one frame. The saved sidecar recordssuccess=falseand the terminal reason;empty plans and exceptions are always discarded; and
a saved failure counts toward
max_episodes, so the requested dataset size is not exceeded.
When several save-mode functors are configured, enabling this option on any of them makes the Dataset Manager submit failed rows to all save-mode functors so their outputs stay aligned.
Parameter |
Description |
|---|---|
|
Root directory for saving datasets. Defaults to EmbodiChain’s default dataset root. |
|
Robot identity metadata for the dataset (for example, |
|
Optional task instruction (e.g., {“lang”: “pick the cube”}) |
|
Optional extra metadata (scene_type, task_description, episode_info) |
|
Whether to save videos (True) or images (False). Default: False. |
|
Number of background threads for per-frame PNG writing (lerobot |
|
Number of background processes for image writing (alternative to threads; higher spawn cost, more isolation). Use 0 to rely on threads only. |
|
Optional :class: |
Note
Dataset FPS is derived from env.step_dt and therefore always matches the
actual observation-action sampling cadence. Configure sim_steps_per_control
or target_control_frequency on the environment. LeRobot currently requires
an integer FPS, so the recorder rejects a non-integer derived control frequency
rather than writing inaccurate metadata.
Recorded Data#
The LeRobotRecorder saves the following data for each frame:
task: The episode-level task instruction, constant across its framessubtask_index: Dataset-global index of the active segment instruction; descriptions are stored inmeta/subtasks.parquetobservation.state: Joint positions (proprioceptive state)action: Applied actionobservation.images.{sensor_name}: Camera images (if sensors present)observation.images.{sensor_name}_right: Right camera images (for stereo cameras)observation.mask.{sensor_name}: Native numeric segmentation-mask arraysobservation.mask.{sensor_name}_right: Right-camera segmentation-mask arraysannotation.episode_step: Zero-based frame position inside the episodeannotation.segment_id: Episode-local semantic segment identifierannotation.segment_step: Zero-based frame position inside the segmentannotation.segment_start: 1 on the segment’s first frame, otherwise 0annotation.segment_end: 1 on the segment’s final frame, otherwise 0annotation.terminated: Gym termination flag for this frameannotation.truncated: Gym truncation flag for this frame
All annotation.* fields above and subtask_index are stored as scalar
int64 features. Booleans therefore appear as 0 or 1. The expert rollout
buffer also contains an internal valid mask, but the recorder uses it only
to select each environment’s real frame prefix; there is no
annotation.valid LeRobot column.
For legacy episodes without explicit segment metadata, the overall task is
also registered as their single subtask. subtask_index follows LeRobot
0.4.4’s optional subtask convention, while annotation.segment_id remains
episode-local and is therefore not required to have the same numeric value.
Episode-level completion, success, terminal reason, per-segment status and
segment metadata are additionally written to
meta/embodichain_episodes.jsonl.
Previewing Recorded Episodes#
Use EmbodiChain’s terminal preview when validating both standard LeRobot data and EmbodiChain’s segment metadata:
embodichain preview_lerobot_data \
path/to/lerobot_dataset \
--episode 0 \
--expect-segments 3
The path may be a parent containing auto-numbered datasets when --latest is
added. --expect-segments is an optional assertion, not a filter: validation
fails if the selected episode does not contain exactly that many segments.
A failed episode can still pass structural validation; in that case the
preview prints Sidecar : success=False while checking that its terminal
annotations and sidecar agree.
Use LeRobot’s official Rerun CLI for an interactive view of standard camera, state, and action streams:
lerobot-dataset-viz \
--repo-id organization/local-dataset-name \
--root path/to/exact/lerobot_dataset \
--mode local \
--episode-index 0
The LeRobot 0.4.4 viewer does not render subtask_index or EmbodiChain’s
annotation.* fields; use the terminal preview to verify segment boundaries,
episode steps, and terminal flags. See
Inspect Recorded LeRobot Data for the
complete three-segment example, exit codes, local-root behavior, and .rrd
export commands.
Depth and mask features keep the dtype and shape declared by the sensor
observation space. Masks are always stored as numeric LeRobot array features
(exact, lossless) so use_videos affects only the RGB image features.
Depth has two storage modes:
Numeric (default):
observation.depth.{sensor_name}numeric arrays, exact but storage-heavy for dense high-resolution depth.Compressed sidecar (when
depth_video.enable=True): depth is written asgray12le/HEVC MP4s alongside the dataset (see below), and the numeric feature is dropped unlesskeep_numeric_fallback=True.
Compressed depth sidecar#
When depth_video.enable=True and an HEVC encoder (libx265) is available,
LeRobotRecorder writes each episode’s depth maps as a single-channel
gray12le video encoded losslessly with HEVC. This is issue #424 Path A:
an EmbodiChain-owned depth writer that works on Python 3.10–3.12 with LeRobot
0.4.4, without modifying the installed LeRobot package.
Depth is quantized to 12-bit codes (logarithmic by default) and packed into the
gray12lepixel format. Withlossless=True(default) the 12-bit codes survive the encode/decode round-trip bit-exactly; the only error is the configurable float32 → 12-bit quantization step (typically sub-millimetre).Depth never enters LeRobot’s RGB-only image/video pipeline; it lives in a sidecar tree next to the dataset:
<dataset_root>/ ├── data/ # LeRobot: state / action / mask ├── videos/ # LeRobot: RGB videos ├── depth_videos/ # depth sidecar │ └── <sensor>[_right]/episode_000000.mp4 └── depth_meta.json # quantization params + per-episode index
The metadata schema (
is_depth_map,video.depth_min/max/shift/use_log,video.codec,video.pix_fmt) is aligned with the official LeRobot 0.6.0 depth pipeline, so sidecar videos remain readable by the official reader.If no HEVC encoder is available, recording silently falls back to numeric depth features (PR #422) rather than failing.
Segmentation masks are always kept as exact numeric features; they are never run through a lossy video codec.
Attention
12-bit quantization is not bit-exact relative to the original float32 or
uint16 depth input. Applications requiring exact source values must set
keep_numeric_fallback=True (stores both the sidecar video and the raw
numeric feature) or keep depth numeric.
Read sidecar depth back with :func:~embodichain.data_pipeline.depth_video.load_depth_dataset,
which composes the LeRobot dataset with a :class:~embodichain.data_pipeline.depth_video.DepthVideoLibrary:
from embodichain.data_pipeline.depth_video import load_depth_dataset
dataset, depth = load_depth_dataset("/path/to/dataset_root")
# depth in metres, shape (1, H, W), aligned to the LeRobot episode/frame index:
frame = dataset[0]
depth_map = depth.get(
episode_index=int(frame["episode_index"]),
sensor_key="camera",
frame_index_in_episode=int(frame["episode_data_index"]),
)
Dataset Recording vs Video Recording#
Need |
Use |
Why |
|---|---|---|
Training or imitation-learning data |
Saves structured observation, action, and metadata for downstream pipelines. |
|
Quick qualitative inspection or demos |
Saves MP4 videos from a dedicated camera without creating a training dataset. |
Saving Strategies#
Saving is the part of data collection that most often bottlenecks the simulator. Each completed episode triggers a save on env.reset(); with num_envs=N parallel envs truncating together, the recorder must persist N episodes worth of frames. Two independent levers control how expensive that is:
Per-frame image writing —
LeRobotRecorderwrites each camera frame to PNG synchronously insideadd_frame(compress_level=6). Setimage_writer_threads(Opt A) to offload these writes to a thread pool.Per-episode conversion + flush — by default the whole convert +
add_frame+save_episodeloop runs inline onenv.reset(), blocking the sim.AsyncLeRobotRecorder(Opt B) clones the finished episode’s buffer slice and runs that loop on a background worker, soenv.reset()returns immediately.
Situation |
Use |
|---|---|
Single env, or debugging / minimal memory |
|
Many parallel envs, sim must keep stepping |
|
Want most of the speedup without a background thread |
|
Tip
Benchmark (scripts/benchmark/data_pipeline/benchmark_lerobot_save.py, 4 envs x 2 episodes x 100 steps, 480x640, 800 frames/variant):
Variant |
t_total |
speedup |
sim blocked? |
|---|---|---|---|
|
57.4 s |
1.0x |
yes (~55 s) |
+ |
22.0 s |
2.6x |
yes (but faster) |
|
56.6 s |
~1.0x |
no (drain-bound at finalize) |
|
20.8 s |
2.8x |
no |
The sync stall grows linearly with num_envs (each reset saves N envs serially). The async recorder’s sim stall stays near zero regardless of num_envs — that is the main reason to prefer it for parallel collection.
Attention
AsyncLeRobotRecorderclones each finished episode (including camera frames) to CPU before enqueuing. For very high resolutions or many envs, monitor RSS — the worker normally keeps up, but a slow disk can let the queue grow.A single background worker touches the
LeRobotDataset(which is not thread-safe), so episode order is preserved FIFO. Always letenv.close()/dataset_manager.finalize()run so the worker drains before the dataset is finalized.Finalization is a durability barrier for episodes already submitted through
mode="save". It does not turn a live rollout into an episode; pending frames are discarded. Background and storage failures are raised to the caller, and repeated finalization is safe.Depth video and EmbodiChain metadata sidecars are part of the same durability result: encoder-close or metadata-write failures are surfaced even when the corresponding LeRobot episode has already committed, and its episode index is never reused by a later queued write.
env.close()callssim.destroy(), which exits the process without returning to Python. In scripts that build multiple envs, run each in its own subprocess and write results before closing.
Usage Example#
from embodichain.lab.gym.envs.managers.cfg import DatasetFunctorCfg
# Example: Record episodes in LeRobot format
dataset = {
"lerobot_recorder": DatasetFunctorCfg(
func="embodichain.lab.gym.envs.managers.datasets.LeRobotRecorder",
save_failed_episodes=True,
params={
"save_path": "/path/to/dataset/root",
"robot_meta": {
"robot_type": "dexforce_w1",
},
"instruction": {
"lang": "pick the cube and place it on the target",
},
"extra": {
"scene_type": "table",
"task_description": "pick_and_place",
"episode_info": {
"rigid_object_physics_attributes": ["mass"],
},
},
"use_videos": False,
},
),
}
Recording Workflow#
Initialization: The Dataset Manager initializes the functor with the configured parameters
Data Collection: During episode rollout, the functor receives observations and actions
Save Trigger: Reset selected rows to commit them. Successful rows always save; failed rows save only when
save_failed_episodes=TrueFinalization: After all episodes, call
finalize()to drain committed writes and finalize storage metadata
# Inside environment loop
if episode_done:
dataset_manager.apply(mode="save", env_ids=completed_env_ids)
# After collection completes. This raises if a committed write failed.
dataset_manager.finalize()
Parallel Collection (async recorder)#
For num_envs > 1 data collection, switch the functor to AsyncLeRobotRecorder and enable async image writing. Everything else (config structure, on-disk format, save/finalize calls) is identical:
from embodichain.lab.gym.envs.managers.cfg import DatasetFunctorCfg
dataset = {
"lerobot_recorder": DatasetFunctorCfg(
func="embodichain.lab.gym.envs.managers.async_datasets.AsyncLeRobotRecorder",
save_failed_episodes=True,
params={
"save_path": "/path/to/dataset/root",
"robot_meta": {"robot_type": "dexforce_w1"},
"instruction": {"lang": "pick the cube"},
"extra": {"scene_type": "table", "task_description": "pick_and_place"},
"use_videos": False,
"image_writer_threads": 4, # Opt A: async per-frame PNG writes
},
),
}
The async recorder drains its background worker during finalize(), so make
sure env.close() (or dataset_manager.finalize()) runs at the end of
collection. No new episode can be submitted after finalization begins.
Compressed depth recording#
To record camera depth as compressed sidecar videos (issue #424, Path A), pass a
depth_video config. Depth is encoded losslessly as gray12le/HEVC;
stereo cameras get a _right sidecar automatically.
from embodichain.lab.gym.envs.managers.cfg import DatasetFunctorCfg
dataset = {
"lerobot_recorder": DatasetFunctorCfg(
func="embodichain.lab.gym.envs.managers.datasets.LeRobotRecorder",
params={
"save_path": "/path/to/dataset/root",
"robot_meta": {"robot_type": "dexforce_w1"},
"instruction": {"lang": "pick the cube"},
"extra": {"scene_type": "table", "task_description": "pick_and_place"},
"use_videos": False,
"depth_video": {
"enable": True,
"depth_min": 0.05, # metres mapped to quantum 0
"depth_max": 5.0, # metres mapped to quantum 4095
"shift": 3.5, # log-mode offset (metres)
"use_log": True, # logarithmic quantization
"lossless": True, # 12-bit codes preserved bit-exactly
"input_unit": "m",
"output_unit": "m",
# "keep_numeric_fallback": True, # also keep raw numeric depth
},
},
),
}
Dataset Manager Modes#
The Dataset Manager supports the following modes:
save: Save completed episodes for specified environment IDsfinalize: Drain explicitly committed writes and finalize dataset resources
See DatasetManager for more details.