Refiner is Macrodata's open-source Python library for reading, transforming, and writing robotics datasets.
It provides one pipeline model for working with robot episodes, frames, videos, metadata, and model-based processing. Use it to convert formats, transform data, run inference, and write structured outputs on your own infrastructure.
This repository includes reference implementations from our published research. At Macrodata, we develop proprietary models and pipelines for robotics data processing. Contact us directly for access.
Install:
pip install macrodata-refinerThis gives you:
- the Python package as
refiner - the CLI as
macrodata
Launch a local pipeline:
import refiner as mdr
def add_preview(row):
return row.update(
preview=" ".join(row["text"].split()[:20]),
)
(
mdr.read_jsonl("input/*.jsonl")
.filter(mdr.col("lang") == "en")
.with_columns(
text=mdr.col("text").str.strip(),
text_len=mdr.col("text").str.len(),
)
.map(add_preview)
.write_parquet("s3://my-bucket/english-cleanup/")
.launch_local(
name="english-cleanup",
num_workers=2,
)
)- a consistent row and episode model for robot trajectories, frames, videos, metadata, tasks, and statistics
- readers and writers for LeRobot, HDF5, Zarr, MCAP, Parquet, JSONL, and other common data formats
- composable transforms and model inference for converting and enriching data
- open-source reference operations for motion trimming, subtask annotation, reward scoring, and hand tracking
- access to storage backends supported by
fsspec, including S3, GCP, and Hugging Face - in-process debugging and local multi-worker execution
Read the Refiner open-source documentation.
Start here:
Build a dataset:
- join the Macrodata Discord: https://discord.gg/S8kZtmBR2x