Egocentric Data
Human-centered video and demonstrations captured from real-world environments, first-person and continuous, providing high-signal training data for embodied AI.
Real-world human data that helps intelligent machines see, understand, learn, and act.

Human-centered video and demonstrations captured from real-world environments, first-person and continuous, providing high-signal training data for embodied AI.
Structured demonstrations of real-world tasks and workflows, performed by people at natural speed and annotated for robot learning.
Object interactions, action labels, timestamps and task-level annotations aligned frame by frame for physical AI training pipelines.
Datasets shaped for humanoid robots, manipulators and embodied AI systems operating in unstructured, real-world spaces.
From single-camera footage to full VR teleoperation, we capture, structure and annotate real-world and robot data in the format your models actually train on: mono RGB, stereo, stereo+wrist, depth, teleoperation, VR/Quest/Pico, iPhone and OTS robot data.
Single-camera RGB video for lightweight, high-volume data collection across everyday tasks and environments.
Calibrated stereo camera pairs for depth-aware scene understanding, giving models real 3D spatial context from two viewpoints.
Stereo head cameras combined with a wrist-mounted camera for close-range grasp and fine-manipulation detail. It's the standard rig for bimanual robot demonstrations.
Depth-aligned RGB streams for precise 3D geometry, object pose estimation and scene reconstruction.
Robot demonstrations recorded through direct human teleoperation, with joint and end-effector actions labeled frame by frame.
Immersive teleoperation data captured through Meta Quest and Pico VR headsets, giving natural, embodied control for complex manipulation tasks.
Portable, high-fidelity capture using iPhone camera and LiDAR rigs for fast, in-the-wild data collection without custom hardware.
Demonstrations collected on commercially available, off-the-shelf robot arms and grippers, delivering production-ready data without bespoke rigs.
Real-world activity recorded at source with calibrated multi-sensor rigs.
Sessions segmented into clean, time-aligned streams and reviewed for integrity.
Actions, objects, trajectories and timestamps labeled and reviewed in Foxglove, against a task taxonomy.
FoxgloveTraining-ready datasets delivered end to end as MCAP, JSON / JSON Schema, or your own schema, with full metadata and provenance intact.
MCAP · JSON · MetadataMulti-sensor capture in live environments: egocentric rigs, instrumented workspaces and robot-in-the-loop demonstrations.


