
Data Acquisition Package
Spatial AI Compute Kit
Data Acquisition Package is an EGO + UMI wearable multi-view capture system designed for robotics foundation models and embodied AI training data generation.
By synchronously recording the head-mounted first-person view and close-range views of both hands, the system turns natural human operations into high-quality robot learning data for annotation, reconstruction, and policy training.
System Components
| Module | Recommended Configuration | Function |
|---|---|---|
| Head Camera | 1 × Insight 9 | The headset with Insight 9 weighs 360g in total. It captures Ego-view RGB-D, head motion and 6DoF pose for global scene understanding. |
| Hand Cameras | 2 × Insight 3 | Each wrist-mounted Insight 3 module weighs only 158g. The two hand cameras capture UMI close-range hand interactions, complementing the Ego view with fine-grained manipulation data. |
| Backpack Unit |
| The backpack unit weighs approx. 1.8kg and integrates edge computing, local storage, and mobile power. It runs the acquisition program and stores ROS 2 bags, RGB-D, IMU, VIO pose, calibration parameters, and task metadata. Local storage can be upgraded based on recording duration and data volume requirements. The battery supports approx. 150–170 minutes of continuous real-world data capture. |
| Software Environment | Ubuntu + ROS 2 data acquisition software | Provides unified device connection, recording control, acquisition status management, data directory organization, clip management, and post-processing workflow. |
Post-Processing Software
| Capability | Description |
|---|---|
| Multi-Camera Coordinate Alignment | Automatically unifies multi-view data from the head camera and left/right hand cameras into a common spatial reference frame, eliminating inconsistencies caused by independent pose estimation across cameras — ensuring spatial consistency for joint annotation, scene reconstruction, and policy training. |
| Recording & Clip Editing | Supports unified recording and clip-based management of multi-stream data (RGB, depth, IMU, VIO pose). Quickly removes low-quality segments caused by frame drops, occlusion, or inconsistent actions — improving dataset usability and reducing manual screening costs from the source. |
| Trajectory Scoring | Leverages onboard VSLAM pose estimation from Insight cameras to automatically evaluate motion smoothness, trajectory continuity, and view completeness. Provides objective quality criteria for data ingestion and helps flag suboptimal samples efficiently. |
| Trajectory Optimization | Uses camera-provided high-quality pose data to perform automated smoothing, pose correction, and temporal alignment on jittery, drifting, or low-score trajectories. Optimized outputs are ready for downstream training, cutting manual post-processing and speeding up model iteration. |
Highlights
EGO + UMI Multi-view Capture
The head view captures the task environment, while the hand views preserve manipulation details, improving the completeness and usability of egocentric operation data.
Onboard VSLAM Output
RGB, depth, IMU, and VIO pose are recorded simultaneously, enabling annotation, reconstruction, and training workflows beyond 2D video.
Open Interfaces
Raw data can be connected to hand detection, object detection, pose estimation, RGB-D reconstruction, action segmentation, and robot learning data conversion pipelines.
Scalable Deployment
Cameras, Jetson, NVMe storage, and a battery backpack form a standardized acquisition unit, making it easier to replicate hardware setups and standardize collection workflows.