HardwareHomie Toolkit

Introducing HOMIE
Gen2: Unlocking
Experience Scaling
Law
for Physical AI

Introduction

Physical AI is missing the raw material for its next scaling curve: human xperience™.

Ropedia is building the infrastructure to capture, structure, and scale that experience — turning how humans see, move, and interact with the physical world into training-ready intelligence.

With HOMIE Gen2, we unlock the Experience Scaling Law for Physical AI.

Key stats

Technical Specifications

Coverage

0 Channel

Spatial audio

0 FPS

Global shutter

0P

Video quality

Up to 0TB

Storage

0g

Light weight

ParameterValueParameterValue
Dimensions300×240×138 mmWeight380 g
Cameras4Video Resolution1080P @ 30fps
Video FormatMP4Camera FOVHorizontal 162°,
Vertical 109°
IMU200 fpsBluetoothBluetooth 5.2
(BR/EDR & BLE)
Built-in Mics4 Digital MEMS MicrophonesStoragemicroSD Card (up to 2 TB)
BuzzerPrompt Sound for Different ScenariosExternal PortsUSB Type-C ×1 (Data),
USB Type-C ×1 (Full Function),
PPS Port ×2
Power Input5V DC InputBattery3.6V 3000mAh Lithium Battery (Built-in)
Battery LifeUp to 90 min recording / ~4 hrs standbyWi-FiWi-Fi 5 (IEEE 802.11 a/b/g/n/ac),
2.4 GHz & 5 GHz

Available Delivery Options

01 /

HOMIE Gen2 Hardware

The capture device, ready to deploy.

The capture device, ready to deploy. Delivers the HOMIE Gen2 headset with local data storage and multimodal sensing hardware — everything needed to start capturing in the real world.

02 /

HOMIE Gen2 + Hardware Management

Stay connected. Stay in control.

The hardware management platform extends Gen2 with remote device oversight — monitor device status, push OTA firmware updates, and keep fleets of devices in sync across varied environments.

03 /

HOMIE Gen2 +
Hardware Management +
Data Capture Suite

From capture to data, in one stack.

Data Capture Management Platform

Distribute capture tasks to devices, receive and review uploaded recordings, and manage the full data pipeline from a centralized platform.

Data Capture App

Connect to HOMIE Gen2, receive assigned capture tasks, and control recording directly from the field.

Chapter 01 · Foundations

Human Xperience™ — beyond perception.

Why physical AI needs 4D data.

We use Xperience deliberately. The X stands for the many dimensions of human experience — what humans see, hear, touch, move, and interact with as they live and act in the real world.

Human Xperience is not just perception. It is perception tightly coupled with action, intention, and physical consequence — what we call 4D data: spatial structure, motion, interaction, and change over time.

To capture it faithfully, we must start from the perspective where experience originates.

Capturing Xperience from the egocentric perspective

If we want AI systems that understand and operate in the physical world, we must teach them from the viewpoint that best reflects real human behavior: the egocentric perspective.

From the egocentric perspective, human experience naturally includes:

  • Rich human–object interactions
  • Spatial context and scene affordances
  • Social signals from other people
  • Fine-grained motor actions unfolding over time

This is exactly the information an intelligent embodied AI system needs to learn how to move, manipulate, and reason about the physical world.

Why egocentric scales

Egocentric capture scales. A body-mounted system, without external infrastructure or environmental instrumentation, enables truly in-the-wild data collection across homes, workplaces, factories, hospitals, and beyond.

This scalability unlocks the diversity, realism, and long-tail behaviors that are critical for robust generalization in robotics and embodied AI.

Chapter 02 · The Platform

HOMIE Gen 2

A full-stack platform for in-the-wild human experience capture.

To make large-scale human Xperience capture possible, we built HOMIE (Human-centric OMni Interaction & Experience) — a full-stack hardware-software platform designed for frictionless, lossless human-centric data capture at scale and low cost.

HOMIE Gen 2 head-mounted capture device

Hardware — head-mounted, multi-modal egocentric data capture

HOMIE's hardware is engineered for accurate spatial understanding and long-term real-world use:

  • Lightweight, ergonomic head-mounted form factor
  • Multi-modal sensing for rich physical and spatial signals
  • Robust ego-motion tracking and localization
  • Long battery life for all-day in-the-wild capture
  • Precise sensor synchronization for high-fidelity 3D reconstruction

Our goal is simple: make capturing human Xperience as natural as wearing glasses.

Immersive Human Experience—See the World From Where Humans Experience It

Most robotic datasets observe people from the outside. HOMIE captures from within the experience. Four wide-angle cameras provide 162° of coverage each, creating a full 360° view around the wearer.

This means the dataset captures not only what the person is looking at, but also what is happening around them — including objects, spatial relationships, movement, and environmental changes outside the immediate field of view.

Because physical tasks rarely happen in a single direction, 360° coverage preserves the context that conventional forward-facing capture misses.

What humans see. What humans do. Where humans are. What the world does in response.

All captured as one continuous and immersive experience.

Software — Spatial foundation models for automatic annotation

Raw experience alone is not enough. Intelligence requires structure.

HOMIE includes a suite of proprietary spatial foundation models that automatically transform captured Xperience into machine-readable, model-ready annotation, including:

  • Spatial localization and stereo depth estimation
  • Hand–object interaction tracking
  • Full-body motion capture (mocap)
  • Panoramic scene perception and understanding
  • …and more to come

Together, these models convert unstructured human experience into structured interactive intelligence datasets — ready for training the next generation of robotics models, world models, and embodied AI systems.

Output — Rich Multimodal Understanding

Physical intelligence cannot be learned from a single sensor.

HOMIE Gen2 synchronizes multimodal signals into a common temporal and spatial representation, allowing models to understand human behavior as a complete sequence rather than a collection of disconnected streams.

The Xperience data engine can include 10+ different modalities:

  • Rectified Stereo Video
  • Depth Maps
  • Full-body motion capture (mocap)
  • 21 Hand Keypoints
  • 52 Full-Body Keypoints
  • Three-Tier Text Annotations: Main Task / Subtask / Current Action
  • 3D Gaussian Representations
  • Scene Meshes
  • 2D Object Detection
  • 3D Object Detection & Perception
  • 4D Object Tracking
  • …and more to come — Check our dataset Xperience 10M

Chapter 03 · RE-LIVING THE DATA

ReXperience.

How embodied AI learns from human Xperience.

Intelligence does not emerge from passive observation alone. It emerges through ReXperience: the process by which AI systems repeatedly relive, model, and internalize human Xperience.

We see immediate impact of ReXperience across three frontier directions in physical AI:

  1. 01

    World models - predictive intelligence from human Xperience

    Egocentric human Xperience provides realistic trajectories for learning how environments evolve and how actions lead to consequences. This grounds world models in real perception and interaction, improving physical prediction and reasoning. At scale, it enables world models that generalize across diverse, unstructured real-world scenarios.

  2. 02

    Real2Sim - better robot training simulations from human-captured reality

    Simulation remains essential for robot learning, but its fidelity is fundamentally limited by the realism of its assets. Human Xperience supplies natural motion patterns, contact dynamics, and task distributions that simulations often miss. By grounding simulation in real human behavior, Real2Sim produces training environments that are more representative, diverse, and effective for robotics.

  3. 03

    VLA models - scaling generalization through egocentric action grounding

    Large-scale egocentric Xperience exposes Vision-Language-Action (VLA) models to how humans actually perceive, act, and describe tasks in real environments. This grounding leads to much stronger generalization across new tasks, instructions, and scenes. As Xperience scales, so does the model's ability to adapt and act intelligently in the open world.

Chapter 04 · Closing

Our mission.

Defining the data foundation for physical intelligence.

If we want AI that moves like us, understands like us, and helps us in everyday physical tasks, then its intelligence must be built from human experience itself. At Ropedia, we are building the data infrastructure that captures, structures, and transforms human Xperience into the foundation for physical intelligence.

Closing thought

Intelligence begins with Xperience. We're here to scale it.

// TALK TO US

Building world models,
robots, or embodied agents?

If you are building world models, robotics systems, or embodied agents and care deeply about learning from real-world human experience data we'd love to talk.

Contact Us