← Back to list

Two Form Factors, One Mission: Why ThirdEye Now Comes in a Wrist Mount and a Frame Cap

FileMarket AI Data Labs · 2026-05-26 18:03 · 0 claps · 8.3 min read
#ai #robotics #robotics-automation #robotics-technology #robots
Open on Medium ↗
Wiki topics: AI · AI · General

Two Form Factors, One Mission: Why ThirdEye Now Comes in a Wrist Mount and a Frame Cap

The most important design decision in egocentric data collection hardware is not the sensor. It is not the lens. It is not the processing architecture.

It is where you put the camera.

This sounds obvious. It is also the decision that most data collection approaches get wrong — and the reason FileMarket AI has spent significant engineering time developing two distinct form factors for ThirdEye rather than settling on one universal mounting solution.

The wrist mount and the sleek frame cap are not variations of the same product. They are designed for different data collection tasks, different industrial contexts, and different training objectives. Understanding why each form factor exists requires understanding what the data they capture is actually used for.

The Data Perspective Problem

Every egocentric data collection system makes a fundamental choice about perspective: where does the camera sit, and what does it see?

A head-mounted camera captures the world from the perspective of the operator’s eyes. It sees what the human sees — the full field of view that includes the workspace, the tools, the objects being manipulated, the environment beyond the immediate task, and the hands moving through that field of view. This is the perspective that matches a deployed humanoid robot’s onboard camera position — typically located in the head or chest, looking out at the world from approximately human eye height.

A wrist-mounted camera captures something different. It sees the world from the perspective of the operator’s hand — close, directed at whatever the hand is interacting with, capturing the contact points, the grip geometry, and the spatial relationship between fingers and objects at a resolution and proximity that a head-mounted camera cannot provide. This is the perspective that matters most for dexterous manipulation training — the data that teaches robot hands how to actually handle objects.

These two perspectives are complementary. They capture different aspects of the same physical interaction. And they are both necessary for training robots that can do the full range of tasks that physical AI deployment requires.

ThirdEye Option B — The Sleek Frame Cap

The Sleek Frame Cap is the head-mounted form factor. It positions the ThirdEye camera module at forehead level, capturing the full first-person field of view that a deployed humanoid robot would use when operating in the same environment.

The design philosophy prioritizes long-session comfort and deployment simplicity. The frame cap is lightweight — worn like a standard cap, adjusted with a rear strap, with the camera module positioned at the front left of the frame where it captures a natural forward field of view without obstructing the operator’s sightlines. Operators report that they stop noticing the device within the first few minutes of a session — which is the threshold that matters for collecting natural, uncontrived behavior rather than footage of someone who is consciously aware of being filmed.

The vest system that pairs with the frame cap houses the Android device and battery pack in dedicated pockets with cable management routes that keep the connection between camera and phone tidy and secure during active work. The vest is the data collection infrastructure that surrounds the camera — holding everything in place, managing the power and data routing, and allowing the operator to work freely without thinking about equipment management.

Together, the frame cap and vest form a complete wearable data collection system. A worker puts on the vest, clips on the cap, connects the USB-C cable, opens the FileMarket data collection app, and starts a session. The setup takes under five minutes. The session can run for a full industrial shift.

The data this system captures is the foundation of egocentric robot training datasets. Every task the operator performs — picking up a component, operating a machine, navigating a workspace, handling materials — is recorded from the perspective that a robot operating in the same environment would use. The spatial relationships, the visual cues, the movement patterns that experienced workers have developed over years of practice — all of it captured continuously, at 4K resolution, from the operator’s own viewpoint.

ThirdEye Option A — The Wrist Mount

The wrist mount is the manipulation form factor. It positions the same ThirdEye camera module on the back of the operator’s dominant wrist, pointing outward and slightly downward to capture whatever the hand is reaching toward and interacting with.

The engineering insight behind the wrist mount is specific to the problem of dexterous manipulation training. A head-mounted camera captures hand-object interactions as a secondary element of a broader scene. The hands appear at the bottom of the frame, the objects they interact with are visible but at a distance, and the fine-grained contact dynamics — the exact finger positions at the moment of grasp, the force distribution across the contact surface, the micro-adjustments that skilled workers make when handling irregular objects — are difficult to resolve at head-camera scale.

A wrist-mounted camera eliminates this resolution problem. Positioned on the wrist and pointing at whatever the hand approaches, it captures the close-range view of manipulation that dexterous robot policies need to learn from. When an operator reaches for a component, the wrist camera sees the component from approximately the distance that the fingertips are from it. When the operator grasps it, the camera captures the grasp geometry. When the operator places it, the camera captures the placement precision.

This is the manipulation data that is hardest to collect with conventional egocentric systems and most valuable for training dexterous robot hands. The close-range, hand-perspective view captures physical interaction detail that no external camera — and no head-mounted camera — can provide.

The wrist mount uses the same USB-C connection and Android-based collection infrastructure as the frame cap. The camera module itself is identical — the same IMX415 sensor, the same wide-angle lens, the same 4K capture capability. What changes is the mounting position and the resulting perspective.

Why Two Form Factors Rather Than One

The decision to develop two distinct form factors rather than one universal solution reflects a genuine understanding of what different physical AI training tasks require.

Many data collection systems attempt to solve the perspective problem with a single mounting position — typically head-mounted — and accept the limitations of that perspective for all data types. This works reasonably well for locomotion and navigation data, where the head perspective captures the relevant information. It works poorly for dexterous manipulation training, where the head perspective misses the close-range contact dynamics that manipulation policies need to learn from.

Some systems attempt to solve this with multiple cameras — a head mount and a wrist mount operating simultaneously. This increases data richness but also increases deployment complexity, equipment cost, and the cognitive overhead on the operator who is wearing and managing multiple devices during an active work session.

ThirdEye’s two-form-factor approach is a different answer to this tradeoff. For general industrial task data collection — workplace navigation, object handling at normal working distances, machine operation, material transport — the frame cap provides the appropriate perspective at low deployment overhead. For dexterous manipulation data collection specifically — assembly work, tool use, close-range object handling, tasks where finger-level precision matters — the wrist mount provides the close-range perspective that the frame cap cannot capture.

The choice of form factor is a data collection design decision, made based on what training task the data will support. Different data types require different perspectives. ThirdEye provides both.

The Deployment Infrastructure

Both form factors share the same underlying data collection infrastructure. USB-C connection to any Android device. The FileMarket data collection application handling session management, real-time quality monitoring, metadata tagging, and upload. Edge processing — no cloud dependency, no network requirement at point of capture. Full session data buffered locally and uploaded when connectivity is available.

This infrastructure architecture is designed for the reality of industrial deployment environments. Factories and workshops often have limited wireless connectivity. The tasks that generate the most valuable training data — machinery operation, assembly work, material handling — happen in the parts of the facility where wireless infrastructure is weakest. A data collection system that depends on continuous connectivity fails precisely where the most valuable data is being generated.

The FileMarket system captures everything locally and manages upload asynchronously. The data collection session is not affected by connectivity. The data gets where it needs to go when the session ends and the device is back in range of a reliable network.

The vest system extends operational duration by housing a high-capacity battery pack alongside the Android device. A full industrial shift — eight hours of continuous data collection — is supported without battery swap. The cable management integrated into the vest keeps the connection between camera and phone secure during active work, preventing the cable snag and disconnection issues that plague improvised data collection setups.

The Scale Architecture

Both ThirdEye form factors are designed to be deployed across many operators simultaneously, not operated by a single researcher in a lab.

This is the architectural distinction that matters most for Physical AI data collection at scale. A research-grade egocentric capture system is designed for one careful operator generating high-quality demonstrations in a controlled environment. ThirdEye is designed for twenty operators running simultaneous sessions across a factory floor, each generating continuous real-world interaction data across their entire shift.

The simplicity of setup — under five minutes, no technical knowledge required — is not a convenience feature. It is a scalability requirement. If each deployment requires a trained technician to configure and troubleshoot, the cost of running simultaneous sessions across many operators becomes prohibitive. If the device can be explained to a new operator in five minutes and reliably produces usable data without further intervention, the deployment scales as fast as you can source operators.

The Android-based processing architecture is a scalability choice as well. Android devices are available everywhere, at a range of price points, supported by an enormous ecosystem of accessories, chargers, and replacement parts. Building the data collection infrastructure on Android means that the device side of the system scales without procurement overhead — operators can use their own phones, or the operation can standardize on any Android device that meets the minimum specification.

This is the data collection infrastructure that enables large-scale egocentric dataset production from real industrial environments. Not one researcher collecting carefully chosen demonstrations. Dozens of workers, across multiple facilities, generating continuous real-world interaction data that reflects the full diversity of industrial labor.

That diversity is what makes the resulting datasets genuinely valuable for Physical AI training. Foundation models for physical AI generalize better when trained on data from many operators, many environments, many task variations. ThirdEye’s scalable deployment architecture is designed to produce exactly that kind of data.

What This Points To

The wrist mount and the frame cap are two form factors. They are also two data types — full first-person scene data and close-range manipulation data — that together cover the spectrum of what physical AI training requires.

The robots being trained on this data will not just walk and navigate. They will handle objects, use tools, and perform the dexterous manipulation tasks that make robots genuinely useful in industrial and domestic environments. Those capabilities require training data that captures manipulation at the resolution and perspective that the form factors are designed for.

FileMarket AI operates a data collection factory in Kathmandu, Nepal. The ThirdEye system — both form factors — is the hardware layer that enables that factory to produce the data that physical AI needs.

The robots of the next decade will be trained on what is collected now. ThirdEye is how that collection happens.

Learn more about FileMarket AI Data Labs: https://filemarket.ai

Data collection and partnership inquiries: humanloop@filemarket.ai


메타데이터
post_id
1bebe4ba6de6
slug
two-form-factors-one-mission-why-thirdeye-now-comes-in-a-wrist-mount-and-a-frame-cap-1bebe4ba6de6
url
https://medium.com/@filemarketai/two-form-factors-one-mission-why-thirdeye-now-comes-in-a-wrist-mount-and-a-frame-cap-1bebe4ba6de6
canonical_url
https://medium.com/@filemarketai/two-form-factors-one-mission-why-thirdeye-now-comes-in-a-wrist-mount-and-a-frame-cap-1bebe4ba6de6
author_url
https://medium.com/@filemarketai
status
ok
fetched_at
2026-06-09 15:37:30