We Went Into a Real Factory. Here Is What We Collected and Why It Matters.
We cannot name the company.
We Went Into a Real Factory. Here Is What We Collected and Why It Matters.
We cannot name the company.
They asked to remain anonymous. What we can say is that their production floor involves some of the most repetitive, dexterous, high-precision manual work in consumer goods manufacturing — tasks where skilled workers perform the same sequence of motions thousands of times per shift, handling materials that vary in texture, weight, and compliance, maintaining consistent quality standards that require exactly the kind of fine motor coordination that robots are being trained to replicate.
This is the second time our team at FileMarket AI has gone into a real industrial facility and deployed ThirdEye — our head-mounted 4K egocentric capture device — on production floor workers. The first was Jugal Garment, where we documented publicly. This one stays anonymous. The data, however, is real.
The Model
The FileMarket industry collaboration model is straightforward. We come to the facility. We brief the workers on the device and what it captures. We set up ThirdEye — five minutes per worker, no technical knowledge required — and connect it to our data collection app. The workers do their jobs exactly as they would on any other day. We collect the data. The workers are compensated through our FileFactory platform for their contribution. The company gets to participate in the Physical AI ecosystem without modifying their production process or exposing proprietary operational details.
No lab. No simulation. No staged demonstrations. No researcher standing in a model environment performing a choreographed task sequence.
Real factory. Real production shift. Real workers doing real work.
This distinction matters more than it might seem. The gap between data collected in a staged demonstration environment and data collected during actual production is significant and systematic. In a staged environment, the demonstrator knows they are being recorded. They perform the task as instructed, carefully, with awareness of the camera. They take breaks when the session ends. They do not deal with the variation, the time pressure, the accumulated fatigue, or the improvised adaptations that characterize real production work.
In an actual factory during an actual shift, none of this applies. Workers move at production speed. They handle the variation in materials and conditions that real production involves. They make the micro-adjustments and compensations that years of experience have taught them. They exhibit the full behavioral repertoire that skilled industrial work requires — not the simplified, self-conscious version of it that demonstration sessions produce.
This is the data that reflects how skilled manufacturing work is actually done. And this is the data that physical AI models need to generalize beyond the controlled conditions of laboratory training to the variable, demanding reality of industrial deployment.
What Dexterous Industrial Work Involves
The specific manipulation tasks on this production floor are not unusual for consumer goods manufacturing. They are precisely the kind of tasks that appear simple to observe and turn out to be technically demanding to replicate.
The materials involved have variable physical properties — different surface textures, different compliance profiles, different weight distributions depending on the specific item and the stage of the production process. A worker handling these materials is continuously sensing and adapting to their properties in real time, adjusting grip force, approach angle, and movement speed based on tactile and visual feedback that they have learned to interpret through thousands of hours of practice.
The task sequences involve precise spatial positioning — placing materials in specific orientations, aligning components to tolerances that affect the quality of the finished product, coordinating both hands for tasks that require simultaneous manipulation of multiple elements. These are bimanual manipulation tasks with quality requirements that create real stakes for every motion.
The production pace creates genuine time pressure. Workers are not performing for a camera at a pace they find comfortable. They are meeting production targets that determine their performance evaluation. The speed of their movements, the efficiency of their motion sequences, the way they handle unexpected variation — all of it reflects the optimization for speed and quality that professional production work demands.
From a robot training perspective, this environment generates data of a type that is extremely difficult to produce in any other setting. The combination of material variety, spatial precision requirements, bimanual coordination, and production-pace time pressure creates a training distribution that reflects industrial deployment conditions far more accurately than any staged equivalent.
Why Factory Data Is the Hardest to Collect
The physical AI data collection ecosystem has converged on several approaches to generating training data, each with distinct strengths and limitations.
Simulation generates data at scale and at low marginal cost. Physics engines can produce enormous volumes of robot interaction data quickly. The limitation is accuracy: the gap between simulated physics and real physics is particularly severe for the kind of material-handling tasks that consumer goods manufacturing involves. Cloth, soft materials, flexible components — the objects that make up much of real industrial work — are among the hardest to simulate accurately. Policies trained on simulation data for these tasks often fail in deployment because the simulated material behavior does not match the real material behavior.
Teleoperation generates robot-perspective demonstration data with natural human intent. A skilled operator guiding a robot through a task produces data that reflects how the task should be done. The limitation is scale: teleoperation requires dedicated infrastructure, a trained operator, and session-by-session management that makes large-scale data production expensive and slow.
Crowdsourced data collection — gig workers recording household tasks at home — generates diverse data at relatively low cost. The limitation is consistency and specificity: the task domains that consumer goods production requires are not the tasks that workers in their homes are performing.
Egocentric data collection in real production facilities solves these problems simultaneously. The materials are real. The physics is correct by definition. The task domain is exactly the deployment domain. The behavioral repertoire is the full range of skilled industrial work, not a simplified demonstration version. The workers are operating at production scale, generating data volume that is proportional to the facility’s actual output.
The challenge is access. Production facilities have legitimate concerns about competitive intelligence, worker privacy, production disruption, and the operational overhead of accommodating an external data collection team. These concerns are real and they explain why most robotics data collection happens in controlled environments rather than real factories.
FileMarket AI’s model is designed to address these concerns directly. We work within the facility’s operational constraints, not around them. We brief workers thoroughly on what is being captured and how it will be used. We compensate workers for their contribution through our FileFactory platform, giving them direct participation in the economic value their data creates. We do not capture anything except what ThirdEye sees from the worker’s first-person perspective — no proprietary equipment, no process documentation, no competitive information.
The result is data collection that the facility can accept without disruption and workers can participate in with clear understanding of what they are contributing.
Growing the Network
This is FileMarket AI’s second industry collaboration. Jugal Garment was the first — a garment manufacturing facility in Kathmandu where we documented the collaboration publicly. This facility asked for anonymity, which we are respecting.
The two collaborations together represent the beginning of a data collection network across Nepal’s manufacturing sector. Different industries. Different task domains. Different materials and manipulation challenges. All generating egocentric training data from real production environments.
Nepal has a manufacturing sector that spans garments, consumer goods, food processing, metalwork, electronics assembly, and more — industries where skilled manual labor performs exactly the kinds of dexterous manipulation tasks that Physical AI needs to learn from. The density and diversity of this manufacturing base, combined with a workforce that brings genuine skill and production experience, makes Nepal a genuinely valuable location for industrial egocentric data collection.
The FileFactory platform is the infrastructure that makes this network scalable. Workers register, receive their ThirdEye device, complete their collection sessions during normal working hours, and receive payment through the platform. Facility managers see the operational impact: minimal disruption, worker compensation handled externally, and participation in a global AI data ecosystem that their facility’s output is helping to build.
Every facility that joins the network adds a new task domain to the training distribution. Every new task domain makes the resulting datasets more diverse and more generalizable. The robots trained on these datasets encounter fewer situations that fall outside their training distribution — because the training distribution includes the full range of real industrial work that real factories actually do.
What Comes Next
The anonymous facility this week will not be the last. We are actively expanding our industry collaboration network across Nepal’s manufacturing sector, with conversations underway across multiple industries and facility types.
Each collaboration adds to the training distribution. Each new task domain covered reduces the gap between what physical AI can learn from lab data and what it needs to know to operate in real factories.
The robots coming to manufacturing facilities over the next decade will be trained on data collected right now. The question is whether that data reflects the full complexity and diversity of real production work — or whether it reflects the simplified, staged, lab-optimized approximation that most current datasets provide.
We are building the former. Real factory. Real data.
If your facility is interested in partnering with FileMarket AI for egocentric data collection, reach out. Any industry. Any location.
humanloop@filemarket.ai
Learn more about FileMarket AI Data Labs: https://filemarket.ai
FileFactory platform: https://filefactory.filemarket.ai



메타데이터
- post_id
- 85330e5ce20b
- slug
- we-went-into-a-real-factory-here-is-what-we-collected-and-why-it-matters-85330e5ce20b
- url
- https://medium.com/@filemarketai/we-went-into-a-real-factory-here-is-what-we-collected-and-why-it-matters-85330e5ce20b
- canonical_url
- https://medium.com/@filemarketai/we-went-into-a-real-factory-here-is-what-we-collected-and-why-it-matters-85330e5ce20b
- author_url
- https://medium.com/@filemarketai
- status
- ok
- fetched_at
- 2026-06-09 15:37:30