Data requirements move faster than data operations

Frontier labs change their data requirements constantly. The challenge is that data operations do not adapt at the same speed. Human review teams cannot absorb a sudden spike in footage, and every new reviewer needs weeks of training. A new requirement means rewriting guidelines, retraining the team, and reprocessing thousands of hours of footage, while old habits linger. And even under the same guideline, no two reviewers interpret it quite the same way.

We believe agentic systems can close that gap.

Why idle time is harder than it looks

Consider one simple problem: separating real work from idle time in large volumes of egocentric video.

Motion signals alone are not enough. A worker scrolling through their phone, for example, can generate significant hand movement without performing a useful task. Short-window VLM analysis can also miss the broader context needed to interpret an action correctly.

Imagine a technician carrying a tire across an auto repair shop. If they are bringing it from storage to the lift, the clip captures a real step in the repair. If they are simply walking around with it, the same motion may add little value.

The motion looks similar. The context makes it meaningful.

How the system works

A good curator would inspect what happened before and after. An agentic system should be able to do the same. For each recording, the system works in five steps:

  1. Survey: sweeps the whole shift once and flags moments that may be idle.
  2. Classify: identifies the recording's scene, job, skill, and difficulty.
  3. Zoom: takes a closer look at each flagged moment.
  4. Double check: retrieves neighboring footage, consults IMU and hand-pose signals, and gathers more evidence when uncertain.
  5. Result: writes each confirmed idle stretch, with a start and an end, back into the dataset.

In the video, the system flags three moments that look idle. A closer look shows two are work; only one, 11:24 to 11:58, is idle.

The system decides its next step from the evidence it finds, rather than following a fixed pipeline.

Our collaboration

Linewise is building the agentic video analysis system and the harness it runs in: secure data access, temporal navigation, sensor and model tools, structured writeback, provenance, and evaluation.

Inco is running the system's inference. Per GPU, it analyzes 31× more video per hour than our native pipeline, at 1/15 the cost per hour of video. That is what makes this kind of iterative, context-seeking reasoning practical at scale.

Compute, not headcount

Because the system scales with compute rather than headcount, it absorbs spikes at the same cost per hour, picks up a new standard with a re-run instead of a retraining cycle, and applies one definition to every hour of video.

Our goal is to push routine preprocessing toward zero human involvement, while keeping expert judgment focused on requirements, edge cases, and quality oversight.

Frontier research is moving faster. The data layer needs to move with it.

More from the blog