Human–AI Divergence
in Ego-Centric Action Recognition under Spatial and Spatiotemporal Manipulations

Sadegh Rahmaniboldaji1*, Filip Rybansky2*, Quoc C. Vuong2,
Anya C. Hurlbert2, Frank Guerin1, Andrew Gilbert1

1 University of Surrey   2 Newcastle University

* Equal contribution

International Journal of Computer Vision (IJCV)

What do humans and AI need to see
to recognise an action?

The same video can support very different recognition strategies.

By progressively cropping first-person videos and disrupting their temporal order, we reveal a sharp human recognition boundary—and a different pattern of failure and recovery in AI.

Research pipeline: balanced videos undergo spatial and temporal manipulation, then human and AI classification, followed by quantitative and qualitative evaluation.
Controlled changes to the same videos let us compare what information matters for humans and AI.

Research overview

Recognising an action is about more than identifying the objects in a scene. This study investigates which spatial and temporal cues support human recognition, and whether contemporary AI models use the same evidence.

Using Epic-ReduAct, derived from 36 EPIC-KITCHENS videos, we compare human observers with Side4Video across progressively reduced video crops and temporally scrambled sequences. Human-defined Minimal Recognisable Configurations (MIRCs) identify the smallest spatial crops that remain recognisable. Further reductions probe the boundary at which recognition breaks down.

Humans show sharp declines when semantically critical cues are removed. The model changes more gradually, sometimes improving as distracting context disappears. Together, the spatial and temporal analyses reveal different recognition strategies and suggest directions for more efficient, human-aligned video understanding.

Two manipulations, two windows into recognition

01 · Spatial information

How little is enough?

Progressively reduce the visible region until a recognisable crop becomes unrecognisable.

Human observers depend on sparse, meaningful object configurations and interactions. Side4Video often relies on distributed context and mid-level visual features.

02 · Temporal information

Does the order matter?

Scramble temporal blocks while preserving the spatial content and local motion within each block.

Human sensitivity depends on the action and the cues that remain. The model often changes little, revealing class-dependent sensitivity to temporal structure.

Finding the recognition boundary

A crop is a MIRC when at least 50% of human observers recognise it, but its smaller child crops fall below that threshold.

In this closing a container example, the parent crop remains recognisable to 65% of participants. Further spatial cropping reduces recognition to 15–30%. Temporal scrambling of the parent reduces it to 40%.

The temporal condition breaks global order by rearranging five video blocks while preserving local motion within each block.

Closing a container: original video, recognisable parent crop at 65 percent, four spatial sub-MIRCs at 15–30 percent, and scrambled video at 40 percent.
Spatial and temporal reductions of the same action. Figure 2 from the paper.

Key results

36Original egocentric videos
474Human-defined MIRCs
1,896Spatial sub-MIRCs

Humans lead on the original videos

Action-recognition accuracy (%) on the selected, unmanipulated video set.

Video subsetHumansSide4VideoQwen 3.5 9B
Easy100.0088.8961.11
Hard100.0055.5627.78
Overall100.0072.2244.44

These are results on the study’s curated 36-video set, not the full EPIC-KITCHENS benchmark. Qwen is included in the baseline comparison; the main reduction analyses compare humans with Side4Video.

Histogram of spatial recognition gaps: human gaps are generally larger and more positive; AI gaps cluster closer to zero.
Spatial recognition gaps for humans and the AI model. Figure 10 from the paper.

A larger drop does not mean worse recognition

MIRCs are defined around the human recognition boundary. Removing a small but critical cue can therefore produce a large human performance drop, even though humans outperform the models on the original videos.

Recognition Gap measures the change between a parent and child video. Average Reduction Rate relates the recognition loss to the amount of spatial information removed.

Human scores measure the proportion of correct responses; AI scores use confidence in the target verb. Their differences describe distinct responses to the manipulation, rather than interchangeable measures of accuracy.

When seeing less helps the model

Successively cropped frames of the action put: original video, MIRC level, and sub-MIRC level, with red rectangles indicating the next crop.
Original video, MIRC and sub-MIRC for the action put. Figure 9 from the paper.

Removing background distractors can allow the AI model to recover the correct action. This contrasts with the human recognition boundary and illustrates why the study examines failure and recovery, rather than accuracy alone.

Epic-ReduAct

A controlled testbed for studying what survives as visual information is removed. The dataset contains Easy and Hard subsets, each with 18 original videos, hierarchical spatial reductions, and temporally scrambled MIRC variants.

Spatial analysis

7,676 spatial samples, including 474 MIRCs and 1,896 spatial sub-MIRCs, support comparisons across reduction levels.

Spatiotemporal analysis

474 temporally manipulated MIRC videos probe sensitivity to global action order. Of these, 345 fall below the human recognition threshold.

Frequently asked questions

What is a MIRC?

A Minimal Recognisable Configuration is the smallest spatial crop of a video that still supports reliable human recognition. The study uses a 50% recognition threshold; smaller, unrecognisable child crops are spatial sub-MIRCs.

Does temporal scrambling remove all motion?

No. The main manipulation rearranges five contiguous blocks. It preserves local motion within each block while disrupting the global sequence, separating the effects of local motion and temporal order.

Why can AI confidence increase after cropping?

A crop may remove distracting context or change the visual features available to the model. Such recovery can reveal a different use of evidence from humans; it does not by itself demonstrate human-like understanding.

What should future video models learn from this?

The findings motivate greater attention to meaningful active objects and interactions, and evaluation at human-defined recognition boundaries. MIRC-based training is a proposed future direction, rather than a new trained model introduced here.

How broadly do the findings generalise?

The study focuses on a curated set of 36 kitchen-action videos and primarily analyses Side4Video. Extending the comparison to other domains and model families is an important next step.

BibTeX

Citation for the publicly available preprint.

@misc{rahmaniboldaji2026humanaidivergence,
  title = {Human--AI Divergence in Ego-Centric Action
           Recognition under Spatial and Spatiotemporal
           Manipulations},
  author = {Sadegh Rahmaniboldaji and Filip Rybansky and
            Quoc C. Vuong and Anya C. Hurlbert and
            Frank Guerin and Andrew Gilbert},
  year = {2026},
  eprint = {2603.08317},
  archivePrefix = {arXiv},
  primaryClass = {cs.CV},
  url = {https://arxiv.org/abs/2603.08317}
}