ProcObject-10K: Benchmarking Object-Centric
Procedural Understanding in Instructional Videos

Michigan State University

NeurIPS 2026 Evaluations & Datasets Track

Abstract

Procedural activities are fundamentally driven by object state transitions, yet existing instructional video benchmarks remain action-centric and cannot evaluate whether models reason about how objects evolve toward task completion. In this work, we introduce ProcObject-10K, the first benchmark that jointly evaluates object-centric reasoning and temporal evidence grounding in instructional videos, across both egocentric and exocentric views. It comprises 10,522 open-ended VideoQA pairs grounded in 1,799 video clips, spanning 137 tasks across 9 domains and five reasoning types covering preconditions, state evolution, counterfactuals, mistakes, and readiness. Benchmarking 13 leading MLLMs reveals a substantial answering-grounding gap: models produce plausible answers while failing to localize the supporting evidence (mIoU < 45%), exposing their reliance on linguistic priors rather than fine-grained object dynamics.

As a step toward closing this gap, we further provide an object-centric supervised fine-tuning baseline with pseudo object-level supervision and spatial-temporal constraints. Models fine-tuned on ProcObject-10K not only improve on the benchmark itself, but also transfer effectively to other grounded VideoQA and embodied planning tasks. The dataset, annotations, and evaluation toolkit are publicly released to support future research on object-centric procedural understanding.

Five Object-Centric Reasoning Types

Earlier procedural benchmarks mostly ask which action comes next. Every ProcObject-10K question is about an object's state, has an open-ended answer, and comes with the time spans in the video that support that answer.

Frames are from test-set questions. Click a card to open that question in the Explorer.

The Benchmark

A grounded VideoQA task: the model watches a procedural clip, answers an open-ended question about an object, and outputs the time spans that support its answer.

video clip + question about an object → MLLM → {"answer": "…", "evidence": [[t0, t1], …]}

Construction Pipeline

Four-stage data generation pipeline of ProcObject-10K
  1. 1
    Procedural video collection. Egocentric cooking (CaptainCook4D, EgoPER), egocentric assembly (HoloAssist) and exocentric Internet videos (COIN). The first three include real execution errors. Corrupted, near-duplicate and ≥30-min videos are removed.
  2. 2
    Action-sequence sampling. A sliding window over N ∈ {2, 3, 4, 5} consecutive action segments yields clips with coherent object state changes; Qwen3.5 writes dense, object-centric captions for each.
  3. 3
    Grounded QA generation. Qwen3.5, with a prompt and question templates for each type, turns the captions and sparse frames into a question, an answer and supporting time spans.
  4. 4
    Verification & filtering. Questions that an LLM answers well from the text alone are dropped (about 30%, ∼4,500). GPT-4o mini checks that the answer and the evidence agree. Two reviewers independently check every QA pair; pairs that fail are corrected by hand.

Statistics

10,522QA pairs
1,799source videos
211 hof video
137 / 9tasks / domains
72.3 smean clip length
1.74evidence spans per QA
9,472 / 1,050train / test, video-disjoint
Videos per domain
QA pairs per type, by temporal-search pattern
Clip length (QA pairs per 5-second bin)
Evidence spans per question

Multi-hop Reasoning: evidence spread over several separate spans. Needle-in-a-Haystack: one short, critical moment inside a long clip. Charts are computed from the released annotations.

Comparison with Existing Datasets

DatasetViewObject-centricQA typeEvidenceMistakesVideo domain
COINExo✗–✗✗Instructional
ChangeItExo✓–✗✗Instructional
EgoSchemaEgo✗Multi-choice✗✗Human activity
EgoPEREgo✗–✗✓Instructional
CaptainCook4DEgo✗–✗✓Instructional
ProMQAExo✗Open-ended✗✓Instructional
REXTIMEExo✗Multi-choice✓✗Generic
VideoInferExo✓Open-ended✓✓Human activity
MultiHop-EgoQAEgo✗Open-ended✓✗Human activity
TrackVerseExo✓–✗✗Generic
EPIC-KITCHENS-100-MQAEgo✗Open-ended✓✓Instructional
ProcObject-10KEgo + Exo✓Open-ended✓✓Instructional

Explore the Benchmark

Ten test questions, two per type, with the ground-truth evidence and the saved benchmark predictions of four models. Play the clip and watch the playhead pass through each model's predicted spans; IoU is computed against the ground truth, exactly as in the benchmark.

0.0 / 0.0 s

Shaded band = ground-truth evidence. Click or drag on the timeline to seek.

Reference answer

GPT-5.4-Mini, question only (no video)

Benchmark Results

13 models on the 1,050-question test split (4 blind LLMs that see only the question, 3 closed-source and 6 open-source MLLMs), plus Qwen3-VL-4B with our object-centric fine-tuning.

1Answers look fine; evidence does not. On the full test split no model reaches 45% mIoU, while the three larger blind LLMs, which never see the video, still score 3.0–3.3 with the judge.
2Needles are harder than multi-hop. Every model localizes a single short moment worse than evidence spread over several spans.

Main Results

AnsweringGrounding (%)
ModelS.↑B.↑J.↑mIoU↑mIoP↑mIoG↑

Metrics. S. = sentence similarity, B. = BERTScore F1, J. = LLM-as-judge score (0–5, the mean of GPT-5-mini, Qwen3 and Llama-3.2 judges over four dimensions: contextual integration, detail orientation, contextual understanding and temporal understanding). Grounding uses set-level, interval-merged IoU between predicted and ground-truth spans; mIoP and mIoG divide by the predicted and the ground-truth length instead. Bold / underline = best / second best, as in the paper (Table 2).

By Question Type

Radar charts of LLM-judge score and evidence IoU for each QA type
Judge score (left) and evidence IoU (right) for each of the five QA types (paper Fig. 4). Mistake and readiness questions, which mostly hinge on one short moment, have the lowest IoU across models. Open full size

An Object-Centric Fine-Tuning Baseline

Standard SFT supervises only the text of the answer, never which objects, or which moments, the answer rests on. We add that supervision cheaply: pseudo object labels from a VLM and an open-vocabulary detector, and two light auxiliary heads that predict where and when the question's objects appear.

Object-centric SFT: pseudo labels supervise a spatial and a temporal head; the heads are removed at inference Pseudo object labels · training split only Qwen3-VL-4B fine-tuning QA pair question + answer Qwen3.5 extracts object phrases “bowl” · “raisins” · “spoon” Grounding DINO boxes + confidence on each input frame Soft patch masks M̃ box overlap per patch, in [0,1]T×P Object presence ỹt 1 if mean box confidence s̄t > τ, else 0 Video T sampled frames Vision encoder P patches per frame LLM + question {"answer", "evidence"} generated text Lgen Spatial head αt,p Temporal head βt Lspl + Ltmp pseudo labels as targets Dashed: training only, removed at inference.
  1. Pseudo object labels (training split only)
  2. QA pair question + answer
  3. ↓
  4. Qwen3.5 extracts object phrases: “bowl”, “raisins”, “spoon”
  5. ↓
  6. Grounding DINO boxes + confidence on each input frame
  7. ↓
  8. Soft patch masks M̃box overlap per patch Object presence ỹt1 if s̄t > τ
  9. Qwen3-VL-4B fine-tuning
  10. Video → vision encoder T frames, P patches each
  11. ↓
  12. LLM + question → {"answer", "evidence"} → Lgen
  13. Training-only branch from the vision encoder
  14. Spatial head αt,p→ Lspl vs M̃ Temporal head βt→ Ltmp vs ỹ
  15. The dashed heads are removed at inference.
A spatial head after the vision encoder predicts, for every patch, whether it overlaps a supporting object (binary cross-entropy against the soft masks M̃, with a per-frame weight wt from the average box confidence). A temporal head predicts whether each frame shows the supporting objects (binary cross-entropy against ỹ). Both heads are dropped after training, so the fine-tuned model is a plain Qwen3-VL checkpoint that generates the answer and its evidence, at no extra inference cost.
L = Lgen + λspl Lspl + λtmp Ltmp

The answer-generation loss stays the main signal. The auxiliary weights are small and warmed up during training, so the constraints shape the visual representation without overwhelming answer generation.

Transfer Beyond the Benchmark

Fine-tuned only on ProcObject-10K, the model is applied zero-shot to another grounded VideoQA benchmark and to embodied planning. ‡ = fine-tuned with the generative loss only, without the spatial and temporal constraints.

Grounded VideoQA · MultiHop-EgoQA

General procedural questions over long egocentric videos, scored with that benchmark's own protocol (LLM score from 0 to 10). Fine-tuning on ProcObject-10K beats GeLM-7B, a model trained on MultiHop-EgoQA itself, on mIoU.

Full table (paper Table 4)
MethodInputmIoPmIoGIoU@0.3mIoUSent. Sim.LLM Score

Embodied Planning · ALFRED

The fine-tuned model replaces the language planner inside FLARE, a framework designed for ALFRED's long-horizon household tasks. As a zero-shot planner it slightly exceeds FLARE with GPT-4 on success rate (41.33 vs 40.88 on unseen scenes).

Full table (paper Table 5)
SR = success rate, GC = goal-condition success (%).
Test seenTest unseen
MethodSRGCSRGC

BibTeX

If you find ProcObject-10K useful, please cite:

@inproceedings{guo2026procobject,
  title     = {{ProcObject-10K}: Benchmarking Object-Centric Procedural Understanding in Instructional Videos},
  author    = {Guo, Wenliang and Kong, Yu},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track},
  year      = {2026},
  url       = {https://arxiv.org/abs/2512.03479}
}

Please also cite the source datasets: CaptainCook4D, EgoPER, HoloAssist and COIN.