Outcome oriented
Judge the completed state.
The full sequence of intermediate actions is optional. Visible evidence of the intended result is not.
Research preview · Under review
Can a video generator actually finish the task—not merely make a convincing video? SemComp-Bench evaluates outcome achievement together with task-relevant semantic grounding.
Reference frame
Completed outcome
01 · The problem
Existing evaluation often emphasizes fidelity, motion, or prompt alignment. Semantic task completion asks a stricter question: did the generated outcome happen, and does it remain meaningfully connected to the reference?
Outcome oriented
The full sequence of intermediate actions is optional. Visible evidence of the intended result is not.
Semantically grounded
Identity, material, appearance, layout, or scene attributes are preserved only when they matter to the task.
02 · SemComp-Data
Every instance links a reference frame, brief and detailed instructions, and an outcome-centric clip from the same source video—keeping the target feasible and visually verifiable.
03 · Curation pipeline
Remove narration-dependent content and categorize visually self-contained tasks.
Localize and conservatively verify reference–outcome frame pairs.
Build a compact clip around the grounded outcome timestamp.
Produce aligned brief and detailed instructions with explicit grounding constraints.
04 · SemComp-Bench
Structured binary VLM judgments separate task success from rendering reliability and provide criterion-level evidence.
Outcome achievement
Generation reliability
05 · Results
The best OA score remains below 40%, while the strongest GR score exceeds 90%. Visually reliable generation does not guarantee that the instructed outcome was achieved.
HunyuanVideo-1.5-720P-I2V
Seedance 2.0
Within-scene coherence pass rate
SemComp-Core · Detailed instructions
Toggle the metric to compare outcome achievement with generation reliability.
Takeaway
“Looking right” and “getting it done” are different capabilities.
SemComp-Bench makes that gap measurable through authentic reference–outcome pairs and evidence-grounded evaluation.
06 · Citation
The current manuscript is anonymous. Replace the author field and enable public resource links before publishing.
@article{semcompbench,
title={SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation},
author={Anonymous},
}