Research preview · Under review

SemComp-Bench Benchmarking Semantic Task Completion in Video Generation

Can a video generator actually finish the task—not merely make a convincing video? SemComp-Bench evaluates outcome achievement together with task-relevant semantic grounding.

Anonymous submission

Semantic task Concrete example
Before and after photographs showing a one-dollar banknote folded into an origami turtle Reference frame Completed outcome
Instruction Fold the banknote into a turtle
Same entity preserved Outcome achieved + semantically grounded
1,273
Structured instances
6
Real-world domains
60
SemComp-Core cases
2
Complementary scores

01 · The problem

A polished video is not necessarily a completed task.

Existing evaluation often emphasizes fidelity, motion, or prompt alignment. Semantic task completion asks a stricter question: did the generated outcome happen, and does it remain meaningfully connected to the reference?

A

Outcome oriented

Judge the completed state.

The full sequence of intermediate actions is optional. Visible evidence of the intended result is not.

B

Semantically grounded

Keep what matters.

Identity, material, appearance, layout, or scene attributes are preserved only when they matter to the task.

02 · SemComp-Data

Real tasks, paired from full-context videos.

Every instance links a reference frame, brief and detailed instructions, and an outcome-centric clip from the same source video—keeping the target feasible and visually verifiable.

Arts & Precision Beauty & Fashion Crafts & DIY Food & Cooking Gardening & Pets Sports & Fitness
4.03 sAverage outcome-centric clip
Brief + detailed instruction pair
4Reference alignment types
27Frames sampled for evaluation

03 · Curation pipeline

From raw video to a verifiable evaluation triplet.

  1. 01

    Candidate filtering

    Remove narration-dependent content and categorize visually self-contained tasks.

  2. 02

    State mining

    Localize and conservatively verify reference–outcome frame pairs.

  3. 03

    Video extension

    Build a compact clip around the grounded outcome timestamp.

  4. 04

    Instruction structuring

    Produce aligned brief and detailed instructions with explicit grounding constraints.

04 · SemComp-Bench

Two scores. Nine interpretable checks.

Structured binary VLM judgments separate task success from rendering reliability and provide criterion-level evidence.

OA

Outcome achievement

All four criteria must pass.

  • 01 Outcome realization
  • 02 Semantic grounding
  • 03 Grounded entity consistency
  • 04 Global visual continuity
OA = Aor × Asg × Agec × Agvc
GR

Generation reliability

Five failure modes, averaged.

  • 01 Physical plausibility
  • 02 Visual clarity
  • 03 Artifact-free rendering
  • 04 Spatiotemporal coherence
  • 05 Text & interface integrity
GR = mean(Gpp, Gvc, Gafr, Gwsc, Gti)

05 · Results

Reliability and task completion tell different stories.

The best OA score remains below 40%, while the strongest GR score exceeds 90%. Visually reliable generation does not guarantee that the instructed outcome was achieved.

Best OA37.8%

HunyuanVideo-1.5-720P-I2V

Best GR91.8%

Seedance 2.0

Main bottleneck32.8–73.9%

Within-scene coherence pass rate

SemComp-Core · Detailed instructions

Model comparison

Seedance 2.0
20.0%
Wan2.2-TI2V-5B
23.3%
Wan2.2-I2V-A14B
28.3%
CogVideoX1.5-5B-I2V
14.4%
SkyReels-V2-I2V-14B
22.8%
HunyuanVideo-1.5-I2V
37.8%
Phantom-1.3B
3.9%

Toggle the metric to compare outcome achievement with generation reliability.

Takeaway

“Looking right” and “getting it done” are different capabilities.

SemComp-Bench makes that gap measurable through authentic reference–outcome pairs and evidence-grounded evaluation.

06 · Citation

Prepared for release after review.

The current manuscript is anonymous. Replace the author field and enable public resource links before publishing.

@article{semcompbench,
  title={SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation},
  author={Anonymous},
}