Ego2Act

Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Patrick Amadeus Irawan1* Iskandar Muda Rizky Parlambang2† Rava Maulana2† Qinrong Cui1 Erland Hilman Fuadi1 Zayd M. K. Zuhri1 Nanda Ryaas Absar1 Ahmed Elshabrawy1 Wilfried Ariel Mulyawan3 Shoubin Yu4‡ Yue Zhang4‡ Mohit Bansal4‡ Alham Fikri Aji1*

1Mohamed bin Zayed University of Artificial Intelligence 2Independent Researcher 3Nanyang Technological University 4University of North Carolina at Chapel Hill

*Lead and corresponding authors   †Major contributors   ‡Equal advising

“Given starting environment scene and a goal, can video model demonstrate a series of manipulations to get the job done?”

Same start image and goal for a correct human recording and six video models. Each number is the model’s overall human-rated Final score, and each video is one example generation.

TL;DR We record 110 everyday manipulation tasks and ask six video generators to finish each one from a single start frame. Even the best model, Seedance-2.0, completes only 67.8% of the task and scores 64.0 Final under human rating, and 88.7% of failed steps happen because the model skips a step or only partly carries it out, which leaves later steps without the state they depend on. Ego2ActJudge scores these videos without a reference and agrees with human ratings (r = 0.69) more closely than existing evaluators.

Full Abstract

Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocentric video generation that requires planning to simulate proper execution to accomplish high-level goals by carrying out multiple real-world manipulations. We introduce Ego2Act, a goal-directed benchmark featuring 2,640 videos from 110 real-world tasks across day-to-day settings, varying object clutter and multi-step complexity. Given an initial scene image and a high-level goal, Ego2Act evaluates whether video generation models can produce realistic egocentric videos of a hand manipulating objects to carry out the task. To support scalable evaluation, we also introduce Ego2ActJudge, a reference-free evaluation pipeline that achieves better task completion and physics plausibility evaluation alignment with human consensus compared to relevant baselines. Our findings reveal that models’ generated simulations often skip or partially execute steps, leaving later steps missing dependent states, which leads to unfulfilled goal. Furthermore, models consistently fail at fine-grained physical dynamics, particularly during complex object manipulation and persistent world modeling. We hope Ego2Act provides a rigorous testbed for advancing video models toward physically plausible, goal-directed simulation.

The Ego2Act Benchmark

Every case starts from one real scene and one everyday goal, and grows into a full set of human and generated videos that Ego2ActJudge scores.

  1. 110Start image and goalreal egocentric tasks in everyday settings
  2. 660Human recordings3 correct and 3 wrong attempts per task
  3. 1,980Generated videos6 generators with 3 seeds each
  4. 2,640Videos scoredEgo2ActJudge gives each one a Task, Physics and Final score

Benchmark Statistics

Figure 2 showing the composition and descriptive statistics of Ego2Act
Figure 2. Composition and descriptive statistics of Ego2Act. (a) Domain composition of the 110-case benchmark. Panels (b–e) summarize the feature distributions, namely (b) action-family co-occurrence and coverage, (c) human control videos duration, (d) involved solid-object counts, and (e) case-mean observed operations. Co-occurrence means joint presence within a case, and observed operations describe number of subaction per case. (f) unique action family count per case.

Evaluation Results

How the six models rank

Six image-to-video generators, ranked separately by human raters and by Ego2ActJudge.

Human raters* Manual

Final score, 0–100

  • #1ByteDance logoSeedance-2.064.0
  • #2Kuaishou logoKling-v3-Pro59.6
  • #3MiniMax logoMiniMax-H355.3
  • #4xAI logoGrok-1.549.9
  • #5Alibaba logoWan-2.744.3
  • #6NVIDIA logoCosmos-3 Nano3.9

Ego2ActJudge Auto

Final score, 0–100, all 110 cases

  • #1ByteDance logoSeedance-2.083.3
  • #2Kuaishou logoKling-v3-Pro72.5
  • #3Alibaba logoWan-2.771.4
  • #4xAI logoGrok-1.569.5
  • #5MiniMax logoMiniMax-H359.3
  • #6NVIDIA logoCosmos-3 Nano13.3

*Human ratings cover a 25-case subset (600 videos).

How far do generators get?

Not very far. Human raters give the best model, Seedance-2.0, a Final score of 64.0, and Cosmos-3 Nano only 3.9. The map shows, per generator and action type, the share of subgoals Ego2ActJudge rates complete (Task) or valid (Physics); hover a cell for its count.

OverallRelocationContainmentOpeningAttachmentShapeMaterial transferSurface treatmentCuttingDevice activation
Seedance-2.07272707258828073100100
Kling-v3-Pro54564255475041697180
Wan-2.758614554496353737133
Grok-1.5586246534057587688100
MiniMax-H344463952295024543867
Cosmos-3 Nano1081021620221125

Leaderboard

GeneratorHuman raters (25-case panel)Ego2ActJudge (Full cases)
TaskPhysicsFinalTaskPhysicsFinal
Human (correct)100.0100.0100.093.998.195.4
Human (wrong)59.694.173.457.891.670.1
Seedance-2.067.863.064.079.991.183.3
Kling-v3-Pro59.364.259.664.088.972.5
Wan-2.755.541.444.367.880.571.4
Grok-1.556.647.949.967.879.469.5
MiniMax-H361.255.855.353.776.259.3
Cosmos-3 Nano16.95.23.915.438.813.3

Human (correct) and Human (wrong) are human recordings of correct and incorrect executions; Human (correct) recordings are verified successful during dataset construction and receive the rubric maximum under human rating. Cosmos-3 Nano is the smaller Nano variant of Cosmos 3, run locally, so its scores describe that variant only.

Models and settings

Table. The six video models, how each was run and the exact version used. All models received the same prompt template, 480p output and seeds 101, 202 and 303.

ModelAccessVersionNotes
Seedance-2.0OpenRouterbytedance/seedance-2.016:9, up to 15 s
Kling-v3-ProOpenRouterkwaivgi/kling-v3.0-pro16:9, up to 15 s
Wan-2.7OpenRouteralibaba/wan-2.716:9, up to 10 s, its maximum
Grok-1.5OpenRouterx-ai/grok-imagine-video-1.516:9, up to 15 s
MiniMax-H3Local, SGLangMiniMaxAI/MiniMax-H34:3 native, up to 15 s, 24 fps, 35 steps
Cosmos-3 NanoLocal, SGLangnvidia/Cosmos3-Nano16:9, up to 15 s, 24 fps, 35 steps, guidance 4.0, flow shift 10.0

Local models ran on NVIDIA A100 hardware. Cosmos-3 Nano is the Nano variant of Cosmos 3, so its results describe that variant only.

Full cross-benchmark leaderboard (Table 2)

Table 2. Cross-Benchmark Performance Leaderboard on Ego2Act. Sub-component Task / Physics splits are reported alongside average of per video geometric mean or benchmark official score, with human controls excluded from best (bold) and second-best (underlined) model rankings.

Gen. / Eval.Ego2ActJudge OursHuman JudgeWRAPQSGRBenchWMBenchVideoScore
T / PAgg.T / PAgg.Ori.T / POri.T / POri.T / POri.T / POri.
Human (correct)93.9 / 98.195.4100.0 / 100.0100.095.695.5 / 98.697.376.7 / 80.475.763.4 / 78.776.283.3 / 84.387.5
Human (wrong)57.8 / 91.670.159.6 / 94.173.459.761.7 / 80.676.247.4 / 73.843.658.5 / 80.575.883.9 / 85.388.2
Grok-1.567.8 / 79.469.556.6 / 47.949.978.782.5 / 95.188.667.4 / 78.264.373.7 / 73.175.988.2 / 87.990.6
Seedance-2.079.9 / 91.183.367.8 / 63.064.088.690.2 / 97.393.874.1 / 80.372.469.7 / 73.074.789.3 / 89.892.2
Kling-v3-Pro64.0 / 88.972.559.3 / 64.259.677.077.1 / 94.286.465.0 / 78.961.967.3 / 74.174.690.5 / 91.293.4
Wan-2.767.8 / 80.571.455.5 / 41.444.380.580.7 / 93.488.367.1 / 76.963.875.1 / 75.978.192.2 / 94.095.8
MiniMax-H353.7 / 76.259.361.2 / 55.855.368.669.8 / 87.280.651.7 / 66.348.175.2 / 81.381.495.8 / 97.198.2
Cosmos-3 Nano15.4 / 38.813.316.9 / 5.23.935.228.0 / 45.845.223.7 / 37.221.049.1 / 78.871.384.1 / 93.192.8

Human Judge scores come from the 600-video alignment panel. Human (correct) and Human (wrong) are human recordings of correct and incorrect task executions; Human (correct) recordings are verified as successful executions during dataset construction and receive the rubric maximum.

Example Cases

Pick a case to compare a correct human recording with one generation (seed 101) from each model, all given the same start image and goal. Badges are Ego2ActJudge Final scores (0–100): ≥70 60–69 30–59 <30.

Ego2ActJudgeGoal + generated videono reference video neededSubgoal planwith prerequisite orderTASK COMPLETIONPHYSICS PLAUSIBILITYT1 · Initiationthe action startsT2 · Processcarried out correctlyT3 · End stategoal state reachedP1 · Continuityobjects stay the sameP2 · Causationmotion has a causeP3 · Interactioncontact is plausibleP4 · Persistenceresult stays putT = mean level / 3P = mean level / 4Final S = √(T · P)both must hold

Agrees with human raters at r = 0.69, better than every baseline, and ranks the six models almost the same way (ρ = 0.94).

Table 3(b). Ego2ActJudge alignment with human consensus on the Final score. Best in bold, second best underlined.

MetricWRA†PQSGRBenchWMBVScoreSimple VQAEgo2ActJudge
r ↑0.420.310.400.16-0.100.320.69
τ ↑0.320.200.280.11-0.150.210.46
CCC ↑0.340.160.390.08-0.020.220.61

† WR-Arena scores the Task axis only.

More analysis
Human ratingEgo2ActJudge
Mean Task and Physics score per model: human rating versus Ego2ActJudge002020404060608080100100TaskPhysicsAll videosAutomated judgeupward bias● human rating → ● Ego2ActJudge
The judge is more lenient on physics than humans. Faint dots are single videos, big dots are model averages (amber = human raters, blue = Ego2ActJudge). The blue judge averages sit clearly above the amber human ones, since the judge adds only +8.5 on Task but +24.0 on Physics (53.2 → 77.2), so it catches missed steps well but overlooks physical errors that humans notice. 423 videos rated by both.
HumanEgo2ActJudge
Worst seed, average and best-of-3 Final score per model0255075100Seedance-2.0Kling-v3-ProMiniMax-H3Grok-1.5Wan-2.7Cosmos-3 Nano
3-seed stability. Final score per model, drawn from the worst seed through the average (dot) to the best of 3 seeds and sorted by the human average.
Rubric ablation: average MAE per variant, lower is better
Rubric ablation. Average MAE under rubric ablation; lower is better. The full rubric scores gates sequentially, stopping at the first failure, whereas Independent scores all gates.

Common Failure Patterns

Most failed steps are simply skipped or left half-done (88.7% of failed Task subgoals). Attaching things (38.6% complete) and transferring material (41.9%) are the hardest. Below are representative examples, grouped by error type, and the ones with annotated frames come from the paper's appendix.

Click a failure type to see representative examples of common errors

Task failures

Skipped prerequisite operations 5 examples

A necessary enabling action is absent, while a dependent operation proceeds.

  • NVIDIACosmos-3 Nano67

    coffee beans seed 202

    Some coffee beans transferred from inside transparent jar into inside orange jar, with no beans spilled outside either jar.

    Why it failed Beans are poured onto the still-seated orange lid before an opening is established.

  • xAIGrok-1.50

    three cups and water seed 202

    Only one of the visible colored cups contains some water, the water in that cup is stirred, and there is no visible spill outside the cups.

    Why it failed Stirring is attempted on the bases of inverted mugs without first turning them upright.

  • KuaishouKling-v3-Pro0

    air freshener dispenser seed 101

    The green air freshener is installed in the automatic dispenser, and the white air freshener and green cap are inside the green tote bag.

    Why it failed The turquoise can is lifted and inserted into the raised cover while its green cap remains attached, and the old white can remains in the actuator bay.

  • ByteDanceSeedance-2.076

    staple paper seed 101

    The round yellow paper is stapled to the white A4 paper with the fully inserted stapler with its staples filling.

    Why it failed The stapler is pressed with its base resting on top of the paper, without feeding the papers between the base and staple outlet.

  • AlibabaWan-2.758

    oil spray pouch seed 202

    All fully closed spray and closed toiletries are inside the pink pouch.

    Why it failed The bottle is put into the pouch with its cap still open, so the closing step it depends on is skipped.

Incomplete outcomes 5 examples

The relevant operation is attempted, but the required final state or object relation is not established.

  • KuaishouKling-v3-Pro59

    move juice seed 101

    Move all of the orange juice from a plastic bottle to a reusable bottle by first emptying the reusable bottle that is filled with water into the sink.

    Why it failed The reusable bottle is not fully emptied first, and the orange juice is only partly transferred.

  • KuaishouKling-v3-Pro18

    bottle cap seed 101

    The turquoise container, black bottle, and gray water bottle are all upright on the table and closed with their matching lids: the turquoise lid on the turquoise container, the black-and-white lid on the black bottle, and the small gray cap on the gray water bottle.

    Why it failed The turquoise lid is put on the red cup rather than the turquoise container.

  • MiniMaxMiniMax-H363

    toothbrush case seed 101

    Closed toothpaste and folded toothbrush stored inside the closed travel case.

    Why it failed Unfolded toothbrush protrudes from the closed case.

  • AlibabaWan-2.737

    shoes socks arrange seed 101

    Blue socks placed inside the blue shoes and grey socks placed inside the black sandals, with the blue shoes positioned to the right of the black sandals

    Why it failed The grey sock is never placed in the black sandal, and the blue sock stays in it.

  • ByteDanceSeedance-2.091

    desk mat tablet seed 303

    Rolled-up desk mat to the left of the tablet, with the book on top of the tablet.

    Why it failed The mat remains partly spread beneath the tablet/book after rolling is attempted.

Physics failures · state

World-state inconsistency 5 examples

Object identity, appearance, material, geometry, or assembly state is not preserved across frames.

  • NVIDIACosmos-3 Nano0

    eyeshadow palette seed 303

    All makeup items properly closed, with the palette stored inside the white pouch, and all items neatly stored inside the large black pouch on the left.

    Why it failed The palette interior changes into an orange cylindrical object while the rectangular housing remains visible.

  • xAIGrok-1.5100

    kendama can seed 202

    Kendama ball properly attached to the standing kendama stick that was originally laid down, with the yellow can stacked on top of the other can.

    Why it failed The yellow can acquires a green Sprite label as it is raised onto the other can.

  • KuaishouKling-v3-Pro0

    coffee beans seed 202

    Some coffee beans transferred from inside transparent jar into inside orange jar, with no beans spilled outside either jar.

    Why it failed Holes appear in the previously solid white and orange lids during spoon contact.

  • MiniMaxMiniMax-H358

    screwdriver wrench seed 202

    All of appliances are inside the closed green container, with the orange glove placed on top of the closed container.

    Why it failed The hammer being held changes into orange gloves over the closed case.

  • ByteDanceSeedance-2.0–

    dock switch game seed 202

    Both Joy-Cons are attached to the Nintendo Switch, the console is placed in the Switch dock, and the game card is inserted with the game card slot cover closed.

    Why it failed The fixed console screen changes into a hinged flap.

Object-count inconsistency 5 examples

The scene inventory changes without a corresponding addition, removal, division, or assembly event.

  • xAIGrok-1.50

    router laptop cable seed 202

    Gray LAN cable plugged into the black laptop, with the orange cable disconnected from the JBL speaker and laptop and connected instead to the black charger block. The closed laptop, router, JBL speaker, and charger block are stacked from bottom to top in that order.

    Why it failed Two JBL speakers coexist although the initial scene contains one.

  • KuaishouKling-v3-Pro87

    lava toast microwave seed 101

    The lava toast is placed on the plate with the chocolate on top, and both are inside the microwave.

    Why it failed Toast appears on the plate while another toast remains in the original tray.

  • ByteDanceSeedance-2.00

    three caps seed 303

    Each marker is fitted with its matching color cap.

    Why it failed A new marker suddenly appears on the table.

  • MiniMaxMiniMax-H367

    figurine toy box seed 202

    Toy cars inside the closed green box, with the AC remote and bowl on top of the box and the cat figurine on top of the bowl.

    Why it failed While the toy car is moved, the small bowl next to it disappears.

  • NVIDIACosmos-3 Nano29

    makeup pouch seed 303

    The cotton is moved from its pack to be inside the clear container, and the closed cyan lipstick is inside the pouch.

    Why it failed A cyan lipstick appears out of nowhere.

Physics failures · interaction

Causal and boundary violations 5 examples

Motion lacks an appropriate trigger, or an object crosses a closed or otherwise inaccessible boundary.

  • NVIDIACosmos-3 Nano42

    vim lunchbox seed 202

    Prepare a closed and latched clear lunchbox containing the wrapped snack and capped bottle (at first cap separated).

    Why it failed Bottle protrudes through the closed lunchbox lid.

  • KuaishouKling-v3-Pro65

    paper binder book seed 101

    Blank loose-leaf paper attached to the ring binder, with the scissors inside the closed red book resting on top of the closed ring binder.

    Why it failed The scissors pass through the closed binder.

  • MiniMaxMiniMax-H327

    light bulb glasses seed 202

    Detached spiral light bulb stored inside the closed red trash can, with the glasses still placed on top of the trash can.

    Why it failed The bulb is lowered through the closed red bin lid rather than through an opened mouth.

  • ByteDanceSeedance-2.0–

    figurine toy box seed 101

    Toy cars inside the closed green box, with the AC remote and bowl on top of the box and the cat figurine on top of the bowl.

    Why it failed The lid closes through the upright cat figurine, leaving it protruding above the continuous closed lid.

  • AlibabaWan-2.782

    screwdriver wrench seed 202

    All of appliances are inside the closed green container, with the orange glove placed on top of the closed container.

    Why it failed The glove is pulled through the closed solid lid without opening the case.

Manipulation and coupling failures 5 examples

Contact-intensive manipulation violates plausible attachment, separation, or material response.

  • NVIDIACosmos-3 Nano20

    blue pouch eyedrops seed 101

    The eye drops is in the blue pouch, then the e-reader and blue pouch that contains the eyedrop are put inside the black bag, and that black bag is finally closed.

    Why it failed While the fingers pull at the e-reader top edge, its rigid casing bends and opens into a soft pouch-like mouth containing loose items.

  • NVIDIACosmos-3 Nano0

    laptop setup seed 303

    The laptop is opened and connected to power, with the mouse to its right and the headphones in front of its keyboard side.

    Why it failed As the hand lifts the headphones from the closed laptop, the rigid lid bows and its edge buckles before resolving into the open display.

  • MiniMaxMiniMax-H333

    cooking oil seed 303

    Cup containing one tablespoon of cooking oil.

    Why it failed The printed solid base of the inverted ceramic mug is lifted off like a removable lid.

  • MiniMaxMiniMax-H316

    make coffee seed 202

    The empty cup filled with instant coffee and water, with the mixture stirred using the spoon.

    Why it failed The solid base of the inverted ceramic mug is pulled off like a removable lid.

  • MiniMaxMiniMax-H358

    pour milk seed 101

    The cup contains some milk from the sealed milk box.

    Why it failed The hand pinches the printed solid bottom of the inverted mug and lifts it away as a separate disc, opening the ceramic body without fracture.

Mechanism violations 5 examples

Articulated or constrained components move in ways that are incompatible with their structure.

  • NVIDIACosmos-3 Nano62

    stack book laptop seed 202

    The closed book is stacked on top of the closed laptop.

    Why it failed The laptop display folds down with its active screen facing outward on top rather than inward toward the keyboard.

  • KuaishouKling-v3-Pro58

    laptop open seed 101

    Laptop lid open, with the key and pen placed together on the left side of the laptop.

    Why it failed During opening the screen rotates diagonally and ends attached along the right side of the keyboard deck, whose trackpad remains at the front.

  • KuaishouKling-v3-Pro71

    laptop setup seed 202

    The laptop is opened and connected to power, with the mouse to its right and the headphones in front of its keyboard side.

    Why it failed The front notch of the closed laptop faces the viewer, and its right edge then rises around the left edge while the base footprint remains in place and the keyboard orientation changes.

  • MiniMaxMiniMax-H382

    four in a row seed 101

    Put the final move of the game to win as either side

    Why it failed A disc is inserted through a front grid aperture rather than via the board top-entry channels.

  • ByteDanceSeedance-2.065

    matryoshka battery seed 303

    The batteries are inside the closed matryoshka doll; the battery cover is attached to the AC remote; and the doll is on top of the remote.

    Why it failed The doll head pivots open on an invented hinge rather than separating at the nesting-shell joint.

Summary

  1. humanmodel64.0best human-rated Final

    A large gap to real execution

    Even Seedance-2.0 reaches only 64.0 under human rating, and Cosmos-3 Nano just 3.9.

  2. 88.7% of failed steps

    Steps get skipped

    Most failures are omissions, where the needed step is skipped or left half-done, so later steps have nothing to build on.

  3. 38.6% attachment done

    Contact-rich changes are hardest

    Attaching things (38.6%) and transferring material (41.9%) have the lowest completion of any action type.

  4. 0.69r with humans

    The judge tracks humans

    Ego2ActJudge agrees with human consensus far better than the best baseline (RBench, 0.40), close to a single human rater (0.76).

  5. 0.91AUC, success vs failure

    It separates success from failure

    Pick any successful and any unsuccessful human recording and the judge scores the successful one higher 91% of the time (Mann–Whitney p < 10−60; mean 95.4 vs 70.4, 305 vs 295 videos), so it tells success from failure.

  6. 24.0Physics points above humans

    The automated judge has an upward bias

    Compared with human raters, Ego2ActJudge adds only +8.5 on Task but +24.0 on Physics, so it catches missing steps but overlooks subtle physical errors that people notice.

BibTeX

@article{ego2act2026,
  title   = {Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation},
  author  = {Irawan, Patrick Amadeus and Parlambang, Iskandar Muda Rizky and Maulana, Rava and Cui, Qinrong and Fuadi, Erland Hilman and Zuhri, Zayd M. K. and Absar, Nanda Ryaas and Elshabrawy, Ahmed and Mulyawan, Wilfried Ariel and Yu, Shoubin and Zhang, Yue and Bansal, Mohit and Aji, Alham Fikri},
  journal = {arXiv preprint arXiv:2610.01092},
  year    = {2026}
}