Benchmark Lab Methodology

How We Score Video Coherence: Temporal Consistency and Motion Quality Explained

Tom Hewson  · 

Abstract visualization of video frame consistency analysis across time

Video generation models need a different scoring framework than image models. We describe the three-axis methodology behind our video coherence leaderboard.

When we added video generation to our benchmark platform, we initially tried to adapt our image scoring pipeline. This seemed reasonable: video is a sequence of images, and our image scoring covered perceptual quality, diversity, and prompt adherence. Apply it frame-by-frame, aggregate over time, done.

That approach failed quickly. The unique failure modes of video generation are temporal, and they are invisible to a frame-level image scorer. A video can contain individually high-quality frames and still be completely unusable because objects teleport between frames, lighting changes discontinuously, or motion blur is applied inconsistently. We had to build a different scoring architecture from the ground up.

What video coherence actually means

Coherence is the right word for what distinguishes a usable video from a sequence of related images. A coherent video maintains consistency of scene elements, lighting, and physics across the time dimension. An incoherent video might look fine if you freeze any single frame but feels wrong when played because the temporal relationships break down.

We identify three measurable axes of coherence: temporal consistency, motion quality, and frame coherence. These sound related but they capture distinct things. You can fail on motion quality while passing on frame coherence. You can pass on temporal consistency while failing on motion quality. The composite video score is a weighted average of all three, with weights tuned based on our calibration against human preference judgments.

Temporal consistency

Temporal consistency measures whether scene elements that should persist across frames actually persist. We are measuring whether a specific object present at frame 1 is still recognizably the same object at frames 16, 32, and 64 in a 64-frame sequence. Identity flickering (where an object briefly ceases to exist or changes fundamental attributes between frames) and geometric drift (where the spatial relationship between objects shifts without an explained camera move or motion event) are both failures of temporal consistency.

We measure this using a combination of object tracking and feature matching. Our pipeline extracts a set of persistent salient regions from the first frame using a feature detector, tracks those regions through the subsequent frames, and measures how much each region's feature representation shifts between adjacent frames. Large feature shifts that do not correspond to expected motion indicate consistency failures.

The challenging case is motion itself: a feature shift caused by genuine expected motion (a hand reaching into frame, a camera pan) should not be penalized as an inconsistency. We handle this by maintaining a motion model per video that distinguishes expected feature drift from discontinuous jumps. A smooth, continuous feature trajectory scores as consistent regardless of total displacement. A discontinuous jump, even a small one, scores as inconsistent.

Motion quality

Motion quality is about whether the motion in the video looks physically plausible. This is distinct from whether the scene is consistent. A video can have perfectly consistent scene elements while containing motion that looks unnatural: objects that accelerate or decelerate in unrealistic patterns, character limbs that bend at non-anatomical angles, or camera movements that violate basic optics.

We score motion quality on two sub-axes: motion smoothness (does motion follow continuous, physically plausible trajectories?) and motion naturalness (does the motion pattern match known physical dynamics for the type of motion depicted?). Smoothness is measurable purely from the optical flow field: measure the second derivative of the flow (jerk), and high jerk indicates unnatural motion. Naturalness is harder and requires some domain-specific calibration. We have physics priors for common motion types (falling objects, water, fire, human locomotion) and check generated videos against those priors for the applicable category.

Motion quality is the dimension where current video generation models show the widest variance. Some models produce very smooth motion with poor naturalness (the motion is consistent but looks wrong). Others produce high-naturalness motion with some smoothness failures (the physics is right but the animation is jerky). Getting both right simultaneously is the hard problem.

Frame coherence

Frame coherence is the closest to what our image scorer measures, applied at the video level. We are checking whether individual frames are high quality and whether that quality is maintained consistently throughout the video, not just at cherry-picked frames.

The most common failure mode here is quality degradation in later frames. Video generation models sometimes produce high-quality early frames and increasingly artifact-heavy later frames as the generation context extends. A model that scores 80 on our image quality metric for frames 1-16 and 55 for frames 48-64 of a 64-frame video has strong frame coherence variance, and we penalize that variance even if the average across all frames is acceptable. An unreliable model is not useful for production workflows even if its average quality looks reasonable.

We also check within-frame quality distributions, specifically checking whether different spatial regions of the same frame show dramatically different artifact levels. Background regions that degrade significantly while subject regions remain clean indicate a compositional quality inconsistency that real-world uses will notice.

Our test prompt set for video evaluation

We evaluate on 200 prompts per full evaluation cycle. The prompt set is stratified across four motion categories: static scene with lighting change, camera movement through a scene, subject motion in a fixed camera, and complex scenes with both camera and subject motion. We maintain this stratification because models differ substantially in how they handle different motion types, and a prompt set that over-represents any one category would produce rankings that only hold for that category.

Prompt length and specificity also vary in our set. We include both highly specific prompts ("a ceramic mug on a wooden table, morning sunlight from the left side, slow zoom in over 4 seconds") and abstract prompts ("calm, meditative atmosphere, slow movement"). Specific prompts test whether models can follow instructions for video; abstract prompts test whether models can generate plausible motion without explicit direction. Models rank differently on these two subtypes, and we report the breakdown in our API response.

Current limitations of automated video scoring

Our scoring pipeline is automated throughout. We do not use human evaluators for the component scores, only for calibration of our weights and validation of our automated measures. This is a practical constraint: running human evaluation at weekly cadence across 200 prompts and multiple models is not feasible for a small team. The tradeoff is that our automated scores may miss failure modes that humans would notice immediately but that are hard to operationalize as a measurement.

The clearest current gap is in semantic coherence: whether the narrative or scene logic of the video makes sense. A video of someone "pouring water from a jug into a glass" might score perfectly on temporal consistency, motion quality, and frame coherence while having the water flowing upward into the jug rather than downward. Our current pipeline would not catch that. We are working on a semantic consistency check using a vision-language model to evaluate the described-versus-generated scene for a subset of prompts, but this is not yet in production scoring. We note this limitation explicitly in the leaderboard methodology for video evaluation.

See the video coherence rankings

Current leaderboard scores across temporal consistency, motion quality, and prompt alignment.

View Video Leaderboard