Skip to benchmark content
MotionBenchHosted by Baz Studio
Methodology v1.0 draft

A ranking is only as useful
as its failure policy.

MotionBench separates creative preference from technical completion. A beautiful broken result is evidence, but it is not an eligible official sample.

01

Object under test

The benchmark evaluates a named model route inside a frozen Baz generation pipeline. The route may control orchestration, code generation, or both. Every run records the real upstream model identifier and its role. It does not claim to measure a provider in the abstract.

02

The Core suite

The official leaderboard uses three task families—identity motion, interface narrative, and data story—across landscape and portrait. Each route receives repeated trials. Prompts, runner version, assets, tool policy, and acceptance criteria are frozen per snapshot.

IdentityLogo, rhythm, lockup
InterfaceHierarchy, continuity, UI motion
Data storyAccuracy, pacing, legibility
03

Eligibility gates

Agent completion

The agent must end cleanly; a transport-level 200 response is not success.

Real mutation

At least one persisted scene mutation and one persisted scene are required.

Validation

Project validation and task-specific requirement checks must both pass.

Runtime

No visible Scene Error sentinel or equivalent runtime failure may appear.

Exact timeline

Requested frames, fps, duration, dimensions, and audio policy must reconcile within published tolerance.

Export

A playable MP4 and its export receipt must exist.

Cost receipt

Provider raw cost, token usage, and billed price must reconcile or be explicitly marked unavailable.

04

Scores

The headline creative score comes from blind pairwise human preferences fit with a regularized Bradley–Terry model and reported with sample count and uncertainty. Technical completion is reported separately. “Both miss” is not silently converted into a tie.

05

Cost and latency

Latency uses wall-clock milestones: queue, generation, validation/repair, export, and total. Raw provider cost is separated from Baz pricing. Medians and tail latency are shown when repeat count supports them.

06

The July pilot is not Core

The current public dataset is one portrait prompt with one trial per route. It exposed important system failures—including false-success signaling, visible runtime errors, duration mismatch, and an instance-local 120 requests/minute limiter—but cannot support a stable global ranking. It is published as a reproducible diagnostic.

Inspect the pilot →