Skip to benchmark content
MotionBenchHosted by Baz Studio
Pilot datasetOne prompt · one trial · 33 routes · not the official Core rankingRead the limits
Open, reproducible evidence · hosted by Baz Studio

One brief.
Every model.
Full playback.

MotionBench measures whether AI models can direct and build editable motion graphics—not merely describe them. Watch every result, inspect every receipt, and decide blind.

FROZEN INPUT01 creative briefsame task · same canvas · same evidence gates
01GPT 5.4 MiniopenaiB42.0s02GPT-5.5 LowopenaiA68.0s03GPT-5.6 LunaopenaiA81.2s
OUTPUTPreference + completion + cost + time
Model routes33
Playable MP4s32
Exact completions31
Clean receipts25
Raw model COGS$14.22
Blind comparison preview

Do not read the label.
Watch the motion.

Pairwise judgment happens before identities are revealed. Full playback matters more than a flattering poster frame.

Preparing a matched comparison…
Evidence, not vibes

Every rank must survive
the receipt.

01

Playable output

The final MP4 is the primary artifact. Posters never substitute for complete playback.

02

Technical completion

Agent terminal state, scene mutations, validation, runtime, exact timeline, and export all have to reconcile.

03

Raw economics

Provider cost and Bazaar price are separated. No hidden margin is presented as model cost.

04

Frozen provenance

Prompt, model route, runner, scorer, and methodology versions live with each immutable snapshot.

Open the black box

Run the same system.
Bring your own brief.

Community comparisons use the same receipt format but never alter the official ranking.

Open Community Lab