Skip to benchmark content
MotionBenchHosted by baz.studio
September 6, 2026 · Local September engine

One brief.Four models.

Each model got the same frozen prompt and brand pack, then directed and built its own result. One trial per model. Unranked, with no blind votes collected. This diagnostic showcase is separate from the published July dataset and is a subset of the September route panel.

All four exports passed encoded checks. A valid export and a completed agent workflow are separate outcomes. Costs below are token-based estimates of raw provider spend. Invoice totals and billed account charges are separate measures. CLI generation time excludes export.

Full video

10s · 300 frames · 30fps · 1920×1080 · silent

GPT-6 Astra

Upstream model · gpt-6-astra
Agent workflow
Complete
Encoded checks
Passed
Provider cost estimate
$2.6041
CLI generation time
358.365 s

AI review noteReadable typography connects an animated preview with editable React code. Abstract motion and long-held layouts limit motion contrast.

Full video

10s · 300 frames · 30fps · 1920×1080 · silent

Claude Fable 5.1

Upstream model · claude-fable-5-1
Agent workflow
Complete
Encoded checks
Passed
Provider cost estimate
$3.9595
CLI generation time
421.481 s

AI review notePreview-to-code connections show editable scenes. A sparse opening and blank transition interrupt the flow. The displayed baz create command is unsupported.

Full video

10s · 300 frames · 30fps · 1920×1080 · silent

Gemini 3.7 Flash

Upstream model · gemini-3.7-flash
Agent workflow
Incomplete
Encoded checks
Passed
Provider cost estimate
$0.3737
CLI generation time
236.415 s

AI review noteReadable headlines and consistent styling have little visual development. The video includes unsupported statistics, speed claims, technical labels, and a baz build command.

Full video

10s · 300 frames · 30fps · 1920×1080 · silent

Gemini 3.8 Flash

Upstream model · gemini-3.8-flash
Agent workflow
Incomplete
Encoded checks
Passed
Provider cost estimate
$0.3179
CLI generation time
272.961 s

AI review noteReadable headlines sit in three held interface layouts. Small technical text and limited motion contrast weaken the result. Simulated editor details include unsupported claims.

Review notes summarize AI assessments of sampled frames. Unsupported commands, statistics, and editor claims inside these model-generated videos are output errors; they are not product facts about baz.studio.

The frozen task

motionbench-agentic-v1 · Same candidate model in both the director and code-generation roles.

Frozen input hashes
Prompt SHA-256
a9b7cda9924763aee2044b037aafa9aa4f2e0982c919cb6b43c3700186186b0b
Brand pack SHA-256
75f639a1a443b4f8cc4b9e3520393d0ce310674f81c39daac6b8b5aa6abfa157
Make a 10-second launch video for baz.studio. Decide the strongest product story, create your own creative brief and execution plan, then build, review, and deliver the finished video. You are directing this - the available patterns and defaults are a toolbox, not a script; use, combine, or ignore them as your own creative judgment requires. No audio.

Inspect the output. Make your own.