Skip to benchmark content
MotionBenchHosted by baz.studio
September 10-second trials56 of 56 routes listed · 56 attempted · 53 measured · unrankedJuly dataset below

Which AI model makes the best motion graphics?

Motion graphics are code. Each model gets the same 10-second brief. Watch recorded outputs and inspect full test coverage, including attempted, interrupted, and not-run routes.

September 6, 2026 · Local September engine · Same frozen prompt and brand pack. Recorded attempts, completion, encoding, and recovery notes are evidence fields, not a quality ranking.

Sample outputs, not a ranking.

September 10-second trials

Current unranked snapshot. 56 attempted routes, 53 measured time and cost rows, 30 completed routes, and 56 planned model routes.

30complete56attempted53measured56route panel2026-09-07 14:03 UTCupdated

Scroll the chart to compare all measured routes and both columns.

Generation time and token-based provider cost cover the displayed attempt. Costs are before baz.studio markup. Earlier attempts and export costs are separate.

Complete · Incomplete · Infrastructure interrupted

This chart shows only rows with measured time and cost. Interrupted and not-run routes stay visible in the full table.

The 10-second benchmark

The September videos above use the full creative task. The July dataset below used a separate, tightly directed 3-second prompt.

  1. Same inputEvery model starts with the same prompt and frozen brand pack.
  2. The candidate runs the whole buildThe candidate is asked to choose the story, plan the scenes, write the code, and review the result. Completion is recorded separately.
  3. Same rendererOne fixed renderer and export path turns every project into its MP4.
  4. Compare the MP4sInspect every full output, completion result, generation time, and cost estimate.

One brief. Every creative decision is the model’s.

This is the frozen 10-second agentic brief. Each model gets the same brand pack and tools. It chooses the story, motion, and execution.

Make a 10-second launch video for baz.studio. Decide the strongest product story, create your own creative brief and execution plan, then build, review, and deliver the finished video. You are directing this - the available patterns and defaults are a toolbox, not a script; use, combine, or ignore them as your own creative judgment requires. No audio.

July diagnostic pilot

One prompt, one trial per route. The July diagnostic asked for 3 seconds; it is historical evidence, not an official ranking. Some routes paired the candidate with a reference model. Open a run to see the exact roles and original prompt.

Time and raw model cost

Lower-left is faster and cheaper. Every point is named; color shows the sampled visual tier, and an open point failed a technical gate.

Tier ATier BTier CTier FExcluded

Scroll horizontally to inspect every labeled route →

Your next launch could be code, too.

Build an editable motion graphic in baz.studio. Or bring your own prompt to Community Lab and compare two to eight models.