Object under test
The benchmark evaluates a named model route inside a frozen Baz generation pipeline. The route may control orchestration, code generation, or both. Every run records the real upstream model identifier and its role. It does not claim to measure a provider in the abstract.
The Core suite
The official leaderboard uses three task families—identity motion, interface narrative, and data story—across landscape and portrait. Each route receives repeated trials. Prompts, runner version, assets, tool policy, and acceptance criteria are frozen per snapshot.
Eligibility gates
The agent must end cleanly; a transport-level 200 response is not success.
At least one persisted scene mutation and one persisted scene are required.
Project validation and task-specific requirement checks must both pass.
No visible Scene Error sentinel or equivalent runtime failure may appear.
Requested frames, fps, duration, dimensions, and audio policy must reconcile within published tolerance.
A playable MP4 and its export receipt must exist.
Provider raw cost, token usage, and billed price must reconcile or be explicitly marked unavailable.
Scores
The headline creative score comes from blind pairwise human preferences fit with a regularized Bradley–Terry model and reported with sample count and uncertainty. Technical completion is reported separately. “Both miss” is not silently converted into a tie.
Cost and latency
Latency uses wall-clock milestones: queue, generation, validation/repair, export, and total. Raw provider cost is separated from Baz pricing. Medians and tail latency are shown when repeat count supports them.
The July pilot is not Core
The current public dataset is one portrait prompt with one trial per route. It exposed important system failures—including false-success signaling, visible runtime errors, duration mismatch, and an instance-local 120 requests/minute limiter—but cannot support a stable global ranking. It is published as a reproducible diagnostic.
Inspect the pilot →