Motion model
evidence board.
September 10-second trials are the current public table. The table lists the whole 56-route panel and keeps measured rows separate from interrupted or not-run routes. The July diagnostic pilot remains available as a separate historical view.
September 10-second trials
Current unranked snapshot. 56 attempted routes, 53 measured time and cost rows, 30 completed routes, and 56 planned model routes.
Scroll the chart to compare all measured routes and both columns.
Generation time and token-based provider cost cover the displayed attempt. Costs are before baz.studio markup. Earlier attempts and export costs are separate.
Complete · Incomplete · Infrastructure interrupted
| Model route | Status | Agent | Generation | Provider estimate | Encoded | Notes | Video |
|---|---|---|---|---|---|---|---|
| GPT-6 Astragpt-6-astra · openai | Complete | completed | 358.4s | $2.60 | Pass | Readable typography connects an animated preview with editable React code. Abstract motion and long-held layouts limit motion contrast. | Watch MP4 |
| Claude Fable 5.1claude-fable-5-1 · anthropic | Complete | completed | 421.5s | $3.96 | Pass | Preview-to-code connections show editable scenes. Sparse transitions interrupt the flow. The displayed baz create command is unsupported. | Watch MP4 |
| Gemini 3.8 Flashgemini-3.8-flash · google | Incomplete | failed | 273.0s | $0.3179 | Pass | Readable headlines sit in held interface layouts. Small technical text, limited motion contrast, and unsupported simulated editor claims weaken the result. | Watch MP4 |
| Gemini 3.7 Flashgemini-3.7-flash · google | Incomplete | failed | 236.4s | $0.3737 | Pass | Readable headlines and consistent styling have little visual development. Unsupported statistics, speed claims, technical labels, and a baz build command appear. | Watch MP4 |
| Grok 4.6grok-4.6 · xai | Complete | completed | 2296s | $2.02 | Pass | Generic placeholders, repeated blur reveals, and mostly held layouts. No unsupported claims were observed in the sampled frames. | Watch MP4 |
| GLM-5.3glm-5.3 · openrouter | Incomplete | cancelled | 2400s | $0.4338 | Pass | Generation cancelled after the 40-minute timeout. A separate export of the unchanged partial timeline is playable, but the video remains unfinished. After the opening launch-video prompt, all sampled frames from 2.5 seconds through the end show only background. The video is unfinished. | Watch MP4 |
| GLM-5.3 Flashglm-5.3-flash · openrouter | Incomplete | cancelled | 2400s | - | - | An earlier attempt was interrupted by infrastructure failure. The fresh trial reached the fixed 40-minute generation timeout and was cancelled. A code-generation response returned empty and another call failed during a network outage. Their provider costs were not retained, so the recorded partial usage cannot establish the total trial cost. No completed export is available. | No video |
| Qwen3.8 Max 0902qwen3.8-max-0902 · openrouter | Incomplete | cancelled | 2400s | - | Pass | An earlier attempt was interrupted by infrastructure. Generation reached the fixed 40-minute limit. The saved composition was exported. A code-generation response returned empty before its usage was recorded. The complete provider cost for this trial is unknown. Sampled frames show a mostly empty paper grid that changes to a dark grid near the end. Small timeline text is visible, but there is no main-canvas product story, proof, or CTA. Provisional AI frame review. | Watch MP4 |
| Qwen3.8 Flashqwen3.8-flash · openrouter | Incomplete | failed | 1808s | - | - | An earlier attempt was interrupted by infrastructure. This trial ended with a provider request timeout after about 30 minutes, before creating a scene. No video was exported. Provider cost is unknown because no response usage was recorded. | No video |
| DeepSeek V4 Pro 0813deepseek-v4-pro-0813 · openrouter | Complete | completed | 1308s | $0.6155 | Pass | Earlier attempt: The local attempt lost provider and database connectivity, then its heartbeat expired. No reviewed export or reconciled full-attempt metrics are available. An automatic production continuation was stopped before further testing. Sampled frames show a clear baz.studio launch story with clean brand language. The main weakness is static pacing with some blurred text. Provisional AI frame review. | Watch MP4 |
| Tencent Hy4 Previewhy4-preview · openrouter | Incomplete | failed | 1816s | $0.4884 | Pass | The agent stopped before completing the workflow. The sampled frames show a clear launch story with strong brand fit. The main weakness is possible slow pacing in the hold sections. Provisional AI frame review. | Watch MP4 |
| Nemotron 3.5 Lightningnemotron-3.5-lightning · openrouter | Incomplete | failed | 392.4s | $0.0583 | Pass | The agent stopped before completing the workflow. Every sampled frame shows a Scene Error overlay: the background scene interpolates a color hex with a numeric interpolator, so the composition does not render as intended. Provisional AI frame review confirmed by operator inspection. | Watch MP4 |
| Claude Haiku 4.5claude-haiku-4-5 · anthropic | Incomplete | failed | 251.6s | $0.2652 | Pass | Sampled frames show clipped graphics, overlapping labels, contradictory zero-spend copy, and invalid CLI syntax. | Watch MP4 |
| Claude Sonnet 5claude-sonnet-5 · anthropic | Interrupted | completed | 526.2s | $0.9822 | Pass | Generation completed. Original QA failed during a local build-cache collision. The unchanged saved scenes passed a separate QA/export recovery. Clean approved copy and brand close. Sparse typing and small placeholder graphics do not show a finished product video or editable scenes. | Watch MP4 |
| Claude Sonnet 4.6claude-sonnet-4-6 · anthropic | Incomplete | failed | 345.5s | $0.6955 | Pass | The agent stopped before completing the workflow. A readable problem-to-CLI story has sparse, mostly held layouts and blank transition samples. The displayed baz create launch-video command is unsupported, and the video gives little visual proof of Skills or editable scenes. | Watch MP4 |
| GPT-5.6 Solgpt-5.6-sol · openai | Complete | completed | 421.9s | $1.63 | Pass | The prompt-to-timeline, skill-selection, and editable-code sequence gives a concrete product story. A clipped Finale card, blurred or faint text, and an altered final wordmark weaken the result. | Watch MP4 |
| GPT-5.6 Terragpt-5.6-terra · openai | Incomplete | failed | 115.0s | $0.4071 | Pass | The agent stopped before completing the workflow. A coherent prompt-to-editable-code story uses the supplied brand palette. Abstract product proof, small labels, and a partially clipped code-panel title limit clarity. | Watch MP4 |
| GPT-5.6 Lunagpt-5.6-luna · openai | Incomplete | failed | 290.5s | $0.0901 | Pass | The agent stopped before completing the workflow. Frame samples show baz.studio branding, approved launch-video claims, a recipe/timeline visual, and an editable-code panel. The opener is nearly black, the early headline is clipped on the right edge, and some interface labels are too small to read. | Watch MP4 |
| GPT-5.5 Lowgpt-5.5-low · openai | Complete | completed | 327.8s | $1.62 | Pass | Prompt and CLI views lead into reusable launch-video skills, a preview/code panel, and a baz.studio close. Blurred transitions, small low-contrast captions, a gray placeholder preview, and partially malformed decorative code weaken the proof. | Watch MP4 |
| GPT-5.5 Highgpt-5.5-high · openai | Complete | completed | 1136s | $3.18 | Pass | Earlier attempt: An infrastructure failure interrupted this attempt after generation started. A coherent prompt-to-recipe-to-editable-code story uses readable type and consistent brand styling. Several headline and editor layouts hold across sampled frames; pacing and transition feel need full playback review. Provisional AI frame review. | Watch MP4 |
| GPT-5.5gpt-5.5 · openai | Complete | completed | 468.6s | $2.23 | Pass | The sampled frames show a clear baz.studio product story with strong brand control. The main weakness is minor crop and blur during transitions. Provisional AI frame review. | Watch MP4 |
| GPT-5.4 Minigpt-5.4-mini · openai | Complete | completed | 213.0s | $0.2046 | Pass | The sampled frames show a branded launch story with a prompt-to-editable-code mechanism. The main weakness is broken and overlapping headline text. Provisional AI frame review. | Watch MP4 |
| GPT-5.4gpt-5.4 · openai | Complete | completed | 329.9s | $1.13 | Pass | Earlier attempt: An infrastructure failure interrupted this attempt after generation started. The sampled frames show a clear baz.studio launch story with code and editor proof motifs. The main weakness is long hold time and some blurred small text. Provisional AI frame review. | Watch MP4 |
| GPT-5.4 Nanogpt-5.4-nano · openai | Complete | completed | 153.1s | $0.0346 | Pass | The sampled frames show a clean text-led baz.studio launch message. The main weakness is limited visual development beyond fades and line motion. Provisional AI frame review. | Watch MP4 |
| Gemini 2.5 Flashgemini-2.5-flash · google | Incomplete | failed | 161.1s | $0.0851 | Pass | The agent stopped before completing the workflow. All 21 sampled frames show a Scene Error screen for the background scene. The exported file meets timing, size, and silence requirements but does not contain a usable launch video. | Watch MP4 |
| Gemini 2.5 Progemini-2.5-pro · google | Complete | completed | 301.5s | $0.3123 | Pass | Frame samples show a minimal baz.studio launch spot with a problem line, prompt-strip motif, success check, and final wordmark. Several samples are blank or contain partial text, and the command-strip text is very small. | Watch MP4 |
| Gemini 3.6 Flashgemini-3.6-flash · google | Complete | completed | 226.3s | $0.3471 | Pass | Prompt, code, cards, and a CLI-style call to action form a clear visual system. Sampled frames include an empty transition and unsupported product labels and named-company references. | Watch MP4 |
| Gemini 3.5 Flash-Litegemini-3.5-flash-lite · google | Complete | completed | 59.6s | $0.0740 | Pass | Clean headline, code panel, and baz.studio wordmark sequence. Visible frames rely on long static holds and include a blank opening frame. | Watch MP4 |
| Gemini 3.1 Progemini-3.1-pro-preview · google | Incomplete | failed | 341.6s | $0.3603 | Pass | The agent stopped before completing the workflow. The approved prompt-to-motion and editable-code claims appear with a code window and final baz.studio wordmark. Sampled frames include blank dark transitions, tight text near a divider, and mostly static holds. | Watch MP4 |
| Gemini 3.1 Pro Toolsgemini-3.1-pro-preview-customtools · google | Incomplete | failed | 336.9s | $0.3650 | Pass | Earlier attempt: An infrastructure failure interrupted this attempt after generation started. The agent stopped before completing the workflow. Sampled frames show a clean prompt-to-code story with strong brand fit, but the motion may rely too much on long holds. Provisional AI frame review. | Watch MP4 |
| MiniMax M2.7minimax-m2.7 · minimax | Complete | completed | 283.0s | $0.0895 | Pass | The sampled frames show a CLI-to-pipeline-to-logo story for baz.studio. The main weakness is long static or near-empty holds. Provisional AI frame review. | Watch MP4 |
| MiniMax M3minimax-m3 · minimax | Complete | completed | 280.9s | $0.1101 | Pass | The sampled frames show a clean baz.studio launch arc with strong brand fit. The main weakness is slow, sparse motion between major beats. Provisional AI frame review. | Watch MP4 |
| DeepSeek V4 Flashdeepseek-v4-flash · openrouter | Complete | completed | 585.0s | $0.0334 | Pass | The sampled frames show clean brand typography and approved claims. The main weakness is sparse motion with blank gaps. Provisional AI frame review. | Watch MP4 |
| DeepSeek V4 Flash 0731deepseek-v4-flash-0731 · openrouter | Incomplete | cancelled | 2400s | $0.0304 | Pass | Generation reached the fixed 40-minute limit. The saved composition was exported. The sampled frames show blank screens and a scene generation failure. The main weakness is that no usable video was produced. Provisional AI frame review. | Watch MP4 |
| DeepSeek V4 Prodeepseek-v4-pro · openrouter | Complete | completed | 680.9s | $0.1116 | Pass | The sampled frames show a clear baz.studio launch arc with typed problem text, a CLI prompt, and the wordmark. The main weakness is sparse motion and an awkward missing space in the CTA copy. Provisional AI frame review. | Watch MP4 |
| Kimi K2.6kimi-k2.6 · openrouter | Complete | completed | 1441s | $0.5639 | Pass | The sampled frames show a clean baz.studio text sequence with approved brand claims. The main weakness is sparse motion and one clipped headline. Provisional AI frame review. | Watch MP4 |
| Kimi K2.7 Codekimi-k2.7-code · openrouter | Incomplete | failed | 741.1s | $0.3143 | - | Latest receipt note: Latest receipt did not pass the benchmark gates: agent_not_completed, requirements_failed, export_failed, mp4_missing, ffprobe_failed. | No video |
| Kimi K3kimi-k3 · openrouter | Complete | completed | 785.0s | $1.78 | Pass | Sampled frames show a clear CLI-to-video-to-code story for baz.studio. The main weakness is generic motion with brief title overlap. Provisional AI frame review. | Watch MP4 |
| GLM-5.2glm-5.2 · openrouter | Complete | completed | 506.9s | $0.3544 | Pass | The sampled frames show a clean baz.studio launch arc with readable text and brand colors. The main weakness is generic product-card imagery and some long held layouts. Provisional AI frame review. | Watch MP4 |
| Qwen3 Coder Plusqwen3-coder-plus · openrouter | Incomplete | failed | 163.1s | $0.2145 | Pass | The agent stopped before completing the workflow. Sampled frames show a sparse baz.studio title and prompt button, then a green accent dot. The main weakness is low story content and limited motion. Provisional AI frame review. | Watch MP4 |
| Qwen 3.7 Plusqwen3.7-plus · openrouter | Complete | completed | 1064s | $0.4587 | Pass | The sampled frames show a clean text-led launch message with approved baz.studio copy and wordmark. The main weakness is sparse motion with blank transition beats. Provisional AI frame review. | Watch MP4 |
| Qwen 3.7 Maxqwen3.7-max · openrouter | Complete | completed | 1098s | $0.7042 | Pass | Sampled frames show a prompt typed into a dark console, a bright transition, icons, and the baz.studio wordmark. The main weakness is a thin story with one overexposed transition that hides text. Provisional AI frame review. | Watch MP4 |
| Qwen 3.7 Flashqwen3.7-flash · openrouter | Incomplete | failed | 283.7s | $0.0520 | Pass | The agent stopped before completing the workflow. The sampled frames show a clean launch explainer with a clear code-to-video story. The main weakness is slow, sparse motion in the middle. Provisional AI frame review. | Watch MP4 |
| Mistral Medium 3.5mistral-medium-3.5 · openrouter | Incomplete | failed | 158.2s | $0.8697 | Pass | The agent stopped before completing the workflow. The sampled frames show a clean text-led launch video with clear brand colors and wordmark. The main weakness is sparse motion and long blank gaps. Provisional AI frame review. | Watch MP4 |
| Mistral Small 2603mistral-small-2603 · openrouter | Complete | completed | 133.3s | $0.0209 | Pass | The workflow completed and the file passes encoded checks, but every sampled frame shows a Scene Error overlay: the background scene interpolates a color hex with a numeric interpolator, so the composition does not render as intended. Provisional AI frame review confirmed by operator inspection. | Watch MP4 |
| Nemotron 3 Ultranemotron-3-ultra-550b-a55b · openrouter | Incomplete | failed | 304.6s | $0.3545 | Pass | The agent stopped before completing the workflow. The sampled frames show a clean text-only launch message with approved baz.studio claims. The main weakness is very limited motion and several blank pauses. Provisional AI frame review. | Watch MP4 |
| MiMo V2.5 Promimo-v2.5-pro · openrouter | Incomplete | failed | 1962s | $0.1596 | Pass | The agent stopped before completing the workflow. The sampled frames show a clean baz.studio launch sequence with wordmark, claims, and a code editor. The main weakness is sparse motion and limited product proof. Provisional AI frame review. | Watch MP4 |
| Gemini 3.5 Flashgemini-3.5-flash · openrouter | Complete | completed | 257.9s | $0.7900 | Pass | Sampled frames show a coherent baz.studio story from source video to recipe schema to editable code, with readable branded type. An automated error-screen check flagged seven pale frames, but manual inspection found no error overlay, so the flag needs playback confirmation. Provisional AI frame review. | Watch MP4 |
| Qwen3.8 Max (OpenRouter)qwen3.8-max-openrouter · openrouter | Incomplete | cancelled | 2400s | $0.7446 | Pass | Generation reached the fixed 40-minute limit. The saved composition was exported. The sampled frames show a partial editable-code sequence, but error screens and a black ending dominate the video. The main weakness is failed scene generation. Provisional AI frame review. | Watch MP4 |
| Muse Spark 1.1muse-spark-1.1 · openrouter | Incomplete | failed | 7.5s | $0.0000 | - | The provider refused the request with HTTP 403: this model requires an 18+ age confirmation on the OpenRouter account. No generation ran and no provider cost was recorded. This is a provider access gate, not a candidate outcome. | No video |
| Tencent Hy3hy3 · openrouter | Complete | completed | 436.9s | $0.0536 | Pass | The sampled frames show a clean baz.studio launch message with code and chart visuals. The main weakness is generic motion structure with long holds. Provisional AI frame review. | Watch MP4 |
| Step 3.7 Flashstep-3.7-flash · openrouter | Complete | completed | 710.4s | $0.1415 | Pass | The sampled frames show a clear baz.studio pitch with correct brand cues. The main weakness is text overlap during key transitions. Provisional AI frame review. | Watch MP4 |
| Grok 4.5grok-4.5 · xai | Complete | completed | 330.2s | $0.6879 | Pass | The sampled frames show a clean baz.studio launch arc with strong typography and brand fit. The main weakness is sparse presentation-style motion with some card overlap. Provisional AI frame review. | Watch MP4 |
| Grok Build 0.1grok-build-0.1 · xai | Complete | completed | 674.3s | $0.2423 | Pass | The sampled frames show a clean text-led baz.studio pitch with strong brand fit. The main weakness is sparse motion and several empty moments. Provisional AI frame review. | Watch MP4 |
| Claude Opus 5claude-opus-5 · anthropic | Complete | completed | 581.6s | $3.14 | Pass | Frames show a clear prompt-to-code-to-video story for baz.studio. Main weakness is long sparse pauses and simple chart motion. Provisional AI frame review. | Watch MP4 |
| Claude Opus 4.8claude-opus-4-8 · anthropic | Incomplete | failed | 265.3s | $1.65 | Pass | The agent stopped before completing the workflow. The sampled frames show a clean prompt to code to video story for baz.studio. The main weakness is restrained motion with a possible blank transition beat. Provisional AI frame review. | Watch MP4 |
This is not a quality ranking. It has no blind preference score and no order implied by cost. Rows without recorded metrics stay in the table but do not appear in the bars. Encoded pass checks the video file, timing, dimensions, and silence; it does not establish visual quality or a complete workflow. Notes are provisional AI frame reviews.
Download this snapshot (JSON) · Open the July diagnostic pilot