Skip to benchmark content
MotionBenchHosted by Baz Studio
Pilot datasetOne prompt · one trial · 33 routes · not the official Core rankingRead the limits
Model field · July 31, 2026

Motion model
evidence board.

Real outputs from one controlled pilot. Filter the routes, play the result, and inspect why technically failed runs are excluded.

Time and raw model cost

Lower-left is faster and cheaper. Every point is named; color shows the sampled visual tier, and an open point failed a technical gate.

Tier ATier BTier CTier FExcluded

Scroll horizontally to inspect every labeled route →

Model routeEvidence
GPT-5.5 Lowopenai · O+CAEligible$0.4468.0sInspect run →
GPT-5.6 Lunaopenai · O+CAEligible$0.053481.2sInspect run →
GPT-5.6 Terraopenai · O+CAEligible$0.6886.9sInspect run →
Claude Opus 4 8anthropic · O+CAEligible$0.94182.6sInspect run →
Claude Opus 4 7anthropic · O+CAEligible$0.99225.3sInspect run →
Claude Opus 4 6anthropic · O+CAEligible$1.15324.3sInspect run →
GPT 5.4 Nanoopenai · O+CAExcluded$0.008954.4sInspect run →
GPT-5.5 Highopenai · O+CAExcluded$0.54294.3sInspect run →
GPT 5.4 Miniopenai · O+CBEligible$0.032042.0sInspect run →
Claude Haiku 4 5anthropic · OBEligible$0.10113.5sInspect run →
GPT 5.5openai · O+CBEligible$0.27168.9sInspect run →
o4-miniopenai · CBEligible$0.77208.6sInspect run →
Claude Sonnet 4 6anthropic · O+CBEligible$0.40214.3sInspect run →
GPT-4.1 Mini FTopenai · CBEligible$0.62223.7sInspect run →
GPT-5.1openai · CBEligible$0.41225.5sInspect run →
GPT 5.2openai · CBEligible$0.45272.2sInspect run →
GPT-5 Miniopenai · CBEligible$0.48288.6sInspect run →
GPT 5 Nanoopenai · CBEligible$0.59362.0sInspect run →
Claude Opus 5anthropic · O+CBEligible$1.25385.3sInspect run →
GLM 5.2openrouter · O+CBEligible$0.52545.5sInspect run →
Kimi K2.7 Codeopenrouter · O+CBExcluded$0.65458.5sInspect run →
Gemini 2.5 Flashgoogle · O+CCEligible$0.095981.2sInspect run →
Gemini 2.5 Progoogle · O+CCEligible$0.15102.3sInspect run →
MiniMax M2.7 Highspeedminimax · OCEligible$0.0243124.6sInspect run →
Deepseek V4 Flashopenrouter · O+CCEligible$0.0531282.3sInspect run →
Claude Haiku 4.5 (dated)anthropic · CCEligible$0.91333.4sInspect run →
Deepseek V4 Proopenrouter · O+CCEligible$0.11364.0sInspect run →
GPT 5.4openai · O+CCExcluded$0.1189.6sInspect run →
Gemini 3.5 Flashopenrouter · O+CCExcluded$0.13196.3sInspect run →
Kimi K2.6openrouter · O+CCExcluded$0.0231899.3sInspect run →
Minimax M2.7minimax · O+CFExcluded$0.0053157.1sInspect run →
Kimi K3openrouter · O+CFExcluded$0.22798.3sInspect run →
Openrouter Fusionopenrouter · O+CFExcluded$1.05875.7sInspect run →

Pilot order is diagnostic, not an official preference ranking. Quality tiers came from sampled human review; no pairwise vote score was collected.