Skip to benchmark contentMotionBenchHosted by Baz Studio Human assessment
Pilot datasetOne prompt · one trial · 33 routes · not the official Core rankingRead the limits
openai · C
GPT-5.1
gpt-5.1-2025-11-13Pilot visual tierB
Good fragment field and clean final hold.
One diagnostic trial. This is not a durable model capability claim.