Skip to benchmark contentMotionBenchHosted by Baz Studio Human assessment
Pilot datasetOne prompt · one trial · 33 routes · not the official Core rankingRead the limits
openai · O+C
GPT 5.4
gpt-5.4Pilot visual tierC
Valid but small and visually conservative in sampled frames.
One diagnostic trial. This is not a durable model capability claim.