Image
MultivationBench: AI Models Still Fail at Understanding Why People Act — and the Numbers Prove It
A new benchmark from HKUST tests eight multimodal AI models on 16,092 motivation-reasoning questions grounded in Maslow's hierarchy. The best model scores 39.2% exact match — humans hit 70.6%. No model can track character motivation correctly through an entire story.