A benchmark released this week asks whether multimodal AI systems can interpret visual situations that people often understand intuitively. Scale AI’s Humanity’s Sixth Sense, or HSS, evaluates models on open-ended questions about images and video. Its authors report that the strongest tested system scored well below the human comparison group. The result is evidence about this benchmark, not a general ranking of intelligence.
Take the prompts our creative team actually uses
The key to maintaining brand quality at scale? Mastering how to communicate with AI models and embedding them into your creative strategy.
Join our honest discussion with Ari Murray, Chief Digital Officer at Salt and Stone, and get 5 tips for expert LLM prompting with ready to run prompts.
HSS contains 522 tasks built from 288 images and 234 video clips. The clips total 17.6 hours. The benchmark divides questions among four broad domains and eleven subdomains. They include spatial, temporal, social, and abstract reasoning. Models must produce an answer rather than select from a short list, so scoring includes interpretation and response generation.
The authors say they created 3,466 candidate tasks and retained 522 after three review rounds. That filtering process is relevant because benchmark performance depends on what its designers include and exclude. The final set is a selected sample, not a random survey of every visual problem encountered in daily life. Its results cannot establish how often similar failures occur outside the tested questions.
Scale reports 25 evaluated models from eight providers. Under the benchmark’s conditions, GPT-6 Astra with maximum reasoning received a 53.6 percent score. The human annotator group reached 93.1 percent, while the model median was 30.9 percent. The reported model measure averages pass-at-one performance across three attempts per task. The report also provides bootstrap intervals over tasks, which express uncertainty within this sample.
200+ Proven Ways to Make Money With AI in 2026
The next wave of millionaires will be people who figured out how to make AI work for them.
The window to get ahead is still open. But not for long.
Here are 200+ proven ways to make money with AI in 2026.
Sign up for Superhuman AI, the free daily newsletter read by 1M+ professionals, and get instant access to all 200+ ways to profit from AI this year.
Those numbers need a clear denominator. They describe answers to HSS tasks, not the fraction of all real-world visual judgments a model can solve. A score can shift when a benchmark changes its images, questions, scoring rules, or model access conditions. The figures also do not say that human observers are uniformly correct in every setting. They compare this particular set of annotations under the authors’ protocol.
The paper reports a video penalty for 23 of the 25 models. Average performance fell by 7.3 points on video relative to images. Social questions were the lowest-performing domain for 21 models, with a mean score of 24.4 percent compared with 34.1 percent across other domains. These are aggregate patterns in the tested set. They suggest useful evaluation targets, but they do not identify a single cause for every error.
The distinction between seeing objects and interpreting a situation is central to the test. A system may recognize people or objects yet miss timing, relationships, implied intent, or a change across frames. HSS turns some of those distinctions into scored questions. That makes the benchmark useful for probing model behavior, while leaving open whether its labels capture the full range of human context.
The evidence also has an independence limit. Scale developed the benchmark and reports the model evaluations. Its report and leaderboard are primary records of those results, but they are not independent replications. A third-party page mirrors leaderboard entries; it does not rerun all models. The comparison should therefore be treated as a reported evaluation until separate researchers reproduce the protocol and scores.
Most importantly, HSS does not measure general intelligence or determine whether a system is close to artificial general intelligence. It measures a narrow collection of visual reasoning tasks. Its human-model gap is meaningful within that scope, while claims about broad capability require other forms of evidence. The strongest next test would publish task-level methods, allow independent replication, and examine whether improvements transfer to new images and videos. Until then, HSS is a diagnostic snapshot, not a verdict on multimodal AI as a whole.
Some teams never seem to stop moving. They're on Attio, the agentic CRM.
Every customer signal is captured in one shared context layer, always current and compounding. Agents and workflows build pipeline, chase every buying signal, and move deals forward, an always-on revenue engine running alongside your team.
With Attio, you’ll get:
Leads automatically prioritised and routed to the right rep
Expansion and risk signals caught the moment they land
Follow-ups written in your voice, already there when you arrive
Teams like Parallel, Turbopuffer, and Wordsmith build on Attio. Are you one of them?





