AI models flub these intelligence tests. Can you fare any better?
AI Summary
Искусственный интеллект продолжает развиваться, демонстрируя значительные успехи в решении головоломок и игр. Однако, несмотря на улучшения, современные модели все еще сталкиваются с трудностями, особенно в визуальных задачах и при изменении классических загадок. Ученые из Колумбийского университета отметили, что к началу 2025 года некоторые модели могут почти идеально решать известные головоломки, но все же остаются области, где люди превосходят ИИ.
Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur Samuel about an algorithm that learned to play checkers. Chess and the Chinese board game Go are famous AI test beds too.
Judged purely on its puzzling skills, AI is improving a lot—and quickly. In late 2024, a team of scientists from Columbia University showed that even the best models could figure out only 18% of the infamous New York Times Connections puzzles; by early 2025, some models could solve them near perfectly every time.
But puzzles do more than just highlight the inexorable advance of AI capabilities. Seeing where models succeed and fail—and where we humans still beat them—can provide a useful window into the technology’s strengths and weaknesses. Despite advances, today’s models still fumble: Subtle changes in classic riddles often trip them up, and visual puzzles are a particular weak spot.
Here you’ll have the chance to test your wits on puzzles that have stumped models at one time or another. Some might be as tricky for you as they were for the AI; others are so simple that they’ll have you doubting whether AI is really intelligent at all. Each one highlights at least one way in which machine and human cognition differ. If you ace the test, you’ll have proved that you can out-puzzle an AI—at least for now.
Spatial Reasoning
Let’s start with a domain where humans have a huge advantage: spatial reasoning. If you’ve ever taken an IQ test, you may have done a mental rotation problem. These puzzles ask you to determine whether different images represent the same objects from different angles. Though today’s language models typically have the ability to analyze visual inputs, they still fail abysmally at these puzzles. For all the talk of how world models can help AI understand physical environments, LLMs still don’t seem to be able to manipulate 3D objects the way spatial thinkers like architects and mechanical engineers can.
Mental Rotation
Instructions: Choose the answer that shows the object in the prompt, but from a different angle. In each case, there’s only one correct answer!
Memory & Adaptability
Frontier LLMs have extraordinary memories; they were exposed to a monstrous volume of facts during training and can recite many of them faithfully. That’s an asset for outcompeting humans at trivia, but it can also be a liability. When a puzzle closely resembles one a model saw during training, the model may whiz by key differences and respond with what it memorized.
This held true in a 2024 study in which researchers from Google and the University of Illinois Urbana-Champaign trained and tested models on slight variations of a classic type of puzzle called Knights and Knaves. In these problems, some characters always tell the truth and others always lie, and you have to figure out who’s who. The same principle may be at work in a test called SimpleBench. These questions resemble more complicated problems that models likely encountered in training. Humans spot the trick, but even top-tier models trip.
Knights and Knaves
Instructions: The only thing you need to know to solve these puzzles is that knights always tell the truth and knaves always lie. Determine who’s what on the basis of what each character says.
SimpleBench
Instructions: Read these SimpleBench problems carefully, and you should be able to figure out the answers in no time.
Abstract & Visual Reasoning
AI doesn’t just bungle visual problems in 3D—two dimensions can trip it up as well. That’s a major factor in how well models do on the most famous puzzle-based benchmark, ARC-AGI. These problems require you to infer abstract, general rules from a set of examples. Models do better on ARC puzzles when they receive each grid not as an image but as a string of numbers that encodes the color of each cell.
Research suggests that even when models answer ARC-AGI questions correctly, they often do so using byzantine and non-generalizable rules, whereas humans draw on simple visual concepts. Despite these disadvantages, models have gotten quite good at ARC-AGI over the past year, but some puzzles—such as the one printed here—still stump them.
ARC-AGI
Instructions: Study the three pairs of grids shown below to figure out the rule that dictates how the ones on the left transform into the ones on the right. Then get out your markers or colored pencils and fill in the fourth grid using that rule. (The solution is the same no matter which way the grids are oriented.)
Intuition
It’s not just AI models that fall into traps. We humans have our own cognitive foibles, many of which AI do
Related News
TechnologyGoogle’s new AI transcription edits out your ‘ums’ and ‘ahs’
Google обновил Gemini Audio, добавив новую функцию транскрипции Gemini 3.5 Transcribe, которая автоматически удаляет слова-паразиты и распознает специализированную терминологию на более чем 85 языках. Эта версия значительно улучшает многозадачность и снижает количество ошибок в словах по сравнению с предыдущей моделью Chirp 3. Ожидается, что Google вскоре выпустит модель Gemini 3.5 Pro.
Radar makes podcasts searchable — and usable by AI agents
Компания Particle представила новую платформу для подкастов, которая транскрибирует и анализирует более 130,000 подкастов, делая их разговоры доступными для поиска в интернете. Эта информация также может быть использована AI-агентами через API и MCP.
В России утвердят условия запуска 5G
В России утвердят условия для запуска сетей 5G, позволяя операторам связи «большой четверки» использовать существующие базовые станции для развертывания новых технологий. Это решение должно ускорить внедрение сетей пятого поколения в стране.