In a head-to-head comparison, children aged three to seven have outperformed some of the most advanced artificial intelligence systems in basic reasoning and problem-solving tasks, according to a study published in Perspectives on Psychological Science. The research, led by scientists at the University of California, Berkeley, focused on tool innovation and causal reasoning—areas where AI's reliance on statistical patterns may be a liability.
The study presented several AI models, including OpenAI's GPT-4 and GPT-3.5 Turbo, Anthropic's Claude, and Google's FLAN-T5, alongside the children with a series of challenges. In one test, participants were asked to draw a circle but were only given a ruler, a teapot, or a stove as potential tools. The children chose the teapot—a round object that could serve as a stencil—85% of the time. GPT-4, the best-performing AI, succeeded only 76% of the time, while other models frequently defaulted to the ruler, which is straight-edged and unsuitable for drawing a circle.
The researchers noted that AI's tendency to pick the ruler stems from its training data, where rulers are commonly associated with drawing. However, as the study authors explain, “Discovering novel functions in everyday tools is not about finding the statistically nearest neighbor from lexical co-occurrence patterns. Rather, it is about appreciating the more abstract functional analogies and causal relationships between objects that do not necessarily belong to the same category or are associated in text.”
In a second experiment, both groups interacted with a virtual “blicket detector”—a device that lights up and plays music when certain objects are placed on it. The children quickly figured out which objects triggered the machine, with even four-year-olds spontaneously experimenting and learning the cause-and-effect structure. The AI models, despite extensive training, struggled to infer these novel causal relationships.
Why AI Falls Short on Innovation
The study underscores a fundamental difference: AI excels at imitation and pattern recognition, but innovation—whether designing new tools or using old ones in new ways—requires a type of abstract reasoning that current models lack. The authors suggest that children's curiosity and intrinsic motivation drive their learning, whereas AI systems are trained on vast amounts of human data without the same exploratory drive.
While the study has limitations, such as the lack of a universal definition of intelligence, it offers insight into the distinct nature of human cognition. As AI becomes more integrated into daily life, understanding these differences may help determine where AI can be most useful—and where it cannot replace human ingenuity.