Large language models were built to predict the next word in a sequence, but their surprising similarities with the human mind are taking cognitive science in bold new directions.
Researchers at the University of Michigan are using those parallels to better understand complex human abilities such as language use and reasoning, and, ultimately, to shed new light on creativity, morality and mental health. Richard Lewis and Chandra Sripada study the computational basis of the mind and brain. Lewis, Arthur F. Thurnau Professor and John R. Anderson Collegiate Professor of Psychology, Linguistics and Cognitive Science, examines language processing, decision-making and skilled performance. Sripada, Theophile Raphael Research Professor of Clinical Neurosciences and professor of psychiatry and philosophy, focuses on executive functions, neurodevelopment and decision-making.
Over the past five years, the advent of generative AI has provided the first working examples of how certain complex mental functions may operate in areas of the mind and brain that are otherwise difficult to study. These models are giving researchers new testable hypotheses.
“We’re discovering a lot more convergence between how the human mind and new AI systems work than you might expect. In some ways, that’s surprising because AI hasn’t lived a life like a human. On the other hand, it’s not surprising because AI and cognitive science have a long, rich history of working together. Many of the core ideas these systems were trained on came out of cognitive science, including some of the important ideas behind neural networks.”
[...]
To better understand this working memory, or short-term memory in language processing, Lewis leads transformer reading-time studies. The transformer architecture, the design behind tools like ChatGPT, processes sentences one word at a time by computing relations between each word and what came before.
Because that resembles how humans are believed to handle language in working memory, Lewis examines the transformer’s internal memory-retrieval patterns, rather than just its output, as it processes sentences. Those patterns predict how long a participant’s eyes will spend on particular words, even though the models were never trained for that purpose.
Sripada’s research examines fast and slow thinking, a dual-process concept in which automatic processing routes allow for fast thinking, while deeper processing supports slow thinking.
Cognitive scientists study fast and slow thinking using conflict tasks, such as the Stroop task. Participants read color words printed in potentially conflicting ink colors. For example, the word “red” might be printed in blue ink and they must name the ink color. Reading the word is the more practiced response, so the brain’s fast processing tends to produce “red.” Naming the ink color requires slower processing.
Surprisingly, LLMs have developed a similar duality, even though they were just trained to predict the next word. They possess automatic processing as well as in-context processing, or the ability to draw on deeper resources and detect hidden patterns in a prompt.
“That’s shocking, because I would have thought dual process structure was evolutionarily shaped and is essentially hard-wired,” Sripada said. “This changes our understanding of what this duality is and how it emerges. We’re still studying the same phenomenon we always have, but with a new perspective on what it is and how it emerged.”
