CogArena Benchmark Finds Limited Evidence for Distinct Cognitive Profiles in LLMs
A new arXiv preprint introduces CogArena, a benchmark using 13 procedurally generated paradigms to test whether large language models (LLMs) exhibit stable, distinct cognitive abilities. Evaluating 55 open-weight models, the study finds that most performance correlations are positive and largely explained by a single common factor, with little evidence for stable, multi-dimensional cognitive profiles. Targeted interventions and theory-aligned prompts did not yield selective improvements or robust, interpretable ability distinctions.
Why it matters: The findings challenge the common practice of assigning human-like cognitive labels to LLM performance, raising questions about how model capabilities are interpreted and benchmarked.
Full story at: arXiv Computation and Language ↗