Human vs LLM: A Comparative Performance Analysis in a Custom Strategic Deduction Game
2025 (English)In: Americas Conference on Information Systems, AMCIS 2025, Association for Information Systems , 2025, Vol. 5, p. 3468-3477Conference paper, Published paper (Refereed)
Abstract [en]
Understanding how large language models perform relative to humans in socially interactive, deduction-based tasks is vital for advancing AI applications. This study compares the performance of human players and GPT-40 in Guess vs. AI, a custom strategic deduction game. Drawing on data from 85 completed games, the AI-opponent achieved a significantly higher win rate than human players (63.5%, p = 0.009) and required fewer questions to identify the target (humans: 17, AI-opponent: 9). These findings highlight GPT-40's strengths in systematic reasoning, pattern recognition and efficient decision-making. While showcasing the potential of large language models in structured deduction scenarios, they also emphasize the need for further research into AI adaptability in more socially complex tasks. Future directions include expanding demographic diversity, exploring additional game formats or different large language models and investigating potential human-AI collaborations rather than strictly competitive environments.
Place, publisher, year, edition, pages
Association for Information Systems , 2025. Vol. 5, p. 3468-3477
Keywords [en]
AI vs Human Performance, Deductive Reasoning, Generative AI, Human-AI Interaction, Large Language Models, Prompt Engineering, Strategy Games
National Category
Information Systems
Identifiers
URN: urn:nbn:se:kth:diva-385580Scopus ID: 2-s2.0-105025162305OAI: oai:DiVA.org:kth-385580DiVA, id: diva2:2086812
Conference
2025 Americas Conference on Information Systems, AMCIS 2025, Montreal, Canada, August 14-16, 2025
Note
Part of ISBN 9798331327743
QC 20260716
2026-07-162026-07-162026-07-16Bibliographically approved