kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Mokav: Execution-driven differential testing with LLMs
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Theoretical Computer Science, TCS. ETH Zurich, Switzerland; KTH Royal Institute of Technology, Sweden.ORCID iD: 0000-0003-2183-9633
Sharif University of Technology, Iran.
ETH Zurich, Switzerland.
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Theoretical Computer Science, TCS.ORCID iD: 0000-0003-3505-3383
2025 (English)In: Journal of Systems and Software, ISSN 0164-1212, E-ISSN 1873-1228, Vol. 230, article id 112571Article in journal (Refereed) Published
Abstract [en]

It is essential to detect functional differences between programs in various software engineering tasks, such as automated program repair, mutation testing, and code refactoring. The problem of detecting functional differences between two programs can be reduced to searching for a difference exposing test (DET): a test input that results in different outputs on the subject programs. In this paper, we propose MOKAV, a novel execution-driven tool that leverages LLMs to generate DETs. MOKAV takes two versions of a program (P and Q) and an example test input. When successful, MOKAV generates a valid DET, a test input that leads to provably different outputs on P and Q. MOKAV iteratively prompts an LLM with a specialized prompt to generate new test inputs. At each iteration, MOKAV provides execution-based feedback from previously generated tests until the LLM produces a DET. We evaluate MOKAV on 1535 pairs of Python programs collected from the Codeforces competition platform and 32 pairs of programs from the QuixBugs dataset. Our experiments show that MOKAV outperforms the state-of-the-art, Pynguin and Differential Prompting, by a large margin. MOKAV can generate DETs for 81.7% (1,255/1535) of the program pairs in our benchmark (versus 4.9% for Pynguin and 37.3% for Differential Prompting). We demonstrate that the iterative and execution-driven feedback components of the system contribute to its high effectiveness.

Place, publisher, year, edition, pages
Elsevier BV , 2025. Vol. 230, article id 112571
Keywords [en]
Behavioral difference, Large language models, Test generation
National Category
Software Engineering
Identifiers
URN: urn:nbn:se:kth:diva-369034DOI: 10.1016/j.jss.2025.112571ISI: 001538605500001Scopus ID: 2-s2.0-105011045033OAI: oai:DiVA.org:kth-369034DiVA, id: diva2:1997514
Note

QC 20250912

Available from: 2025-09-12 Created: 2025-09-12 Last updated: 2025-11-13Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopus

Authority records

Etemadi, KhashayarMonperrus, Martin

Search in DiVA

By author/editor
Etemadi, KhashayarMonperrus, Martin
By organisation
Theoretical Computer Science, TCS
In the same journal
Journal of Systems and Software
Software Engineering

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 60 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf