kth.sePublications KTH
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
SelfAPR: Self-Supervised Program Repair with Test Execution Diagnostics
KTH, School of Electrical Engineering and Computer Science (EECS), Computer Science, Theoretical Computer Science, TCS.ORCID iD: 0000-0003-4807-2110
Université Polytechnique Hauts-de-France, France.
Show others and affiliations
2023 (English)In: Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering, Association for Computing Machinery (ACM) , 2023Conference paper, Published paper (Refereed)
Abstract [en]

Learning-based program repair has achieved good results in a recent series of papers. Yet, we observe that the related work fails to repair some bugs because of a lack of knowledge about 1) the application domain of the program being repaired, and 2) the fault type being repaired. In this paper, we solve both problems by changing the learning paradigm from supervised training to self-supervised training in an approach called SelfAPR. First, SelfAPR generates training samples on disk by perturbing a previous version of the program being repaired, enforcing the neural model to capture project-specific knowledge. This is different from the previous work based on mined past commits. Second, SelfAPR executes all training samples and extracts and encodes test execution diagnostics into the input representation, steering the neural model to fix the kind of fault. This is different from the existing studies that only consider static source code as input. We implement SelfAPR and evaluate it in a systematic manner. We generate 1 039 873 training samples obtained by perturbing 17 open-source projects. We evaluate SelfAPR on 818 bugs from Defects4J, SelfAPR correctly repairs 110 of them, outperforming all the supervised learning repair approaches.

Place, publisher, year, edition, pages
Association for Computing Machinery (ACM) , 2023.
Series
ASE ’22
National Category
Computer Sciences
Identifiers
URN: urn:nbn:se:kth:diva-323660DOI: 10.1145/3551349.3556926ISI: 001062775200034Scopus ID: 2-s2.0-85146336362OAI: oai:DiVA.org:kth-323660DiVA, id: diva2:1735153
Conference
the 37th IEEE/ACM International Conference on Automated Software Engineering
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Note

QC 20230214

Available from: 2023-02-08 Created: 2023-02-08 Last updated: 2025-12-05Bibliographically approved
In thesis
1. Improving the Precision of Automatic Program Repair with Machine Learning
Open this publication in new window or tab >>Improving the Precision of Automatic Program Repair with Machine Learning
2023 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Automatic program repair as a research field aims to eliminate software bugs and vulnerabilities in an automatic manner. Automatic program repair holds great promise to reduce the debugging cost and increase the productivity of software development. Test suite is one of the most widely used specifications for automatic program repair to specify correct program behavior and guide patch generation. However, test suite is an incomplete specification with limited input-output data. This results in automatic program repair generating patches that merely satisfy test suite specifications, yet fail to repair buggy programs in general. The generation of a great number of incorrect patches leads to a low precision of automatic program repair.

In this thesis, we focus on improving the precision of automatic program repair from three perspectives: patch generation, patch assessment in practice, and patch assessment for scientific usage. This thesis makes contributions to the following in automatic program repair. First of all, to increase the precision of patch generation, we propose two learning-based automatic program repair approaches to encourage the generation of more correct patches with fewer candidate patches. Second, to increase the precision of patch assessment in practice, we propose to build a probabilistic model based on static code features to discard incorrect patches and thus increase the ratio of correct patches to all generated patches. Third, to increase the patch assessment precision for scientific usage, we propose to use automatically generated test cases to discard incorrect patches.

Place, publisher, year, edition, pages
Stockholm: KTH Royal Institute of Technology, 2023. p. 97
Series
TRITA-EECS-AVL ; 2023:10
National Category
Computer Sciences
Identifiers
urn:nbn:se:kth:diva-323295 (URN)978-91-8040-469-3 (ISBN)
Public defence
2023-02-24, https://kth-se.zoom.us/j/63393781380, Kollegiesalen, Brinellvägen 6, Stockholm, 13:30 (English)
Opponent
Supervisors
Funder
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Note

QC 20230208

Available from: 2023-02-08 Created: 2023-02-07 Last updated: 2026-03-25Bibliographically approved

Open Access in DiVA

No full text in DiVA

Other links

Publisher's full textScopushttps://doi.org/10.1145/3551349.3556926

Authority records

Ye, HeMartinez, MatiasMonperrus, Martin

Search in DiVA

By author/editor
Ye, HeMartinez, MatiasMonperrus, Martin
By organisation
Theoretical Computer Science, TCS
Computer Sciences

Search outside of DiVA

GoogleGoogle Scholar

doi
urn-nbn

Altmetric score

doi
urn-nbn
Total: 126 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf