Endre søk
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf
Improving the Precision of Automatic Program Repair with Machine Learning
KTH, Skolan för elektroteknik och datavetenskap (EECS), Datavetenskap, Teoretisk datalogi, TCS.ORCID-id: 0000-0003-4807-2110
2023 (engelsk)Doktoravhandling, med artikler (Annet vitenskapelig)
Abstract [en]

Automatic program repair as a research field aims to eliminate software bugs and vulnerabilities in an automatic manner. Automatic program repair holds great promise to reduce the debugging cost and increase the productivity of software development. Test suite is one of the most widely used specifications for automatic program repair to specify correct program behavior and guide patch generation. However, test suite is an incomplete specification with limited input-output data. This results in automatic program repair generating patches that merely satisfy test suite specifications, yet fail to repair buggy programs in general. The generation of a great number of incorrect patches leads to a low precision of automatic program repair.

In this thesis, we focus on improving the precision of automatic program repair from three perspectives: patch generation, patch assessment in practice, and patch assessment for scientific usage. This thesis makes contributions to the following in automatic program repair. First of all, to increase the precision of patch generation, we propose two learning-based automatic program repair approaches to encourage the generation of more correct patches with fewer candidate patches. Second, to increase the precision of patch assessment in practice, we propose to build a probabilistic model based on static code features to discard incorrect patches and thus increase the ratio of correct patches to all generated patches. Third, to increase the patch assessment precision for scientific usage, we propose to use automatically generated test cases to discard incorrect patches.

sted, utgiver, år, opplag, sider
Stockholm: KTH Royal Institute of Technology, 2023. , s. 97
Serie
TRITA-EECS-AVL ; 2023:10
HSV kategori
Identifikatorer
URN: urn:nbn:se:kth:diva-323295ISBN: 978-91-8040-469-3 (tryckt)OAI: oai:DiVA.org:kth-323295DiVA, id: diva2:1735113
Disputas
2023-02-24, https://kth-se.zoom.us/j/63393781380, Kollegiesalen, Brinellvägen 6, Stockholm, 13:30 (engelsk)
Opponent
Veileder
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Merknad

QC 20230208

Tilgjengelig fra: 2023-02-08 Laget: 2023-02-07 Sist oppdatert: 2026-03-25bibliografisk kontrollert
Delarbeid
1. SelfAPR: Self-Supervised Program Repair with Test Execution Diagnostics
Åpne denne publikasjonen i ny fane eller vindu >>SelfAPR: Self-Supervised Program Repair with Test Execution Diagnostics
Vise andre…
2023 (engelsk)Inngår i: Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering, Association for Computing Machinery (ACM) , 2023Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Learning-based program repair has achieved good results in a recent series of papers. Yet, we observe that the related work fails to repair some bugs because of a lack of knowledge about 1) the application domain of the program being repaired, and 2) the fault type being repaired. In this paper, we solve both problems by changing the learning paradigm from supervised training to self-supervised training in an approach called SelfAPR. First, SelfAPR generates training samples on disk by perturbing a previous version of the program being repaired, enforcing the neural model to capture project-specific knowledge. This is different from the previous work based on mined past commits. Second, SelfAPR executes all training samples and extracts and encodes test execution diagnostics into the input representation, steering the neural model to fix the kind of fault. This is different from the existing studies that only consider static source code as input. We implement SelfAPR and evaluate it in a systematic manner. We generate 1 039 873 training samples obtained by perturbing 17 open-source projects. We evaluate SelfAPR on 818 bugs from Defects4J, SelfAPR correctly repairs 110 of them, outperforming all the supervised learning repair approaches.

sted, utgiver, år, opplag, sider
Association for Computing Machinery (ACM), 2023
Serie
ASE ’22
HSV kategori
Identifikatorer
urn:nbn:se:kth:diva-323660 (URN)10.1145/3551349.3556926 (DOI)001062775200034 ()2-s2.0-85146336362 (Scopus ID)
Konferanse
the 37th IEEE/ACM International Conference on Automated Software Engineering
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Merknad

QC 20230214

Tilgjengelig fra: 2023-02-08 Laget: 2023-02-08 Sist oppdatert: 2025-12-05bibliografisk kontrollert
2. Neural Program Repair with Execution-based Backpropagation
Åpne denne publikasjonen i ny fane eller vindu >>Neural Program Repair with Execution-based Backpropagation
2022 (engelsk)Inngår i: ICSE '22: Proceedings of the 44th International Conference on Software Engineering, Association for Computing Machinery (ACM) , 2022, s. 1506-1518Konferansepaper, Publicerat paper (Fagfellevurdert)
Abstract [en]

Neural machine translation (NMT) architectures have achieved promising results for automatic program repair. Yet, they have the limitation of generating low-quality patches (e.g., not compilable patches). This is because the existing works only optimize a purely syntactic loss function based on characters and tokens without incorporating program-specific information during neural network weight optimization. In this paper, we propose a novel program repair model called RewardRepair. The core novelty of RewardRepair is to improve NMT-based program repair with a loss function based on program compilation and test execution information, rewarding the network to produce patches that compile and that do not overfit. We conduct several experiments to evaluate RewardRepair showing that it is feasible and effective to use compilation and test execution results to optimize the underlying neural repair model. RewardRepair correctly repairs 207 bugs over four benchmarks. we report on repair success for 121 bugs that are fixed for the first time in the literature. Also, RewardRepair produces up to 45.3% of compilable patches, an improvement over the 39% by the state-of-the-art.

sted, utgiver, år, opplag, sider
Association for Computing Machinery (ACM), 2022
Serie
International Conference on Software Engineering, ISSN 0270-5257
HSV kategori
Identifikatorer
urn:nbn:se:kth:diva-316694 (URN)10.1145/3510003.3510222 (DOI)000832185400122 ()2-s2.0-85130298109 (Scopus ID)
Konferanse
44th ACM/IEEE International Conference on Software Engineering, ICSE 2022, Pittsburgh, 22 May 2022, through 27 May 2022
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)
Merknad

QC 20221109

Part of proceedings: ISBN 978-145039221-1

Tilgjengelig fra: 2022-09-05 Laget: 2022-09-05 Sist oppdatert: 2023-02-08bibliografisk kontrollert
3. Automated Classification of Overfitting Patches with Statically Extracted Code Features
Åpne denne publikasjonen i ny fane eller vindu >>Automated Classification of Overfitting Patches with Statically Extracted Code Features
Vise andre…
2022 (engelsk)Inngår i: IEEE Transactions on Software Engineering, ISSN 0098-5589, E-ISSN 1939-3520, Vol. 48, nr 8, s. 2920-2938Artikkel i tidsskrift (Fagfellevurdert) Published
Abstract [en]

Automatic program repair (APR) aims to reduce the cost of manually fixing software defects. However, APR suffers from generating a multitude of overfitting patches, those patches that fail to correctly repair the defect beyond making the tests pass. This paper presents a novel overfitting patch detection system called ODS to assess the correctness of APR patches. ODS first statically compares a patched program and a buggy program in order to extract code features at the abstract syntax tree (AST) level. Then, ODS uses supervised learning with the captured code features and patch correctness labels to automatically learn a probabilistic model. The learned ODS model can then finally be applied to classify new and unseen program repair patches. We conduct a large-scale experiment to evaluate the effectiveness of ODS on patch correctness classification based on 10,302 patches from Defects4J, Bugs.jar and Bears benchmarks. The empirical evaluation shows that ODS is able to correctly classify 71.9% of program repair patches from 26 projects, which improves the state-of-the-art. ODS is applicable in practice and can be employed as a post-processing procedure to classify the patches generated by different APR systems. 

sted, utgiver, år, opplag, sider
Institute of Electrical and Electronics Engineers Inc., 2022
Emneord
Automatic program repair, Code features, Feature extraction, Maintenance engineering, Overfitting patch, Patch assessment, Predictive models, Software, Syntactics, Tools, Training
HSV kategori
Identifikatorer
urn:nbn:se:kth:diva-308879 (URN)10.1109/TSE.2021.3071750 (DOI)000846878500013 ()2-s2.0-85104193931 (Scopus ID)
Forskningsfinansiär
Swedish Foundation for Strategic Research, TrustfullWallenberg AI, Autonomous Systems and Software Program (WASP)
Merknad

QC 20220216

Tilgjengelig fra: 2022-02-16 Laget: 2022-02-16 Sist oppdatert: 2023-02-08bibliografisk kontrollert
4. Automated patch assessment for program repair at scale
Åpne denne publikasjonen i ny fane eller vindu >>Automated patch assessment for program repair at scale
2021 (engelsk)Inngår i: Empirical Software Engineering, ISSN 1382-3256, E-ISSN 1573-7616, Vol. 26, nr 2, artikkel-id 20Artikkel i tidsskrift (Fagfellevurdert) Published
Abstract [en]

In this paper, we do automatic correctness assessment for patches generated by program repair systems. We consider the human-written patch as ground truth oracle and randomly generate tests based on it, a technique proposed by Shamshiri et al., called Random testing with Ground Truth (RGT) in this paper. We build a curated dataset of 638 patches for Defects4J generated by 14 state-of-the-art repair systems, we evaluate automated patch assessment on this dataset. The results of this study are novel and significant: First, we improve the state of the art performance of automatic patch assessment with RGT by 190% by improving the oracle; Second, we show that RGT is reliable enough to help scientists to do overfitting analysis when they evaluate program repair systems; Third, we improve the external validity of the program repair knowledge with the largest study ever.

sted, utgiver, år, opplag, sider
Springer Nature, 2021
Emneord
Automatic program repair, Automatic patch assessment
HSV kategori
Identifikatorer
urn:nbn:se:kth:diva-292295 (URN)10.1007/s10664-020-09920-w (DOI)000620938900001 ()2-s2.0-85101589275 (Scopus ID)
Forskningsfinansiär
Wallenberg AI, Autonomous Systems and Software Program (WASP)Swedish Foundation for Strategic Research , trustfull
Merknad

QC 20210406

Tilgjengelig fra: 2021-04-06 Laget: 2021-04-06 Sist oppdatert: 2023-02-08bibliografisk kontrollert
5. A Comprehensive Study of Automatic Program Repair on the QuixBugs Benchmark
Åpne denne publikasjonen i ny fane eller vindu >>A Comprehensive Study of Automatic Program Repair on the QuixBugs Benchmark
(engelsk)Manuskript (preprint) (Annet vitenskapelig)
Abstract [en]

Automatic program repair papers tend to repeatedly use the same benchmarks. This poses a threat to the external validity of the findings of the program repair research community. In this paper, we perform an automatic repair experiment on a benchmark called QuixBugs that has been recently published. This benchmark has never been studied in the context of program repair. In this study, we report on the characteristics of QuixBugs, and we design and perform an experiment about the effectiveness of test-suite based program repair on QuixBugs. We study two repair systems, Astor and Nopol, which are representatives of generate-and-validate repair technique and synthesis repair technique respectively. We propose three patch correctness assessment techniques to comprehensively study overfitting and incorrect patches. Our key results are: 1) 13/40 buggy programs in the QuixBugs can be repaired with a test-suite adequate patch; 2) a total of 22 different plausible patches for those 13 buggy programs in the QuixBugs are present in the search space of the considered tools; 3) the three patch assessment techniques discard in total 12/22 patches that are overfitting. This sets a baseline for future research of automatic repair on QuixBugs. Our experiment also highlights the major properties and challenges of how to perform automated correctness assessment of program repair patches. All experimental results are publicly available on Github in order to facilitate future research on automatic program repair.

HSV kategori
Identifikatorer
urn:nbn:se:kth:diva-239890 (URN)
Merknad

QC 20181206

Tilgjengelig fra: 2018-12-04 Laget: 2018-12-04 Sist oppdatert: 2023-02-08bibliografisk kontrollert

Open Access i DiVA

fulltext(1445 kB)733 nedlastinger
Filinformasjon
Fil FULLTEXT04.pdfFilstørrelse 1445 kBChecksum SHA-512
1383487e1b608ec6211102d3f7ad25c72ddd319e2d41c3b53e42ea9b2aafcf6dfd88acc1c9c79732068890d6c7c5aecc1513d8b99490294fa11910ec653d59c0
Type fulltextMimetype application/pdf

Søk i DiVA

Av forfatter/redaktør
Ye, He
Av organisasjonen

Søk utenfor DiVA

GoogleGoogle Scholar
Totalt: 741 nedlastinger
Antall nedlastinger er summen av alle nedlastinger av alle fulltekster. Det kan for eksempel være tidligere versjoner som er ikke lenger tilgjengelige

isbn
urn-nbn

Altmetric

isbn
urn-nbn
Totalt: 2430 treff
RefereraExporteraLink to record
Permanent link

Direct link
Referera
Referensformat
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Annet format
Fler format
Språk
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Annet språk
Fler språk
Utmatningsformat
  • html
  • text
  • asciidoc
  • rtf