Comparison of CNN and LSTM for classifying short musical samples
2023 (English)Independent thesis Basic level (degree of Bachelor), 10 credits / 15 HE credits
Student thesisAlternative title
Jämförelse av CNN och LSTM för instrumentklassificering (Swedish)
Abstract [en]
Applying machine learning to music and audio data is becoming increasingly common. One such area of research is instrument classification, which is the task of identifying the instrument played in a given audio file. In this study, we compared two machine learning model types, LSTM and CNN, on the task of classifying ten different instruments. Additionally, hyperparameter tuning was performed to optimize three parameters used in each model. To test the performance of the models, two test datasets containing instrument samples were used. These were obtained from the NSynth dataset. The first was a randomized equally split dataset based on the NSynth training data. The other one was a curated test dataset created by the NSynth authors. The results for the optimized models showed that the CNN’s accuracy on the randomized dataset was 97.1% and for the LSTM it was 96.0%. The corresponding results for the NSynth test dataset were 74.3% and 71.6%. While the results indicate that the CNN is slightly better, the difference could also be explained by the selected feature extraction technique or more optimized hyperparameters for the CNN. Furthermore, while the accuracy was higher for the CNN on both datasets, a Wilcoxon signed rank test showed that the difference was not statistically significant.
Abstract [sv]
Att använda maskininlärning för musik- och ljuddata blir allt vanligare. Ett exempel på ett sådant forskningsområde är instrument-klassificering. Instrumentklassificering innebär att identifiera vilket instrument som spelas i en ljudfil. I denna studie jämfördes två typer av maskininlärningsmodeller, LSTM och CNN, givet denna uppgift. Därtill utfördes hyperparameteranpassning för att optimera tre parametrar i vardera modell. För att testa modellernas prestanda användes två datamängder tagna från NSynth-datamängden. Den ena var en slumpad datamängd byggd från NSynth-datamängdens träningsdata, medan den andra var en testdatamängd skapad av NSynth-författarna. Resultatet från de optimerade modellerna visade att CNN:en hade en precision på 97.1% på den slumpade datamängden och 74.3% på NSynth-testdatamängden. Motsvarande resultat för LSTM:en var 96.0% och 71.6%. Medan resultaten indikerar att CNN:en presterar bättre kan skillnaden möjligtvis förklaras av den valda metoden för att bearbeta ljuddatat. Alternativt kan det bero på mer optimerade hyperparametrar i CNN-modellen. Dessutom bör det poängteras att CNN:ens högre precision inte är statistiskt säkerställd, även om CNN:en var bättre på båda datamängderna.
Place, publisher, year, edition, pages
2023. , p. 26
Series
TRITA-EECS-EX ; 2023:284
National Category
Computer and Information Sciences
Identifiers
URN: urn:nbn:se:kth:diva-330770OAI: oai:DiVA.org:kth-330770DiVA, id: diva2:1778370
Supervisors
Examiners
2023-07-272023-07-012023-07-27Bibliographically approved