Audio detection using machine learning & transfer learning models

Acar, Mesut; Dağ, Hasan

Audio detection using machine learning & transfer learning models

Files

704061.pdf (2.11 MB)

Date

2021

Authors

Acar, Mesut

Dağ, Hasan

Publisher

Kadir Has Üniversitesi

Organizational Units

Organizational Unit

Management Information Systems

Abstract

In this paper, using datasets ESC-50 & ESC-10 of environmental sounds, machine learning algorithms, and feature extraction methods are used to develop recognition performance. K-NN, SVM, Random Forest are used for comparing the recognition results. The different feature extraction methods in the literature are used to get more meaningful attributes from these datasets and obtain a higher accuracy rate. This approach shows that SVM algorithm has a significantly good result with accuracy scores. The best accuracy scores obtained by classic machine learning algorithms are %42,15 for ESC-50 and %77,7 for ESC-10. In addition to this, the experiments have been done with a pre-trained ResNet neural network as a backbone, which achieves successful results despite the machine learning models. In this study, a higher accuracy rate is achieved from baseline machine learning algorithms in literature and using transfer learning with pre-trained Resnet backbones to reach some state of art results. The accuracy scores are %68,95 for ESC-50 and %87,25 for ESC-10.
Bu çalışmada çevre seslerinden oluşan ESC-50 ve ESC-10 veri seti, çeşitli makine öğrenmesi, transfer öğrenme altyapısı ve farklı öznitelik çıkarımı yöntemleri kullanarak sınıflandırma çalışmaları yapılmıştır. K-NN, SVM, Rastgele Orman makine öğrenimi algoritmaları kullanılmıştır. Farklı öznitelik çıkarım algoritmaları kullanılarak, bu veri seti için makine öğrenmesi algoritmalarında farklı sonuçlar elde edilmiştir. Bu yaklaşımda SVM algoritmasın gözle görülür bir şekilde performansının attığı gözlemlenmiştir. Klasik makine öğrenmesi algoritmaları ile elde edilen en iyi doğruluk puanları ESC-50 için %42,15 ve ESC-10 için %77,7'dir. Buna ek olarak, makine öğrenmesi modellerinden daha başarılı sonuçlar elde eden, omurga olarak önceden eğitilmiş bir ResNet sinir ağı ile deneyler yapılmıştır. Yapılan deneylerde, literatürdeki temel makine öğrenmesi algoritmalarından ve literatürdeki iyi sonuçlara ulaşmak için önceden eğitilmiş Resnet omurgaları ile transfer öğrenmesi kullanılarak daha yüksek bir doğruluk oranı elde edilmiştir. Resnet algoritması ile ESC-50 için %68,95, ESC-10 için ise %87,25 doğruluk oranı elde edilmiştir.

Keywords

:Bilgisayar Mühendisliği Bilimleri-Bilgisayar ve Kontrol, Computer Engineering and Computer Science and Control, Bilim ve Teknoloji, Science and Technology, Makine öğrenmesi, Machine learning, Makine öğrenmesi yöntemleri, Machine learning methods, Yapay zeka, Artificial intelligence

URI

https://hdl.handle.net/20.500.12469/4283

Collections

Tez Koleksiyonu

Full item page

Audio detection using machine learning & transfer learning models

Files

Date

Authors

Journal Title

Journal ISSN

Volume Title

Publisher

Open Access Color

OpenAIRE Downloads

OpenAIRE Views

Research Projects

Organizational Units

Journal Issue

Events

Abstract

Description

Keywords

Turkish CoHE Thesis Center URL

Fields of Science

Citation

WoS Q

Scopus Q

Source

Volume

Issue

Start Page

End Page

URI

Collections