Violence detection in audio: evaluating the effectiveness of deep learning models and data augmentation

doi:10.9781/ijimai.2023.08.007

Utilize este identificador para referenciar este registo: https://hdl.handle.net/1822/89936

Registo completo

Campo DC	Valor	Idioma
dc.contributor.author	Durães, Dalila	por
dc.contributor.author	Veloso, Bruno	por
dc.contributor.author	Novais, Paulo	por
dc.date.accessioned	2024-03-25T10:57:35Z	-
dc.date.available	2024-03-25T10:57:35Z	-
dc.date.issued	2023	-
dc.identifier.issn	1989-1660	-
dc.identifier.uri	https://hdl.handle.net/1822/89936	-
dc.description.abstract	Human nature is inherently intertwined with violence, impacting the lives of numerous individuals. Various forms of violence pervade our society, with physical violence being the most prevalent in our daily lives. The study of human actions has gained significant attention in recent years, with audio (captured by microphones) and video (captured by cameras) being the primary means to record instances of violence. While video requires substantial processing capacity and hardware-software performance, audio presents itself as a viable alternative, offering several advantages beyond these technical considerations. Therefore, it is crucial to represent audio data in a manner conducive to accurate classification. In the context of violence in a car, specific datasets dedicated to this domain are not readily available. As a result, we had to create a custom dataset tailored to this particular scenario. The purpose of curating this dataset was to assess whether it could enhance the detection of violence in car-related situations. Due to the imbalanced nature of the dataset, data augmentation techniques were implemented. Existing literature reveals that Deep Learning (DL) algorithms can effectively classify audio, with a commonly used approach involving the conversion of audio into a mel spectrogram image. Based on the results obtained for that dataset, the EfficientNetB1 neural network demonstrated the highest accuracy (95.06%) in detecting violence in audios, closely followed by EfficientNetB0 (94.19%). Conversely, MobileNetV2 proved to be less capable in classifying instances of violence.	por
dc.description.sponsorship	FCT - Fundação para a Ciência e a Tecnologia(UIDB/00319/2020)	por
dc.language.iso	eng	por
dc.publisher	Universidad Internacional de La Rioja (UNIR)	por
dc.relation	info:eu-repo/grantAgreement/FCT/6817 - DCRRNI ID/UIDB%2F00319%2F2020/PT	por
dc.rights	openAccess	por
dc.subject	Audio	por
dc.subject	Deep learning	por
dc.subject	Human action recognition	por
dc.subject	Machine learning	por
dc.subject	Transfer learning	por
dc.subject	Violence detection in a car	por
dc.title	Violence detection in audio: evaluating the effectiveness of deep learning models and data augmentation	por
dc.type	article	por
dc.peerreviewed	yes	por
oaire.citationStartPage	72	por
oaire.citationEndPage	84	por
oaire.citationIssue	3	por
oaire.citationVolume	8	por
dc.date.updated	2024-03-15T12:01:22Z	-
dc.identifier.doi	10.9781/ijimai.2023.08.007	por
sdum.export.identifier	13512	-
sdum.journal	International Journal of Interactive Multimedia and Artificial Intelligence	por
Aparece nas coleções:	CAlg - Artigos em revistas internacionais / Papers in international journals

Ficheiros deste registo:

Ficheiro	Descrição	Tamanho	Formato
IJIMAI.pdf		2,09 MB	Adobe PDF	Ver/Abrir

Ver registo simples Sugerir correção Estatísticas

Citations

Altmetrics