Human-centered emotion recognition is becoming a key enabler for innovative computing experiences in context-aware environments. As one of the most natural and unobtrusive forms of human expression, speech provides a powerful channel for inferring human emotions. However, the development of reliable AI-based emotion recognition models depends on the availability of high-quality, well-annotated emotional speech datasets. While several datasets exist in English and other widely spoken languages, resources for Italian remain limited, hindering the development of multilingual and culturally adapted systems. This paper presents AI4SER, a novel open-source dataset for Speech Emotion Recognition (SER) in Italian. AI4SER contains high-quality recordings from multiple speakers with balanced emotional content, covering a wide range of acted emotions. All audio files are provided in WAV format at 44.1 kHz and distributed under a CC BY 4.0 license. We describe the data collection process, corpus statistics, and conduct baseline experiments using deep learning methods. The results demonstrate that AI4SER achieves performance comparable to existing Italian datasets, confirming its suitability as a reliable benchmark for fostering research in emotion recognition and affective computing in Italian.

AI4SER: An Italian Speech Emotion Dataset for Human-Centered Affective Applications

Serrano S.
Primo
;
Amadeo M.;Patane Luca;Scarpa M.;Mento C.;Carbone S.;Esposito G.
2026-01-01

Abstract

Human-centered emotion recognition is becoming a key enabler for innovative computing experiences in context-aware environments. As one of the most natural and unobtrusive forms of human expression, speech provides a powerful channel for inferring human emotions. However, the development of reliable AI-based emotion recognition models depends on the availability of high-quality, well-annotated emotional speech datasets. While several datasets exist in English and other widely spoken languages, resources for Italian remain limited, hindering the development of multilingual and culturally adapted systems. This paper presents AI4SER, a novel open-source dataset for Speech Emotion Recognition (SER) in Italian. AI4SER contains high-quality recordings from multiple speakers with balanced emotional content, covering a wide range of acted emotions. All audio files are provided in WAV format at 44.1 kHz and distributed under a CC BY 4.0 license. We describe the data collection process, corpus statistics, and conduct baseline experiments using deep learning methods. The results demonstrate that AI4SER achieves performance comparable to existing Italian datasets, confirming its suitability as a reliable benchmark for fostering research in emotion recognition and affective computing in Italian.
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11570/3360730
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact