Speech Emotion Recognition (SER) has achieved significant progress in widely studied languages such as English, driven by the availability of large-scale benchmark datasets. However, research on Italian remains limited due to the scarcity of recent and complementary emotional speech corpora. This paper investigates a feature-based transfer learning approach for Italian SER using audio models pre-trained on AudioSet. Experiments are conducted on two recent Italian corpora with complementary characteristics: AI4SER, recorded in controlled laboratory conditions, and Emozionalmente, developed through crowdsourcing and characterized by higher acoustic variability. A three-scenario evaluation protocol-within-corpus, cross-corpus, and joint training-is adopted to assess robustness to domain shift. Results indicate that embeddings extracted from deeper and higher-capacity models consistently improve performance and robustness. However, significant degradation is observed in cross-corpus settings, highlighting the strong impact of dataset-specific characteristics. Joint training partially mitigates this effect but does not fully eliminate the generalization gap. These findings provide insights into transfer learning effectiveness for Italian SER under heterogeneous recording conditions.

EXPLORING TRANSFER LEARNING FOR SPEECH EMOTION RECOGNITION IN ITALIAN

Serrano S.
Primo
;
Amadeo M.
Secondo
;
Scarpa M.
Penultimo
;
Spinella S.
Ultimo
2026-01-01

Abstract

Speech Emotion Recognition (SER) has achieved significant progress in widely studied languages such as English, driven by the availability of large-scale benchmark datasets. However, research on Italian remains limited due to the scarcity of recent and complementary emotional speech corpora. This paper investigates a feature-based transfer learning approach for Italian SER using audio models pre-trained on AudioSet. Experiments are conducted on two recent Italian corpora with complementary characteristics: AI4SER, recorded in controlled laboratory conditions, and Emozionalmente, developed through crowdsourcing and characterized by higher acoustic variability. A three-scenario evaluation protocol-within-corpus, cross-corpus, and joint training-is adopted to assess robustness to domain shift. Results indicate that embeddings extracted from deeper and higher-capacity models consistently improve performance and robustness. However, significant degradation is observed in cross-corpus settings, highlighting the strong impact of dataset-specific characteristics. Joint training partially mitigates this effect but does not fully eliminate the generalization gap. These findings provide insights into transfer learning effectiveness for Italian SER under heterogeneous recording conditions.
2026
978-3-937436-90-6
978-3-937436-89-0
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11570/3357452
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact