Speech Emotion Recognition (SER) has achieved significant progress in widely studied languages such as English, driven by the availability of large-scale benchmark datasets. However, research on Italian remains limited due to the scarcity of recent and complementary emotional speech corpora. This paper investigates a feature-based transfer learning approach for Italian SER using audio models pre-trained on AudioSet. Experiments are conducted on two recent Italian corpora with complementary characteristics: AI4SER, recorded in controlled laboratory conditions, and Emozionalmente, developed through crowdsourcing and characterized by higher acoustic variability. A three-scenario evaluation protocol-within-corpus, cross-corpus, and joint training-is adopted to assess robustness to domain shift. Results indicate that embeddings extracted from deeper and higher-capacity models consistently improve performance and robustness. However, significant degradation is observed in cross-corpus settings, highlighting the strong impact of dataset-specific characteristics. Joint training partially mitigates this effect but does not fully eliminate the generalization gap. These findings provide insights into transfer learning effectiveness for Italian SER under heterogeneous recording conditions.
EXPLORING TRANSFER LEARNING FOR SPEECH EMOTION RECOGNITION IN ITALIAN
Serrano S.
Primo
;Amadeo M.Secondo
;Scarpa M.Penultimo
;Spinella S.Ultimo
2026-01-01
Abstract
Speech Emotion Recognition (SER) has achieved significant progress in widely studied languages such as English, driven by the availability of large-scale benchmark datasets. However, research on Italian remains limited due to the scarcity of recent and complementary emotional speech corpora. This paper investigates a feature-based transfer learning approach for Italian SER using audio models pre-trained on AudioSet. Experiments are conducted on two recent Italian corpora with complementary characteristics: AI4SER, recorded in controlled laboratory conditions, and Emozionalmente, developed through crowdsourcing and characterized by higher acoustic variability. A three-scenario evaluation protocol-within-corpus, cross-corpus, and joint training-is adopted to assess robustness to domain shift. Results indicate that embeddings extracted from deeper and higher-capacity models consistently improve performance and robustness. However, significant degradation is observed in cross-corpus settings, highlighting the strong impact of dataset-specific characteristics. Joint training partially mitigates this effect but does not fully eliminate the generalization gap. These findings provide insights into transfer learning effectiveness for Italian SER under heterogeneous recording conditions.Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


