Digital Twin (DT) technologies are increasingly explored in healthcare as a means to model users' physical and behavioral states. However, their integration with affective intelligence, including Speech Emotion Recognition (SER), remains largely unexplored, despite the potential of voice as a natural and non-intrusive source of emotional information. At the same time, training SER models in healthcare contexts raises privacy concerns, as speech data may reveal sensitive information. To address these challenges, this paper proposes a novel framework for emotion-aware DTs that actively participate in a cross-silo Federated Learning (FL) process, collaboratively training a speech emotion classifier without sharing raw data. Each DT acts as a long-lived digital counterpart of a healthcare institution, maintaining local affective profiles extracted from speech interactions, and contributing to a shared global SER model that is refined through FL aggregation. Experimental results demonstrate that the federated DT-based approach consistently improves model generalization across linguistic and institutional boundaries, outperforming isolated local training while ensuring data confidentiality and local data ownership.

Federated Learning for Digital Twin-assisted Speech Emotion Recognition

Serrano S.;Amadeo M.
2026-01-01

Abstract

Digital Twin (DT) technologies are increasingly explored in healthcare as a means to model users' physical and behavioral states. However, their integration with affective intelligence, including Speech Emotion Recognition (SER), remains largely unexplored, despite the potential of voice as a natural and non-intrusive source of emotional information. At the same time, training SER models in healthcare contexts raises privacy concerns, as speech data may reveal sensitive information. To address these challenges, this paper proposes a novel framework for emotion-aware DTs that actively participate in a cross-silo Federated Learning (FL) process, collaboratively training a speech emotion classifier without sharing raw data. Each DT acts as a long-lived digital counterpart of a healthcare institution, maintaining local affective profiles extracted from speech interactions, and contributing to a shared global SER model that is refined through FL aggregation. Experimental results demonstrate that the federated DT-based approach consistently improves model generalization across linguistic and institutional boundaries, outperforming isolated local training while ensuring data confidentiality and local data ownership.
File in questo prodotto:
Non ci sono file associati a questo prodotto.
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11570/3360732
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
social impact