King's College London

Research portal

Exploring Zero-Shot Emotion Recognition in Speech Using Semantic-Embedding Prototypes

Research output: Contribution to journalArticlepeer-review

Xinzhou Xu, Jun Deng, Nicholas Cummins, Zixing Zhang, Li Zhao, Bjorn W. Schuller

Original languageEnglish
JournalIEEE TRANSACTIONS ON MULTIMEDIA
DOIs
Accepted/In press2021

Bibliographical note

Publisher Copyright: IEEE Copyright: Copyright 2021 Elsevier B.V., All rights reserved.

King's Authors

Abstract

Speech Emotion Recognition (SER) makes it possible for machines to perceive affective information. Our previous research differed from conventional SER endeavours in that it focused on recognising unseen emotions in speech autonomously through machine learning. Such a step would enable the automatic leaning of unknown emerging emotional states. This type of learning framework, however, still relied on manual annotations to obtain multiple samples of each emotion. In order to reduce this additional workload, herein, we propose a zero-shot SER framework employing a per-emotion semantic-embedding paradigm to describe emotions in zero-shot SER, instead of using the sample-wise descriptors. Aiming to optimise the relationship between emotions, prototypes, and speech samples, this framework includes two types of learning strategies: Sample-wise learning and emotion-wise learning. These strategies apply a novel learning process to speech samples and emotions, respectively, via specifically designed semantic-embedding prototypes. We verify the utility of these approaches by performing an extensive experimental evaluation on two corpora on three aspects, namely the influence of different types of learning strategies, emotional-pair comparison, and the selections of semantic-embedding prototypes and paralinguistic features. The experimental results indicate that it is applicable to use semantic-embedding prototypes for zero-shot emotion recognition in speech, despite the influence of choosing optimal strategies and prototypes.

View graph of relations

© 2020 King's College London | Strand | London WC2R 2LS | England | United Kingdom | Tel +44 (0)20 7836 5454