TY - GEN
T1 - A Dynamic Confidence-Weighted Ensemble of Speech and Text Models for Emotion Recognition
AU - Byeon, Gongkyu
AU - Oh, Beomseok
AU - Yu, Sunjin
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Recognizing human emotion in conversation is a complex task because people convey feelings through multiple channels, including vocal tone (prosody) and word choice (lexical content). Systems that rely on a single channel, or modality, such as only speech or only text, often fail to capture the complete emotional context. To address this, we propose the Dynamic Confidence-Weighted Ensemble (DCWE). This novel framework intelligently integrates the outputs from a Speech Emotion Recognition (SER) model and a text-based emotion classifier. Unlike static fusion methods, our approach dynamically adjusts each model's contribution based on two key criteria: the pre-assessed difficulty of each emotion class for the model, and the model's real-time prediction confidence. By leveraging the complementary strengths of both models, DCWE achieves superior performance on multi-party conversational data.
AB - Recognizing human emotion in conversation is a complex task because people convey feelings through multiple channels, including vocal tone (prosody) and word choice (lexical content). Systems that rely on a single channel, or modality, such as only speech or only text, often fail to capture the complete emotional context. To address this, we propose the Dynamic Confidence-Weighted Ensemble (DCWE). This novel framework intelligently integrates the outputs from a Speech Emotion Recognition (SER) model and a text-based emotion classifier. Unlike static fusion methods, our approach dynamically adjusts each model's contribution based on two key criteria: the pre-assessed difficulty of each emotion class for the model, and the model's real-time prediction confidence. By leveraging the complementary strengths of both models, DCWE achieves superior performance on multi-party conversational data.
KW - Dynamic Weighting
KW - Emotion Recognition
KW - Ensemble Method
KW - Multimodal Learning
KW - Speech Emotion Recognition (SER)
UR - https://www.scopus.com/pages/publications/105034832071
U2 - 10.1109/ICEIC69189.2026.11386277
DO - 10.1109/ICEIC69189.2026.11386277
M3 - Conference contribution
AN - SCOPUS:105034832071
T3 - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
BT - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
Y2 - 18 January 2026 through 21 January 2026
ER -