Skip to main navigation Skip to search Skip to main content

A Dynamic Confidence-Weighted Ensemble of Speech and Text Models for Emotion Recognition

  • Changwon National University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Recognizing human emotion in conversation is a complex task because people convey feelings through multiple channels, including vocal tone (prosody) and word choice (lexical content). Systems that rely on a single channel, or modality, such as only speech or only text, often fail to capture the complete emotional context. To address this, we propose the Dynamic Confidence-Weighted Ensemble (DCWE). This novel framework intelligently integrates the outputs from a Speech Emotion Recognition (SER) model and a text-based emotion classifier. Unlike static fusion methods, our approach dynamically adjusts each model's contribution based on two key criteria: the pre-assessed difficulty of each emotion class for the model, and the model's real-time prediction confidence. By leveraging the complementary strengths of both models, DCWE achieves superior performance on multi-party conversational data.

Original languageEnglish
Title of host publication2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331580773
DOIs
StatePublished - 2026
Event2026 International Conference on Electronics, Information, and Communication, ICEIC 2026 - Macau, China
Duration: 18 Jan 202621 Jan 2026

Publication series

Name2026 International Conference on Electronics, Information, and Communication, ICEIC 2026

Conference

Conference2026 International Conference on Electronics, Information, and Communication, ICEIC 2026
Country/TerritoryChina
CityMacau
Period18/01/2621/01/26

Keywords

  • Dynamic Weighting
  • Emotion Recognition
  • Ensemble Method
  • Multimodal Learning
  • Speech Emotion Recognition (SER)

Fingerprint

Dive into the research topics of 'A Dynamic Confidence-Weighted Ensemble of Speech and Text Models for Emotion Recognition'. Together they form a unique fingerprint.

Cite this