Designing And Integrating Explainable Ai Modules Into Multimodal Emotion Recognition Architectures For Superior Transparency And Interpretability Across Diverse Data Modalities
DOI:
https://doi.org/10.70082/ev8jck77Keywords:
Explainable AI, Multimodal Emotion Recognition, Interpretability, Transparency, Fusion Architectures, Affective Computing.Abstract
Multimodal emotion recognition systems are good at fusing various types of data, such as audio, visual, and physiological signals; however, they may have black-box opacity that limits their trustworthiness in vital applications including healthcare diagnostics and human-computer interaction. This should improve interpretability of cross-modal decisions, but it is hampered by opacity with respect to bias and reliability. This study develops and injects the explainable AI (XAI) modules into the multimodal emotion recognition frameworks for enhanced modality-based transparency and interpretability. We suggest a novel hybrid framework consisting of transformer-based fusion followed by post-hoc XAI techniques, such as SHAP and LIME on the datasets like IEMOCAP and MELD. Example-specific interpretability is achieved through the generation of modality-specific attention maps and counterfactual explanations at inference time. The proposed system uses federated learning with privacy-friendly training, validated by robustness tests in the noisy environment. Experiments show 12% accuracy improvement above baselines (e.g.∼78.5% F1-score) with 85% user-rated interpretability Also, cross-modal attributions show that the model relies more on audio to detect arousal, which improves debugging iteration. The integration of XAI leads to both trustworthy affective computing as well as open paths towards deployable systems in mental health monitoring and edge devices.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
