Emotional artificial intelligence via identity-free body language recognition and understanding
Li, Deng (2026-08-04)
Väitöskirja
Li, Deng
04.08.2026
Lappeenranta-Lahti University of Technology LUT
Acta Universitatis Lappeenrantaensis
School of Engineering Science
School of Engineering Science, Laskennallinen tekniikka
Kaikki oikeudet pidätetään.
In reference to IEEE copyrighted material which is used with permission in this thesis, the IEEE does not endorse any of Lappeenranta-Lahti University of Technology LUT's products or services. Internal or personal use of this material is permitted. If interested in reprinting/republishing IEEE copyrighted material for advertising or promotional purposes or for creating new collective works for resale or redistribution, please go to http://www.ieee.org/publications_ standards/publications/rights/rights_link.html to learn how to obtain a License from RightsLink.
In reference to IEEE copyrighted material which is used with permission in this thesis, the IEEE does not endorse any of Lappeenranta-Lahti University of Technology LUT's products or services. Internal or personal use of this material is permitted. If interested in reprinting/republishing IEEE copyrighted material for advertising or promotional purposes or for creating new collective works for resale or redistribution, please go to http://www.ieee.org/publications_ standards/publications/rights/rights_link.html to learn how to obtain a License from RightsLink.
Julkaisun pysyvä osoite on
https://urn.fi/URN:ISBN:978-952-412-480-5
https://urn.fi/URN:ISBN:978-952-412-480-5
Kuvaus
ei tietoa saavutettavuudesta
Tiivistelmä
Emotional Artificial Intelligence (EAI) refers to an Artificial Intelligence (AI) system’s ability to perceive, interpret, and respond to human emotions. EAI is essential for humancentered applications such as healthcare, education, and human–computer interaction. However, mainstream emotion recognition systems depend heavily on identity-sensitive signals including facial expressions, raw speech, and biometric data, raising privacy concerns in practice. This dissertation investigates an alternative paradigm: privacy-preserving emotional artificial intelligence through Identity-Free Body Language (IFBL). Specifically, it pursues four objectives to validate this paradigm: effective IFBL perception, efficient IFBL perception, de-identified emotion dataset construction, and Multimodal Large Language Model (MLLM)-based emotion understanding.
First, this dissertation develops an effective IFBL recognition method based on visual-text contrastive learning. The proposed method achieves 66.12% and 65.08% top-1 accuracy on the iMiGUE and SMG benchmarks, respectively, outperforming all previous methods. Ablation studies show that visual-text contrastive learning contributes to a 26.43% improvement over a vision-only baseline. Second, this dissertation proposes the Motion-aware State Fusion Mamba (MSF-Mamba) for efficient IFBL recognition. MSF-Mamba enhances state-space models with a Multiscale Central Frame Difference State Fusion Module that injects motion cues into the state transition. MSF-Mamba achieves state-of-the-art (SoTA) performance on iMiGUE and SMG, surpassing all CNN-, Transformer-, and SSM-based baselines while maintaining high computational efficiency. Third, this dissertation introduces the task of De-identified Multimodal Emotion Recognition and Reasoning (D-MERR) and constructs the corresponding DEEMO dataset. Fourth, this dissertation develops DEEMO-LLaMA, an MLLM that integrates de-identified video, de-identified audio, and transcriptions for emotion recognition and reasoning. DEEMOLLaMA achieves 74.49% accuracy, outperforming the best baseline by 9.75% in accuracy.
In summary, these contributions establish a complete research pipeline. This dissertation demonstrates that privacy-preserving emotional artificial intelligence grounded in identity-free body language is both feasible and effective.
First, this dissertation develops an effective IFBL recognition method based on visual-text contrastive learning. The proposed method achieves 66.12% and 65.08% top-1 accuracy on the iMiGUE and SMG benchmarks, respectively, outperforming all previous methods. Ablation studies show that visual-text contrastive learning contributes to a 26.43% improvement over a vision-only baseline. Second, this dissertation proposes the Motion-aware State Fusion Mamba (MSF-Mamba) for efficient IFBL recognition. MSF-Mamba enhances state-space models with a Multiscale Central Frame Difference State Fusion Module that injects motion cues into the state transition. MSF-Mamba achieves state-of-the-art (SoTA) performance on iMiGUE and SMG, surpassing all CNN-, Transformer-, and SSM-based baselines while maintaining high computational efficiency. Third, this dissertation introduces the task of De-identified Multimodal Emotion Recognition and Reasoning (D-MERR) and constructs the corresponding DEEMO dataset. Fourth, this dissertation develops DEEMO-LLaMA, an MLLM that integrates de-identified video, de-identified audio, and transcriptions for emotion recognition and reasoning. DEEMOLLaMA achieves 74.49% accuracy, outperforming the best baseline by 9.75% in accuracy.
In summary, these contributions establish a complete research pipeline. This dissertation demonstrates that privacy-preserving emotional artificial intelligence grounded in identity-free body language is both feasible and effective.
Kokoelmat
- Väitöskirjat [1215]
