Published January 1, 2026
| Version v1
Journal article
Open
Enhancing Speech Synthesis With Human-Like Emotional Intelligence For Natural And Expressive Communication
Authors/Creators
Description
This paper presents an emotion-aware voice-based conversational therapy assistant that integrates speech recognition, con-versational AI, and emotional text-to-speech synthesis into a unified pipeline. The system captures user speech through a microphone, transcribes it to text, generates context-aware empathetic responses using a large language model (Gemini AI), and synthesizes emotion-ally expressive speech output using IndexTTS2 with zero-shot voice cloning. The architecture follows a modular design comprising four major modules: Voice Input, Processing and AI, Emotion Analysis, and Speech Synthesis. The emotion mapping subsystem identifies user affect and selects an appropriate response emotion to guide TTS output. Evaluation against two baselines (generic neutral TTS and rule-based keyword approach) demonstrates that the proposed model achieves the highest overall score of 74.51, significantly outper-forming both baselines in holistic end-to-end quality. The system balances emotion recognition accuracy, response relevance, and audio naturalness, making it suitable for mental health support, virtual assistants, and human-centered AI applications. The results confirm that combining emotional conditioning with contextual response generation yields substantially better conversational quality than neutral or rule-driven approaches.
Files
IJSRET_V12_issue2_609.pdf
Files
(323.8 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:a0763ad72dabc5b6c549dad71e29fd95
|
323.8 kB | Preview Download |
Additional details
Related works
- Has part
- Journal article: https://ijsret.com/wp-content/uploads/IJSRET_V12_issue2_609.pdf (URL)
- Is identical to
- Journal article: https://ijsret.com/2026/05/06/enhancing-speech-synthesis-with-human-like-emotional-intelligence-for-natural-and-expressive-communication/ (URL)