End-to-end listening agent for audiovisual emotional and naturalistic interactions

dc.contributor.authorEl Haddad, Kevin
dc.contributor.authorRizk, Yara
dc.contributor.authorHeron, Louise
dc.contributor.authorHajj, Nadine
dc.contributor.authorZhao, Yong
dc.contributor.authorKim, Jaebok
dc.contributor.authorTrung, Ngô Trọng
dc.contributor.authorLee, Minha
dc.contributor.authorDoumit, Marwan
dc.contributor.authorLin, Payton
dc.contributor.authorKim, Yelin
dc.contributor.authorÇakmak, Hüseyin
dc.contributor.departmentDepartment of Electrical and Computer Engineering
dc.contributor.facultyMaroun Semaan Faculty of Engineering and Architecture (MSFEA)
dc.contributor.institutionAmerican University of Beirut
dc.date.accessioned2025-01-24T11:29:33Z
dc.date.available2025-01-24T11:29:33Z
dc.date.issued2018
dc.description.abstractIn this work, we established the foundations of a framework with the goal to build an end-to-end naturalistic expressive listening agent. The project was split into modules for recognition of the user’s paralinguistic and nonverbal expressions, prediction of the agent’s reactions, synthesis of the agent’s expressions and data recordings of nonverbal conversation expressions. First, a multimodal multitask deep learning-based emotion classification system was built along with a rule-based visual expression detection system. Then several sequence prediction systems for nonverbal expressions were implemented and compared. Also, an audiovisual concatenation-based synthesis system was implemented. Finally, a naturalistic, dyadic emotional conversation database was collected. We report here the work made for each of these modules and our planned future improvements. © 2018, Universidade Catolica Portuguesa. All rights reserved.
dc.identifier.doihttps://doi.org/10.7559/citarj.v10i2.424
dc.identifier.eid2-s2.0-85073362795
dc.identifier.urihttp://hdl.handle.net/10938/27254
dc.language.isoen
dc.publisherUniversidade Catolica Portuguesa
dc.relation.ispartofJournal of Science and Technology of the Arts
dc.sourceScopus
dc.subjectDyadic conversation database
dc.subjectEmotion database
dc.subjectEyebrow movement
dc.subjectHead movement
dc.subjectLaughter
dc.subjectListening agent
dc.subjectMultimodal synthesis
dc.subjectNonverbal expression detection
dc.subjectNonverbal expression synthesis
dc.subjectSequence-to-sequence prediction systems
dc.subjectSmile
dc.subjectSpeech emotion recognition
dc.titleEnd-to-end listening agent for audiovisual emotional and naturalistic interactions
dc.typeArticle

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
2018-4348.pdf
Size:
834.45 KB
Format:
Adobe Portable Document Format