Human Object Interaction Detection in Paintings using Multi-Task Learning

dc.contributor.advisorAsmar, Daniel
dc.contributor.authorAntoun, Maya
dc.contributor.commembersAbou Ghali, Kamel
dc.contributor.commembersH. ElHajj, Imad
dc.contributor.commembersMetni, Najib
dc.contributor.commembersTekli, Joe
dc.contributor.degreePhD
dc.contributor.departmentDepartment of Mechanical Engineering
dc.contributor.facultyMaroun Semaan Faculty of Engineering and Architecture
dc.contributor.institutionAmerican University of Beirut
dc.date2023
dc.date.accessioned2023-07-18T05:48:41Z
dc.date.available2023-07-18T05:48:41Z
dc.date.issued2023-07-17T21:00:00Z
dc.date.submitted2023-07-05T21:00:00Z
dc.description.abstractHuman Object Interaction (HOI) detection provides valuable insights into the meaning and interpretation of a painting, as the interactions between humans and object reveal information about the scene, characters, and story depicted in the artwork. Automatically detecting HOI in paintings is a challenging task, as the paintings often contain complex scenes with intricate details and variations in artistic style. Additionally, unlike in real-world images, the context and physics of the painting may not follow physical rules, which can further complicate the detection process. The proposed system addresses the complexities of this task, considering the intricate details and variations in artistic style found in paintings. It incorporates a model that captures discriminative information by extracting visual features from detected humans, objects, and the Region of Interest. The model analyzes spatial arrangements to understand the relationships and interactions between elements. Moreover, the model integrates contextual knowledge and semantic relationships using a knowledge graph based on Graph Convolution Network to capture the underlying meaning and story depicted in artwork. However, relying solely on appearance and context may not be enough to accurately infer HOIs in paintings. To overcome this challenge, multitask learning is employed by introducing four supplementary classification tasks. These tasks provide complementary information that enhances the HOI detection process, leveraging shared representations across multiple tasks. The proposed system introduces the SemArt-HOI benchmark dataset, augmenting the SemArt dataset with instance detection annotations and interaction classes. Experimental results demonstrate that the proposed model outperforms the state-of-the-art one-stage transformer-based HOI detection model in both single-task and multi-task settings by 1.19% and 1.51% respectively. Furthermore, the system exhibits superior efficiency, training four times faster and requiring fewer resources. This makes it suitable for practical and large-scale HOI detection in paintings.
dc.identifier.urihttp://hdl.handle.net/10938/24097
dc.language.isoen
dc.subjectHuman Object Interaction Detection
dc.subjectComputer Vision
dc.subjectDeep learning
dc.subjectMulti-Task Learning
dc.titleHuman Object Interaction Detection in Paintings using Multi-Task Learning
dc.typeDissertation
local.AUBID200801873

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
AntounMaya_2023.pdf
Size:
20.4 MB
Format:
Adobe Portable Document Format
Description:

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.65 KB
Format:
Item-specific license agreed upon to submission
Description: