This paper investigates the requirements for intelligible high-quality motion capture of sign language, with a particular focus on face animation. There are several methods for capturing facial animation, yet there is no established standard for which type of facial animation is the most optimal for sign languages. We compare two facial animation approaches-ARKit and Metahuman Animator (MHA) in terms of intelligibility. As expected, a human evaluation study with deaf Swedish Sign Language signers showed that MHA outperforms ARKit because MHA has a higher number of facial controls and an advanced depth data solver. In addition, we introduce a biLSTM-based occlusion infilling technique for MHA data as opposed to linear interpolation. Although biLSTM infilled MHA animations showed no improvement in human evaluation compared to original MHA animations, the model produces linguistically reasonable infillings upon qualitative exploration of the renders.
Part of ISBN 9798400719967
QC 20260311