Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
arXiv:2607.07294v1 Announce Type: new Abstract: Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where the robot must anticipate conversational dynamics rather than merely react to pauses. This work presents a Multimodal Voice Activity Projection (MM-VAP) framework that extends the original audio-only VAP formulation to synchronized audio-visual inputs while preserving its self-supervised future-projection objec...
arXiv cs.RO
·Antonio Cano, Guillermo P\'erez, Luis Merino, Randy Gomez
·
// relacionados
Leia também
Blog
Meta ran ads for an app promising to nudify female politicians
Editorial
MiniMax Music 3: uma canção inteira de cinco minutos, com pesos abertos
Blog
Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
Blog