Building a Multimodal Dataset of Academic Paper for Keyword Extraction
arXiv:2606.31069v1 Announce Type: new Abstract: Up to this point, keyword extraction task typically relies solely on textual data. Neglecting visual details and audio features from image and audio modalities leads to deficiencies in information richness and overlooks potential correlations, thereby constraining the model's ability to learn representations of the data and the accuracy of model predictions. Furthermore, the currently available multimodal datasets for keyword extraction task are pa...
arXiv cs.CL
·Jingyu Zhang, Xinyi Yan, Yi Xiang, Yingyi Zhang, Chengzhi Zhang
·
// relacionados
Leia também
Modelo
nvidia/NVIDIA-NemotronLabs-VoiceChat-11B
Blog
Detecção de Autoapresentações em Depoimentos Legislativos
Blog
"Muitos São os Meus Nomes": A Anatomia do Assistente e Suas Personas por Meio de Autoencoders Esparsos
Blog