SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts
arXiv:2607.06620v1 Announce Type: new Abstract: Recent Multimodal Large Language Models (MLLMs) struggle to bridge the representational gap between 2D semantic understanding and 3D spatial geometry. Existing 3D-aware models either rely on costly 3D-specific data or utilize RGB-only inputs with heuristic sampling and monolithic, shallow fusion, which respectively disrupt essential spatiotemporal connectivity and induce modality contention across diverse spatial tasks. To overcome these bottleneck...
arXiv cs.CV
·Haida Feng, Hao Wei, Haolin Wang, Shiwei Li, Chade Li, Yihong Wu
·
// relacionados
Leia também
Blog
Um Guia de Programação para a Programação de GPU Baseada em Tiles da NVIDIA: De cuTile e Kernels Triton até Flash Attention
Blog
OpenAI's GPT-5.6 Sol Ultra reportedly solves a 50-year-old math problem in under an hour
Blog
Grupos terroristas estão usando todos os principais chatbots de IA para planejamento de ataques e desenvolvimento de armas
Blog