// radar de ia

Multimodal

Papers, modelos e datasets em alta no Hugging Face, além do blog oficial — com leitura editorial em português.

Blog LLMs & Texto

TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging

arXiv:2607.09818v1 Announce Type: new Abstract: Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute corresponding actions as embodied agents. Recently, autoregressive token-based action generation has driven the development of many representative VLA models. However, this paradigm often reduces action generation to next-token prediction, thereby lacking explicit modeling of the spatiotemporal structure of action...

14.07.2026
Blog Multimodal

Source-Lifted Flow Matching for Intervenable Multimodal Imitation

arXiv:2607.10206v1 Announce Type: new Abstract: Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stochasticity is largely passive: repeated sampling may yield diverse behaviors, but users cannot directly choose among valid continuations from the same state. We propose Source-Lifted Flow Matching (SL-FM), a source-intervenable flow-matching policy that exposes such a handle while keeping the velocity field share...

14.07.2026
Blog Robótica & RL

Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models

arXiv:2607.09785v1 Announce Type: new Abstract: Traditionally, continual learning has assumed access to labeled data, yet many real-world applications -- such as lifelong robotics -- require models to adapt continuously from unlabeled streams. This has led to the development of continual self-supervised learning (CSSL), a rapidly growing area that lacks a dedicated, systematic review. In this work, we present a comprehensive survey of CSSL for vision, with connections to emerging vision-language...

14.07.2026
Blog LLMs & Texto

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning

arXiv:2607.10004v1 Announce Type: new Abstract: Unified multimodal models (UMMs) with interleaved reasoning, which generate both textual and visual steps as part of intermediate reasoning traces, have demonstrated great potential for visual mathematical reasoning tasks. However, we identify a key insight in this paradigm: generating intermediate visual reasoning steps is not always beneficial and can even be harmful, as self-generated visual steps may introduce erroneous visual evidence that mis...

14.07.2026
Blog Multimodal

Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation

arXiv:2607.10082v1 Announce Type: new Abstract: Building pixel-level correspondence between event and image data is a fundamental task for multi-sensor systems. However, existing cross-modal matching methods are largely restricted by their reliance on either matching labels or strictly aligned hardware, which limits them to unlabeled and unconstrained real-world scenarios where neither matching ground truth nor prior sensor relationships are available. To address this, we propose a novel two-sta...

14.07.2026
Blog Multimodal

ShapKO: Shapley-Adaptive Modality Knockout for Robust Multimodal Learning

arXiv:2607.09884v1 Announce Type: new Abstract: Multimodal medical models often degrade when inputs are missing, a common scenario in real-world clinical workflows. Separately, even when all modalities are present, modality dominance is observed during training, where optimization over-relies on a highly predictive modality and undertrains complementary sources, resulting in poor robustness under partial availability. While training-time modality knockout improves missing-modality robustness, ex...

14.07.2026
Blog Multimodal

OmniMapBench: Avaliando o Raciocínio Centrado no Visual em Diversos Documentos de Mapas

arXiv:2607.09068v1 Tipo de Anúncio: novo Resumo: Os avanços recentes em LVLMs exigem benchmarks robustos para raciocínio complexo e visualmente fundamentado. Uma limitação crítica é identificada em muitos benchmarks de compreensão de documentos: o conteúdo visual frequentemente pode ser reduzido a texto, permitindo alto desempenho sem uma genuína fundamentação visual. Para enfrentar essa limitação, o OmniMapBench é apresentado para promover o raciocínio centrado no visual em documentos de mapas. O benchmark compreende 2.096 perguntas-respostas anotadas manualmente...

13.07.2026
467 itens no radar