// radar de ia

Multimodal

Papers, modelos e datasets em alta no Hugging Face, além do blog oficial — com leitura editorial em português.

Blog LLMs & Texto

Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning

arXiv:2607.02588v1 Announce Type: new Abstract: Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed incrementally under limited memory. Existing online methods either retain compact visual representations that lack semantic structure, or build higher-level memory stores organized around temporal proximity rather than explicit causal links, leaving multi-hop narrative reasoning to be reconstructed by the LLM at ev...

07.07.2026
Blog Multimodal

MAGE: View-guided Point Cloud Completion with Efficient Modality Alignment and Adaptive Geometry Enhancement

arXiv:2607.02568v1 Announce Type: new Abstract: View-based point cloud completion aims to recover a complete 3D shape from a partial point cloud, guided by a single-view image. However, existing approaches often suffer from limited performance due to weak modality alignment and limited self-geometry enhancement. To overcome these challenges, we propose a unified geometry-aware framework that integrates efficient modality alignment and adaptive geometry enhancement, mainly to address cross-modal ...

07.07.2026
Blog Robótica & RL

Exp2VLA: Enabling Vision-Language-Action for Drone Navigation from Expert Demonstrations

arXiv:2607.03146v1 Announce Type: new Abstract: Vision-language-action (VLA) models open a new path toward intuitive robot control by directly linking perception, language, and action in a single end-to-end framework. Yet for UAVs, practical adoption remains difficult because existing solutions are either computationally heavy or insufficiently capable in complex environments. In this work, we propose a practical expert-distillation pipeline (Exp2VLA) for language-conditioned drone navigation. T...

07.07.2026
Blog Multimodal

Feeling the Unexpected: ResTacVLA for Contact-Rich Manipulation via Residual Tactile Representation

arXiv:2607.03387v1 Announce Type: new Abstract: Tactile perception is indispensable for contact-rich manipulation, yet integrating it into Vision-Language-Action (VLA) models often induces modality collapse, where high-bandwidth visual features overshadow sparse tactile cues. Inspired by Predictive Coding, a neural mechanism where the brain attenuates predictable inputs to prioritize surprising stimuli, we propose ResTacVLA. Rather than treating tactile data as raw input, we reformulate it as a ...

07.07.2026
Blog Multimodal

Reliability-Aware CT-MRI Registration: A Quality Engineering Framework with Stability Analysis and Risk Classification

arXiv:2607.02585v1 Announce Type: new Abstract: Multimodal CT-MRI registration is central to image-guided radiotherapy, surgical navigation, and diagnostic workflows, but most pipelines report only aggregate quality metrics without per-case reliability signals. We propose a reliability-aware framework that converts registration quality into Green/Yellow/Red risk categories using data-learned thresholds. CT images were registered to T1-weighted MRI using rigid and affine transformations on 90 pai...

07.07.2026
Blog LLMs & Texto

H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation

arXiv:2607.02592v1 Announce Type: new Abstract: On-policy distillation (OPD) has recently emerged as an effective post-training paradigm by providing supervision on student-generated trajectories. However, existing OPD methods for multimodal reasoning usually rely on a static teacher routing, assigning each sample to a single teacher based on modality or task type. This ignores that visual grounding and abstract reasoning may dominate different decoding steps, making a single teacher insufficien...

07.07.2026
Blog Áudio & Voz

Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion

arXiv:2607.02862v1 Announce Type: new Abstract: Automatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit significant dialectal differences. Existing methods often optimize ASR or DID individually, resulting in performance trade-offs. In this work, we propose a multimodal framework that jointly improves ASR and DID. Our method employs a Bottleneck Encoder to extract dialectal features from Conformer-based spee...

07.07.2026
Blog LLMs & Texto

MentalThink: Shaping Thoughts in Mental SVG World

arXiv:2607.03530v1 Announce Type: new Abstract: We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink is a think-with-SVG pipeline, where the model learns to generate, render, and interpret scalable vector graphics (SVG) code as an intermediate visual representation for multi-turn reasoning. By creating structured vector sketches, the model can externalize spatial hypothe...

07.07.2026
467 itens no radar