// radar de ia

Multimodal

Papers, modelos e datasets em alta no Hugging Face, além do blog oficial — com leitura editorial em português.

Blog LLMs & Texto

IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation

arXiv:2607.25106v1 Announce Type: new Abstract: Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing approaches typically depend on text-only queries, which become less reliable as semantic specificity increases toward fine-grained object categories. We introduce IMPRINT, a zero-shot plug-and-play framework that enriches textual object queries with web-sourced images to i...

29.07.2026
Blog LLMs & Texto

FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding

arXiv:2607.25266v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of relevant content decreases sharply as video sequence length increases, and exposing the model to more irrelevant content measurably reduces its accuracy. In this paper, we address the problem of maximizing query-relevant information in a frame subset selected at inference time, without training. FO...

29.07.2026
Blog Geração de Imagem

MorphUNet: Alpha-Controlled Biometric Transport for Diffusion-Based Face Morphing Attacks

arXiv:2607.25092v1 Announce Type: new Abstract: Face morphing attacks create synthetic images verifiable against multiple identities, threatening border control and identity verification systems. We introduce MorphUNet, a diffusion morphing framework formulating two-parent generation as alpha-controlled biometric transport: each parent is decomposed into CLIP appearance and ArcFace identity evidence, aligned into a CLIP-compatible token space, with the two contributors preserved as separate iden...

29.07.2026
Blog LLMs & Texto

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

arXiv:2607.24904v1 Announce Type: new Abstract: Standard vision-language models (VLMs) suffer from Moravec's paradox: they excel at complex offline visual reasoning but struggle with simple streaming perception tasks and process them inefficiently. We present Mage-VL, an efficient codec-native streaming foundation model for real-time multimodal understanding and interaction. At its core, our custom tokenizer, Mage-ViT, replaces uniform frame sampling by selectively encoding dynamic, entropy-rich...

29.07.2026
Blog Multimodal

CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition

arXiv:2607.25294v1 Announce Type: new Abstract: Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context to be learned from is multimodal: scientific findings are conveyed through figures and tables, financial indicators are scattered across converted repor...

29.07.2026
Blog LLMs & Texto

Raciocinando com Memória: Uma Estrutura Adaptativa à Granularidade Temporal para Compreensão de Vídeos Longos Sem Treinamento

arXiv:2607.24794v1 Tipo de Anúncio: novo Resumo: Embora os Modelos de Linguagem de Grande Escala Multimodais (MLLMs) demonstrem generalização superior em tarefas fundamentais de vídeo, janelas de contexto restritas limitam sua compreensão de vídeos longos. Para acomodar essa restrição, os modelos normalmente recorrem à seleção de quadros-chave. No entanto, a amostragem uniforme ou a seleção estática guiada por consulta frequentemente ignora o contexto temporal crítico, falhando em se adaptar às diferentes granularidades temporais das consultas. Neste artigo, propomos o ReMem, ...

29.07.2026
Blog Dados & Embeddings

Neurai-VN Benchmark: Standardized Machine Learning Models for Multimodal Digital Phenotyping in Mental Health Classification

arXiv:2607.25232v1 Announce Type: new Abstract: Digital phenotyping (DP) using smartphones and wearable devices has shown considerable potential for mental health monitoring. However, progress remains difficult to evaluate due to heterogeneous datasets, inconsistent preprocessing pipelines. In this study, we present a reproducible benchmark built upon the Neurai-VN dataset, a high-resolution, multimodal dataset comprising passive sensing and active assessment from wearable and smartphone devices...

29.07.2026
Blog Multimodal

Room-Mediated Co-occurrence for Zero-Shot Object-Centric Semantic Navigation via Frontier Scoring

arXiv:2607.25448v1 Announce Type: new Abstract: Zero-shot ObjectNav methods increasingly use vision-language priors, but direct object-object similarity in the latent space is often a weak proxy for spatial co-occurrence. We present an analytical, training-free semantic navigation pipeline that mediates object relationships through a compact room lexicon. Each object label is mapped to a CLIP-derived Room Probability Vector (RPV), and object-target co-occurrence is computed from RPV distribution...

29.07.2026
Blog Robótica & RL

SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models

arXiv:2607.25912v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-language backbones and lack fine-grained 3D understanding of target objects, especially under occlusion, pose variation, scale changes, and precise spatial interaction. We propose an object-centric 3D representation alignment framework built upon $\pi_0$, using SAM3D as a frozen 3D teacher to provide target-obje...

29.07.2026
Blog LLMs & Texto

ProcAgent: Um framework de agentes para orientação em tarefas procedurais na borda com humano no circuito

arXiv:2607.24770v1 Tipo de anúncio: novo Resumo: Tarefas procedurais como montagem de móveis e reparos domésticos impõem demandas cognitivas substanciais, pois os usuários precisam interpretar instruções, acompanhar o progresso da tarefa, raciocinar sobre o estado espacial e recuperar-se de erros enquanto realizam ações físicas. Assistentes multimodais anteriores mostraram-se promissores para orientação procedural, mas a maioria depende de inferência na nuvem e de percepção fixa sempre ativa, o que os torna pouco adequados para do... [sensíveis à privacidade, críticos quanto à latência]

29.07.2026
Blog LLMs & Texto

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization

arXiv:2607.24856v1 Announce Type: new Abstract: Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocalization challenging. To address this issue, we propose DisasterTD, a disaster toponym disambigu...

29.07.2026
Blog LLMs & Texto

MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios

arXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop MyoCardBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: MyoCardBench includes 2,263 items from 13 task-sp...

29.07.2026
467 itens no radar