// radar de ia

Multimodal

Papers, modelos e datasets em alta no Hugging Face, além do blog oficial — com leitura editorial em português.

Blog Multimodal

Trajectory-aware Cross-view Geo-localization with Sequential Observations

arXiv:2607.15491v1 Announce Type: new Abstract: Cross-view geo-localization matches ground-level observations against geo-tagged satellite imagery. Recent methods show that sequential queries such as video clips yield richer spatiotemporal cues than single images, yet they overlook a complementary sequential modality: route descriptions -- which capture the same trajectory at a higher level of abstraction and are often the only input available (e.g., a user directing an autonomous vehicle to a p...

20.07.2026
Blog Multimodal

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

arXiv:2607.15295v1 Announce Type: cross Abstract: We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transformer and modality dropout as masking, the model is trained to align the embeddings of global and per-modality local views, while the SIGReg objective encourages a theoretically optimal distribution. This achieves cross-modal alignment in the latent space, resulting in a remarkably clean architecture with no d...

20.07.2026
Blog LLMs & Texto

SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification

arXiv:2607.15517v1 Announce Type: new Abstract: Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated whether multimodal large language models (MLLMs) can verify identity from SLAP images. We introduce SLAPBench, the first benchmark for MLLM-based four-finger SLAP fingerprint verification, built from NIST SD302b with 7,832 pairs (176 ...

20.07.2026
Blog LLMs & Texto

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

arXiv:2607.15686v1 Announce Type: new Abstract: We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilit...

20.07.2026
Blog LLMs & Texto

Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts

arXiv:2607.15647v1 Announce Type: new Abstract: LEED v4.1 BD+C certification remains a document-intensive process that requires reviewers to read hundreds of pages of project evidence and apply credit-specific threshold logic by hand. This paper investigates whether small, locally deployed language models can perform meaningful screening of LEED documentation and how deterministic symbolic components should share that work. A neuro-symbolic pipeline is introduced that aligns project PDFs to LEED...

20.07.2026
O modelo aberto K3 da Kimi se aproxima do GPT-5.6 Sol e do Fable 5, ao mesmo tempo em que sinaliza o fim da IA chinesa superbarata
Blog LLMs & Texto

O modelo aberto K3 da Kimi se aproxima do GPT-5.6 Sol e do Fable 5, ao mesmo tempo em que sinaliza o fim da IA chinesa superbarata

A Kimi está lançando o K3, um modelo multimodal de pesos abertos com 2,8 trilhões de parâmetros e um milhão de tokens de contexto. Nos próprios benchmarks da empresa, ele chega perto do Claude Fable 5 e do GPT 5.6 Sol, ao mesmo tempo em que supera o Opus 4.8 e o GLM 5.2, em alguns casos por uma ampla margem. O modelo também é significativamente mais caro que seu antecessor. A liberação dos pesos completos está prevista até 27 de julho. O artigo O modelo aberto K3 da Kimi se aproxima do GPT-5.6 Sol e do Fable 5, ao mesmo tempo em que sinaliza o fim da IA chinesa superbarata aparece...

16.07.2026
467 itens no radar