// radar de ia

Áudio & Voz

Papers, modelos e datasets em alta no Hugging Face, além do blog oficial — com leitura editorial em português.

Blog Dados & Embeddings

Um Ensemble Multimodal Calibrado para Reconhecimento de Ambivalência/Hesitação: Descrição do Sistema e Estratégia de Submissão para o Teste Privado

arXiv:2607.12176v1 Tipo de Anúncio: novo Resumo: A ambivalência e a hesitação (A/H) prejudicam as intervenções digitais de mudança de comportamento, e reconhecê-las automaticamente a partir de vídeo é o objetivo do desafio ABAW A/H sobre o conjunto de dados BAH. Descrevemos nosso sistema para a 11ª edição do desafio: um ensemble calibrado e de pesos iguais de três modelos de fusão sobre embeddings congelados de rosto, áudio, texto e pose, que atinge 0,7358 de macro-F1 no conjunto de teste público. O teste privado deste ano, divulgado em um...

15.07.2026
Blog LLMs & Texto

Exploring Agentic Workflows for Generating High Quality Math Visual Aids

arXiv:2607.09839v1 Announce Type: new Abstract: Mathematical diagrams play a crucial role in K 12 education, both as problem components and as scaffolding for student comprehension. However, current AI tools, including Large Language Models (LLMs), struggle to reliably generate accurate and pedagogically sound visual diagrams, even when provided with detailed descriptions. A significant gap therefore remains in the reliable generation of diagrams for middle school mathematics. To address this, w...

14.07.2026
Blog Robótica & RL

RASR: Range-Aware Scale Recovery for Metric UAV Navigation

arXiv:2607.09815v1 Announce Type: new Abstract: Under Global Navigation Satellite System (GNSS) denial, a UAV controller still needs a distance and heading command it can execute, making accurate metric last-meter navigation essential. Dense pair-geometry foundation models transfer relative structure well, yet the distance scale of their raw metric outputs remains poorly calibrated. Under the relative error metric of PairUAV, correcting only the average scale can still leave costly, distance-dep...

14.07.2026
Blog Áudio & Voz

Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR

arXiv:2607.10256v1 Announce Type: new Abstract: This paper investigates how language similarity can improve cross-lingual transfer for automatic speech recognition (ASR) in extremely low-resource settings. Warlpiri, an Australian Aboriginal language, has very limited transcribed speech data, making transfer learning essential. We propose a framework combining acoustic similarity from pre-trained speech models with linguistic similarity based on typology, phoneme inventories, grammatical, and syn...

14.07.2026
Blog LLMs & Texto

Efficiently Adapting Spoken Language Models for the Singaporean Context

arXiv:2607.10092v1 Announce Type: new Abstract: Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially when the original training data is inaccessible and the use case demands multilingual, spoken-query interaction. We adapt an open-source SLM to the Singaporean Home Team context across five speech tasks in Singapore's four official languages, combining LoRA fine-tuning, a surrogate text-QA dataset that guards agai...

14.07.2026
Building a VideoAgent-Style Multi-Agent System: Intent Parsing, Graph Planning, and Tool Routing for Video Editing Tasks
Blog LLMs & Texto

Building a VideoAgent-Style Multi-Agent System: Intent Parsing, Graph Planning, and Tool Routing for Video Editing Tasks

In this tutorial, we reconstruct the VideoAgent workflow as a runnable, API-key-free multi-agent pipeline. We build an intent parser, an agent library, a tool router, a graph planner, and a textual-gradient optimizer that repairs the execution graph. We wire these planning components to FFmpeg, Whisper transcription, scene detection, keyframe sampling, captioning, cross-modal indexing, and beat-synced editing. By the end, we have a system that answers questions about a video, summarizes it, and ...

13.07.2026
216 itens no radar