// radar de ia

Áudio & Voz

Papers, modelos e datasets em alta no Hugging Face, além do blog oficial — com leitura editorial em português.

Blog Áudio & Voz

PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response

PolyAI has introduced Dialog-RSN-1, a dialog model that perceives caller audio directly instead of reading an ASR transcript. It fuses turn-taking, speech recognition, function calling, and response generation into a single audio-native model, keeps TTS separate so the output voice stays controllable, and runs as a request-based LLM rather than an always-on stream. PolyAI reports sub-300ms responses in live deployments. The post PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fus...

31.07.2026
Blog Áudio & Voz

Beyond Feeling Better: Capability-Sustaining Emotional Dialogue as a Longitudinal Research Paradigm

arXiv:2607.27851v1 Announce Type: new Abstract: Emotional dialogue research includes two influential strategy traditions. Empathetic dialogue prioritizes understanding a speaker's emotional experience. Emotional support conversation selects and sequences support for the seeker's current needs. Sustained use introduces a further goal. Effective support should sustain users' capacities for emotion regulation, coping, self-endorsed decisions, and social connection across the interaction lifecycle. ...

31.07.2026
Blog Áudio & Voz

Evaluation of forced alignment of code-mixed speech: the case of Hindi-English

arXiv:2607.25581v1 Announce Type: new Abstract: Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We address 2 problems: (1) free variation involving native vs non-native pairs and (2) phonemic boundary detection for mid-utterance English words. Bootstrapping strategies substantially outperform unmodified lexicons. Acous...

29.07.2026
Blog LLMs & Texto

MusiChat: Composição por Vibe para Criação Musical

arXiv:2607.24873v1 Tipo de anúncio: novo Resumo: Avanços recentes na geração de música por IA permitiram que usuários criassem peças musicais completas a partir de comandos em linguagem natural. No entanto, a maioria dos sistemas existentes segue um paradigma de comandar e regenerar, o que dificulta o refinamento iterativo, pois os usuários precisam recriar composições repetidamente em vez de evoluir diretamente ideias musicais já existentes. Apresentamos o MusiChat, um sistema conversacional de composição por vibe que possibilita a criação musical colaborativa entre humano e IA por meio de...

29.07.2026
Blog Áudio & Voz

S2A2: Audio-Visual Imitation Learning for Manipulation Tasks Using Acoustic Spatial Information

arXiv:2607.26047v1 Announce Type: new Abstract: Acoustic information provides rich cues about object location, material properties, and changes caused by contact or motion. This paper introduces a new set of acoustic-aware manipulation tasks for imitation learning, in which robots must use auditory cues to determine manipulation targets. These tasks require sound source localization and identification for active exploration in robotic manipulation. Also, we propose a multimodal imitation learnin...

29.07.2026
1 / 18 próxima →
216 itens no radar