sarulab-speech/DuplexChat
Dataset em destaque no Hugging Face — 128 downloads. DuplexChat large-scale, two-speaker, full-duplex spoken-dialogue corpus built from public podcast feeds.
Papers, modelos e datasets em alta no Hugging Face, além do blog oficial — com leitura editorial em português.
Dataset em destaque no Hugging Face — 128 downloads. DuplexChat large-scale, two-speaker, full-duplex spoken-dialogue corpus built from public podcast feeds.
Wan-Streamer v0.2 enhances audio-visual interaction by increasing visual resolution while maintaining low latency through optimized thinker-performer architecture with multi-GPU pa…
A speaker-disentangled syllabic tokenizer regresses perturbed student representations toward clean teacher targets to improve syllable boundary detection and speech language modeli…
Com apenas 600 milhões de parâmetros, o Nemotron 3.5 ASR Streaming transcreve 40 variações de idioma palavra por palavra, à medida que a pessoa fala — e sustenta até 17 vezes mais conversas simultâneas que o modelo de streaming anterior da própria NVIDIA.
Com 2,5 bilhões de parâmetros e uma taxa de erro de palavras de 5,63%, o Canary-Qwen da NVIDIA chegou ao topo do Open ASR Leaderboard e virou modelo a copiar — mesmo que o primeiro lugar, como sempre, já tenha trocado de dono.
arXiv:2607.01245v1 Tipo de anúncio: novo Resumo: Apresentamos o Office Comprehension Bench (OCB), o primeiro benchmark público a avaliar conjuntamente sistemas de LLM na compreensão de Word, Excel e PowerPoint sobre formatos de arquivo nativos (.docx, .xlsx, .pptx) e suas variantes. O OCB é composto por duas trilhas. A trilha de Perguntas e Respostas de Fidelidade de Arquivo testa a percepção estrutural e visual de artefatos de escritório — tabelas, gráficos, imagens incorporadas, fórmulas e elementos específicos de cada aplicativo, como cabeçalhos, notas do apresentador e intervalos nomeados. Q de Domínio...
arXiv:2607.01729v1 Announce Type: new Abstract: Deep learning models for speech classification are vulnerable to backdoor attacks, where malicious triggers cause misclassification at inference time. While sample-specific attacks can bypass many defenses, they often rely on poisoned label attack, making them detectable via manual data defense. In this paper, we propose DRL-CLBA, a novel clean label backdoor attack for speech classification that leverages Deep Deterministic Policy Gradient (DDPG) ...
arXiv:2607.01502v1 Announce Type: new Abstract: Recent advances in automatic speech recognition (ASR) have explored different sequence models, including Conformer-based models and newer state space models such as Mamba. Although prior work has evaluated these architectures in multiple languages, their effectiveness in African languages remains underexplored. In this work, we evaluate Mamba for ASR on seven South African languages. In monolingual experiments, each model is trained on 50 hours of ...
arXiv:2607.01470v1 Announce Type: new Abstract: Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural candidates for RL from world feedback: once clinical SMEs encode decision logic into a verifier, that verifier grades unlimited rollouts without per-episode annotation. But applying RL requires a sound feedback channel and sufficient base capability. We audit MedAgentBench v1/v2, find a 41.7\% silent-finish ceili...
arXiv:2607.01238v1 Announce Type: new Abstract: Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P) systems that fail to capture speaker-specific acoustic variation. Prior work demonstrates that grapheme-based models outperform phoneme-based systems at scale, but not in low-resource settings. In this paper, we propose SPARCLE, a ...
arXiv:2607.01733v1 Announce Type: new Abstract: Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increases, the contribution of LLM priors becomes less evident, and simple speech-text joint training under-utilizes textual knowledge. We therefore propose Joint Speech-Text Interleaved Pretraining (JSTIP), an ASR-oriented pre...
Interfaze open-sourced diffusion-gemma-asr-small, a multilingual ASR model that transcribes via diffusion, not autoregression. It adds audio to Google's frozen DiffusionGemma using a ~42M-parameter adapter. One adapter covers six languages, with transcription cost set by denoising steps, not transcript length. The post Interfaze Ships diffusion-gemma-asr-small, an Open-Source Diffusion ASR Model Transcribing Six Languages via DiffusionGemma’s Parallel Denoising Decoder appeared first on MarkTe...