Weekly Large Model Application
Author

Weekly Large Model Application

Sharing to add value to technology

39
Articles
0
Likes
290
Views
0
Comments
Recent Articles

Latest from Weekly Large Model Application

39 recent articles
Weekly Large Model Application
Weekly Large Model Application
Sep 21, 2026 · Artificial Intelligence

Full-Duplex Voice Agents: When to Commit Parameters? Eager, Conservative, Two-Phase Compared

This article analyzes three commit strategies—Eager, Conservative, and Two-phase—for tool calling in full-duplex voice agents, explaining how each handles user self-correction, side-effect safety, and latency trade-offs, and provides selection guidelines based on operation reversibility.

System DesignTool Callingconservative commit
0 likes · 8 min read
Full-Duplex Voice Agents: When to Commit Parameters? Eager, Conservative, Two-Phase Compared
Weekly Large Model Application
Weekly Large Model Application
Sep 20, 2026 · Artificial Intelligence

Voice Agent Architecture: Full-Duplex Frontend + Async Backend Delegation

This article explains a voice agent architecture that separates real-time full-duplex frontend interaction from asynchronous backend computation using delegation IDs, detailing the two-timeline mechanism, three-layer responsibility split, task state machine, and evaluation strategies to avoid stale responses and interruption mishandling.

async backenddelegation patternfull-duplex
0 likes · 8 min read
Voice Agent Architecture: Full-Duplex Frontend + Async Backend Delegation
Weekly Large Model Application
Weekly Large Model Application
Aug 10, 2026 · Artificial Intelligence

LiveKit Turn Detector v1.0 Beats Deepgram Flux and Makes Turn Detection Measurable

LiveKit’s Turn Detector v1.0 replaces the traditional text‑only pipeline with a dual‑branch audio‑semantic architecture, achieving a 9.9% error rate at 300 ms latency—outperforming Deepgram Flux’s 12.9%—and releases an open‑source benchmark (eot‑bench) that turns end‑of‑turn detection into a reproducible engineering problem.

Audio ModelingDeepgram FluxLLM Fusion
0 likes · 11 min read
LiveKit Turn Detector v1.0 Beats Deepgram Flux and Makes Turn Detection Measurable
Weekly Large Model Application
Weekly Large Model Application
Jul 25, 2026 · Artificial Intelligence

How Pipeline Parallelism Cuts AI Voice Latency Below 700 ms

The article explains that keeping end‑to‑end voice‑assistant latency under 700 ms requires a co‑designed pipeline—streaming STT, speculative LLM, and streaming TTS—rather than faster individual models, and it details concrete component choices, budget allocations, and common pitfalls.

AI voiceLLMSTT
0 likes · 9 min read
How Pipeline Parallelism Cuts AI Voice Latency Below 700 ms
Weekly Large Model Application
Weekly Large Model Application
Jul 22, 2026 · Artificial Intelligence

AI Denoising Shifts from Call Beautification to Voice Intelligence (2025‑2026)

The article explains how AI‑driven noise reduction moves beyond simple call‑beautification to become the essential front‑line for speech recognition, voice agents, and localization, detailing traditional DSP shortcomings, three core AI techniques, latency trade‑offs, and a practical selection guide for various real‑time and offline scenarios.

AI denoisingGenerative Modelsreal-time audio
0 likes · 10 min read
AI Denoising Shifts from Call Beautification to Voice Intelligence (2025‑2026)
Weekly Large Model Application
Weekly Large Model Application
Jul 21, 2026 · Artificial Intelligence

From VAD to Flux: Built‑in End‑of‑Turn Eliminates Latency vs Interruption Trade‑off

Real‑time voice agents traditionally struggle between cutting users off and long pauses because turn detection relies on VAD and silence thresholds, but Deepgram's Flux model embeds End‑of‑Turn detection in ASR, delivering lower latency, more reliable turn quality, and a clear selection framework for practitioners.

CovalFluxTurn Detection
0 likes · 11 min read
From VAD to Flux: Built‑in End‑of‑Turn Eliminates Latency vs Interruption Trade‑off
Weekly Large Model Application
Weekly Large Model Application
Jun 23, 2026 · Artificial Intelligence

Inside Artificial Analysis: Independent AI Voice Benchmarks for ASR, TTS, and Speech‑to‑Speech

Artificial Analysis provides an independent, reproducible benchmarking platform for voice AI, offering objective WER scores for ASR, Elo‑based blind‑listening scores for TTS, and three‑dimensional metrics for end‑to‑end speech dialogue, together with detailed methodology, top‑model rankings, and practical guidance for developers to choose the most suitable model and provider for their scenarios.

AI voice evaluationASRArtificial Analysis
0 likes · 14 min read
Inside Artificial Analysis: Independent AI Voice Benchmarks for ASR, TTS, and Speech‑to‑Speech
Weekly Large Model Application
Weekly Large Model Application
Jun 16, 2026 · Artificial Intelligence

Building an Open‑Source TTS Evaluation Framework with ZipVoice, OmniVoice & Latest Benchmarks

This guide explains why TTS evaluation requires a three‑metric “iron triangle” (WER/CER, speaker similarity, and naturalness), introduces community benchmarks such as Seed‑TTS‑eval, TTSDS2, TTS Arena and TTSD‑eval, and provides a concrete six‑stage pipeline and best‑practice checklist for reproducible, production‑ready assessment.

CI pipelineOpen-source benchmarksSeed-TTS-eval
0 likes · 11 min read
Building an Open‑Source TTS Evaluation Framework with ZipVoice, OmniVoice & Latest Benchmarks
Weekly Large Model Application
Weekly Large Model Application
Jun 16, 2026 · Artificial Intelligence

Building a Reproducible, Scalable ASR Evaluation Framework for 2025‑2026

The article outlines why a unified ASR evaluation pipeline—combining a TestSet Zoo, Model Zoo, and standardized Benchmark Pipeline—is essential for fair cross‑model comparison, describes 2025‑2026 trends such as multi‑track metrics and robustness, and provides a step‑by‑step implementation guide with best‑practice warnings.

ASRBenchmarkNeMo
0 likes · 9 min read
Building a Reproducible, Scalable ASR Evaluation Framework for 2025‑2026