Tagged articles

dialogue evaluation

2 articles · Page 1 of 1
Amap Tech
Amap Tech
Jun 23, 2026 · Artificial Intelligence

GrowLoop: Turning Subjective Dialogue Quality into a Rational Benchmark

GrowLoop proposes a self‑evolving loop that uses a few human seed annotations and large‑language‑model meta‑reflection to automatically generate and refine scoring rubrics and test questions for open‑domain dialogue, enabling reliable benchmarking where no fixed standard exists.

LLMbenchmarkdialogue evaluation
0 likes · 23 min read
GrowLoop: Turning Subjective Dialogue Quality into a Rational Benchmark
Meituan Technology Team
Meituan Technology Team
Jan 13, 2022 · Artificial Intelligence

MME-CRS: Multi-Metric Evaluation with Correlation Re-Scaling for Open-Domain Dialogue Evaluation

The paper presents MME‑CRS, a champion method for DSTC10 open‑domain dialogue evaluation that combines seven diverse metrics—fluency, relevance, topic coherence, engagement, and three specificity measures—using a correlation‑re‑scaling algorithm to weight each metric, achieving state‑of‑the‑art Spearman correlation and top rankings across multiple evaluation dimensions.

DSTC10MME-CRSOpen-domain Dialogue
0 likes · 21 min read
MME-CRS: Multi-Metric Evaluation with Correlation Re-Scaling for Open-Domain Dialogue Evaluation