Evaluating routing stability and coordination in swarm-based multi-agent task-oriented dialogue systems

Оценка стабильности маршрутизации и координации в многоагентных системах целеориентированного диалога на основе роевого подхода
Abuzar Khan, Fahad Masood, Abid Iqbal, Ahmad Junaid, Saad Arif, Mohammed Al-Naeem, Ghassan Husnain, Ali Saeed Alzahrani
2026-03-03

MultiWOZ 2.2cascading-error attributioncoordination metricsmulti-agent task-oriented dialoguerouting stability
Conversational systems are becoming a primary interface for services and enterprise automation, and rapid market growth is pushing deployments into safety- and cost-sensitive settings. Reliability remains a bottleneck when interactions span multiple domains: an orchestrator must choose the next specialist, maintain shared dialogue state, and recover from mistakes before they cascade across handoffs. Despite rising interest in swarm-like multi-agent designs, orchestration is rarely evaluated with coordination-centric metrics, making it hard to compare routing policies beyond surface fluency. We present an evaluation-first pipeline for multi-domain task-oriented dialogue on MultiWOZ 2.2 that decouples routing from generation and exposes measurable failure modes. A DeBERTa-based router selects domain specialists, while a FLAN-T5 generator produces structured actions and belief-state updates under a shared memory interface. The protocol tracks delegation correctness, slot-progress coverage, switching and bouncing instability, loop behavior, and recovery after misroutes, and it links early-turn errors to downstream collapse using cascading-error attribution. We further introduce stress tests that simulate reformulation, long-horizon corrections, and tool-latency delays to probe robustness beyond static annotations. Across routing variants, confidence-aware gating yields the strongest stability improvement, achieving routing accuracy of 0.77 while substantially reducing handoff churn, with switching 0.11 and bounce 0.01, relative to a learned baseline with 0.65 accuracy, switching 0.44, and bounce 0.09. At the same time, confidence gating can trade progress for precision when it suppresses belief updates, highlighting an accuracy-progress tension that is important for deployment tuning. Diagnostic summaries identify misrouting and empty-state updates as dominant contributors, while looping is comparatively rare. Finally, applying the same evaluation to SGD shows that coordination challenges persist under schema shift. Overall, the proposed metrics and implementation blueprint provide a reproducible basis for diagnosing coordination failures and selecting orchestration policies for deployment.
1
Confidence gating improves routing stability but can reduce task progress by suppressing belief-state updates, revealing an accuracy–progress trade-off.
2
Confidence-aware routing achieves 0.77 accuracy with switching of 0.11 and bouncing of 0.01, outperforming a learned baseline with 0.65 accuracy, 0.44 switching, and 0.09 bouncing.
3
Misrouting and empty-state updates are the dominant diagnosed contributors to failure, whereas looping is comparatively rare; stress tests examine reformulation, long-horizon corrections, and tool latency.
4
The protocol evaluates delegation correctness, slot-progress coverage, switching, bouncing, loops, recovery after misroutes, and cascading-error attribution across dialogue turns.
5
The study introduces an evaluation-first pipeline that decouples routing from generation and measures coordination-specific failures in multi-agent task-oriented dialogue.

Swarm-based multi-agent task-oriented dialogue orchestration (router and specialist agents coordinating via shared memory) evaluated on MultiWOZ 2.2 and SGD

routing stability, inter-agent coordination, and cascading-error recovery under reformulation, long-horizon corrections, and tool-latency delays

Publication Details
Publication Date
2026-03-03
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Abuzar Khan
Fahad Masood
Abid Iqbal
Ahmad Junaid
Saad Arif
Mohammed Al-Naeem
Ghassan Husnain
Ali Saeed Alzahrani
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%