Dialt
Benchmarks
How Dialt compares.
Full-Duplex-Bench turn-taking is a measure of accuracy handling responses, pauses and interruptions. Every table includes the test date and model versions so results can be interpreted and reproduced as services change.
Full-Duplex-Bench (v1 + v1.5) 2026-08-19
| Provider / model | FDB v1 pause | FDB v1 turn taking | FDB v1.5 interruption | FDB v1.5 backchannel | FDB avg |
|---|---|---|---|---|---|
| Dialt | 95.6 | 100.0 | 100.0 | 98.0 | 98.4 |
| Qwen Audio 3.0 Realtime plus | 97.8 | 98.3 | 97.5 | 100.0 | 98.4 |
| GPT-Realtime-2.1 high | 95.6 | 98.3 | 94.0 | 94.9 | 95.7 |
| GPT-Realtime-2 high | 99.3 | 100.0 | 95.0 | 86.7 | 95.2 |
| Grok Voice Think Fast 2.0 high | 97.8 | 90.8 | 97.0 | 94.9 | 95.1 |
| Gemini 3.1 Flash Live high | 93.4 | 97.5 | 14.5 | 91.8 | 74.3 |
Scores 0–100, higher is better; FDB avg is the mean of the four gates; comparison rows are from the Artificial Analysis leaderboard (2026-08-14) at each model's flagship setting, and Dialt is measured on the same suites at our deployed configuration.
Coming up: voice-tau2 (agentic voice) benchmarks. Retail domain, regular and control speech complexity.