PVQA 1.9 is out. We compared it against ViSQOL on two public MOS corpora. Numbers below.
Setup: take clips from each corpus, compute MOS on each clip with ViSQOL and PVQA, then compare the predicted MOS against the corpus subjective MOS ratings.
r is Pearson, rho is Spearman, “within 0.5” is the share of clips where the predicted MOS is within ±0.5 of the corpus subjective rating.
PVQA is ahead of ViSQOL on all three corpora, across r, rho and within 0.5. The biggest gap is TCD-VoIP: r goes from 0.82 to 0.91, and the share of clips within 0.5 MOS from 68% to 73%. TCD-VoIP is built around real VoIP impairments (packet loss, jitter, codec artefacts), which is where a call quality analyser earns its keep.
NISQA is harder: mixed conditions, background noise, many speakers, varied recordings. The whole industry sits below 0.75 there. PVQA still wins on r, rho and within 0.5, by narrower margins.
Both corpora are public. No hand-picked test set. Anyone can rerun the comparison.
PVQA is Sevana’s non-intrusive perceptual voice quality analyser. For details contact: info@sevana.biz.