This voice biometric PAD SASV architecture diagram reviews two challenge documents, not a product. The ASVspoof 2021 evaluation plan (version 0.4, 16 July 2021) defines logical access (LA), physical access (PA), and speech deepfake (DF). LA and PA score a countermeasure in tandem with automatic speaker verification using minimum t-DCF. DF has no verifier and uses equal error rate. The ASVspoof 5 results paper adds SASV on post-sensor logical access only, with 53 teams, and does not evaluate replay.
What this voice biometric PAD SASV architecture diagram shows
In 2021 LA, bona fide speech and spoofs from text-to-speech (TTS), voice conversion (VC), or hybrid systems (VC fed with synthetic speech) cross PSTN or VoIP. The plan names alaw and G.722, says other codecs are used, and states there is no additive noise. The plan delivers that 2021 LA audio as 16 kHz FLAC without per-file codec metadata. PA in the plan is a room presentation: a high-quality loudspeaker as the talker, a MEMS or condenser microphone, and a replay device as the attack. This diagram does not draw that room. ASVspoof 5 calls Track 2 telephony or VoIP and excludes sensor-level replay.
Then a front-end and a countermeasure score. A verifier score is included only when the task has one. The 2021 baselines are LFCC-GMM, CQCC-GMM, LFCC-LCNN, and RawNet2. Top ASVspoof 5 submissions use waveform or mel front ends, or wav2vec 2.0 and WavLM in the open condition. The 2021 plan cites ISO/IEC 30107 but does not use the label PAD, and it specifies no liveness probe.
The problem this architecture is solving
A speaker matcher and a spoof detector answer different questions. The 2021 glossary defines anti-spoofing as the countermeasure to impersonation of an enrolled user. Bona fide speech is unmodified live speech. An LA spoof is speech modified automatically. A PA spoof is played back and re-recorded. Tandem t-DCF scores the pair while the subsystems stay isolated: participants submit countermeasure scores and organizers combine them with a common verifier. Normalized t-DCF of 1 means a useless countermeasure. The ASV floor is what remains when the countermeasure is perfect.
Main components and trust boundaries
In 2021 the participant never runs the verifier. Training uses only the ASVspoof 2019 countermeasure partitions. External speech, external models, and the 2019 evaluation set are forbidden. Codecs and compression tools are allowed if the system description says so. Each file is scored alone, and evaluation verifier scores are withheld during the challenge. The common verifier is deep speaker embeddings plus PLDA: the extractor is trained on VoxCeleb.
LA and PA share one cost table: target prior 0.9405, non-target 0.0095, spoof prior 0.0500, miss cost 1, false-accept cost 10, spoof false-accept cost 10. The appendix says these values are not a real-world case; they rank systems. DF is scored on its own with equal error rate. ASVspoof 5 closed training uses only its training partition (Track 2 may add VoxCeleb2). Open models must not overlap evaluation speakers or utterances; the paper cites LibriSpeech and VCTK as compliant and LibriLight as not. Track 2 returns one SASV score for target, nontarget, and spoof classes. Only target trials should be accepted.
Request or data path, step by step
On a 2021 logical-access trial the channel is PSTN or VoIP, and delivery is 16 kHz FLAC with codec identity hidden. That file format is in the 2021 evaluation plan only. ASVspoof 5 encodes part of the evaluation set with MP3, opus, amr, speex, m4a, Encodec, MP3 combined with Encodec, or a simulated mobile-to-PSTN path. Track 2 drops Encodec, MP3, m4a, and Encodec combined with MP3. The countermeasure emits one real-valued score. Higher favors bona fide speech.
The verifier runs only if the task has one. In 2021 its threshold is the organizer system's equal-error point. Development verifier scores are released; evaluation scores are not. ASVspoof 5 Track 2 compares the probe with enrolment utterance(s) of the claimed identity. Baselines are a fusion of AASIST with an ECAPA-TDNN pretrained on VoxCeleb2 (B03) and an end-to-end embedding score (B04). A team may build only the countermeasure and use the reference verifier.
The 2021 protocol does not ask the participant to build a gate. Its primary metric is minimum normalized ASV-constrained t-DCF; equal error rate is secondary. ASVspoof 5 allows tandem, score fusion, embedding fusion, or end-to-end. Track 2's primary metric, minimum a-DCF, needs one SASV score, with constants near 1.58 and 0.84. If separate scores are submitted, the paper also reports minimum t-DCF and t-EER. ASVspoof 5 Track 1 uses a beta near 1.90.
Replay is 2021 PA only: 2019 simulated training, mostly real evaluation replay, room size 10 to 65 square metres, talker distance 50 to 200 cm, attacker factors unlabeled. Do not attach it to ASVspoof 5. DF is TTS or VC under general compression (mp3 and m4a known, bitrate hidden), equal error rate, no verifier, including attacks aimed at a human listener.
The diagram: labeled boxes and failure or isolation edges
The drawing shows a microphone or channel, front-end features, and a countermeasure score. A verifier is included only when the task has one. LA, PA, and DF are evaluation tasks, not boxes. Replay into the mic is on the diagram, dashed and labeled PA edge only. TTS and voice conversion into the channel are on it too, dashed and labeled LA side only. A logical-access-only hybrid and named adversarial-attack callouts are source distinctions, and they are not on this diagram.
In 2021 the countermeasure score and the verifier score stay separate until organizers combine them. ASVspoof 5 may use tandem, score fusion, embedding fusion, or an end-to-end score. Tandem or fusion is one option, not the only path. A false accept is a spoof scored above threshold. A miss rejects bona fide speech. Normalized t-DCF of 1 means a useless countermeasure; the ASV floor remains if it is perfect. A hard block before speaker verification is one later option, not the 2021 submission rule.
What the source does not claim (preview, case study, or limits)
The 2021 file is an evaluation plan, not a latency budget or a scale claim. Priors are ranking assumptions, not a measured deployment. The plan calls equal-error-rate reporting deprecated by the ISO/IEC standards it cites, yet still uses equal error rate as the secondary LA and PA metric and as the DF metric. It specifies no template store, match-on-device design, or API.
The results paper does not define a production cascade and does not claim that open-condition foundation models generalize. Replaying spoofs in a room is future work, not 2021 PA. A neural codec is not itself a spoof class. Speech is English read speech from Multilingual Librispeech, almost 2,000 speakers. Neither source states a preview, beta, or general-availability label, a CRD field, or a latency number. A gate before speaker verification is one SASV option, not the 2021 submission rule.
FAQ
What is the difference between ASVspoof 2021 logical access, physical access, and speech deepfake?
Logical access sends bona fide speech and TTS, voice-conversion, or hybrid spoofs over PSTN or VoIP, with coding and transmission and no additive noise, and uses minimum t-DCF with a common verifier. Physical access is real-room replay, trained on 2019 simulated replay, also under minimum t-DCF. Speech deepfake detects TTS and voice conversion under general compression, with no verifier, using equal error rate. The 16 July 2021 plan scores the three tasks independently.
Does tandem t-DCF mean the countermeasure must reject audio before speaker verification runs?
Not in the 2021 plan. Participants submit countermeasure scores only; organizers combine them with an isolated verifier. The primary metric is minimum normalized ASV-constrained t-DCF, and 1 means a useless countermeasure. ASVspoof 5 allows tandem, score fusion, embedding fusion, or end-to-end, and Track 2 minimum a-DCF needs only one SASV score. A hard block-before-verify gate is a later option, not the 2021 protocol.
Does ASVspoof 5 evaluate replay inside spoofing-robust speaker verification?
No. The paper limits ASVspoof 5 to post-sensor logical-access attacks and leaves sensor-level replay as future work. Track 2 accepts only target trials among target, nontarget, and spoof, on telephony or VoIP, and drops Encodec, MP3, m4a, and Encodec combined with MP3. Replay is the 2021 physical-access task, not an ASVspoof 5 trial.
Conclusion
Replay is the 2021 physical-access attack, separate from text-to-speech and voice conversion. Speech deepfake has no verifier. Do not treat challenge priors or leaderboard costs as a production false-accept rate. Cite the ASVspoof 2021 evaluation plan and the ASVspoof 5 results paper. Browse more architecture diagrams on the ByteDiagram blog.
Diagram the voice PAD and SASV cascade
Map the microphone or channel, front-end features, and a countermeasure score. Include a verifier only when the task has one. Tandem or fusion is one ASVspoof 5 option, not the only path.
Open Diagram Editor