In late September 2026, coverage of arXiv paper 2605.07655 put Bharat ABIS — a research multimodal biometric search system — back in the news. The evergreen lesson for architects is not a product launch claim. It is the pipeline: how a multimodal ABIS architecture turns face, ten fingerprints, and two irides into one fixed-length template, then runs sharded GPU biometric 1:N deduplication against a national-scale gallery.
This post walks that enrollment path end to end, with the reported operating points from the paper, and draws a hard line between prototype metrics on a sampled gallery and production Aadhaar ABIS replacement claims (which the paper does not make).
What an ABIS does at enrollment vs authentication
An Automated Biometric Identification System (ABIS) has three logical pieces: a biometric pipeline that builds templates, a similarity-search module that compares a probe to the gallery, and a scheduler that shards work across servers. Enrollment de-duplication is a 1:N search: does this new person already exist among N enrolled identities? Authentication is usually a 1:1 check against a claimed ID, often with a single finger, iris, or face and a near-real-time budget.
Those workloads diverge in representation. Dedup favors a salient, multimodal template even if extraction is heavier. Auth can use a lighter unimodal vector after enrollment succeeds. The paper stresses that Aadhaar’s daily authentications are mostly single-modality with the 12-digit ID — not a full multimodal 1:N sweep.
Capture: face, 10 fingerprints, two irides
Capture is multimodal by design for coverage and accuracy at large N:
- Fingerprints — ten digits as two four-finger slaps plus a thumb pair on FTIR slap scanners at 500 dpi.
- Iris — left and right images captured together in near-infrared.
- Face — a frontal visible-spectrum photograph.
Stations may mark failure-to-acquire exceptions. The system is expected to enroll residents even when quality is poor or digits are missing — a core constraint that pushes fusion across modalities rather than hard reject-on-quality gates for fingerprints.
That same constraint is why face joined fingerprints and iris for de-duplication as face recognition matured: one weak modality must not collapse national uniqueness checks. Your diagram should show all three capture boxes feeding the pipeline in parallel, not a single “biometrics” blob.
Per-modality pipeline: segmentation, quality, PAD, embedding
Before search, each modality runs the same stage order: segmentation → quality → presentation attack detection (PAD) → embedding. Models are open-source-style CNNs fine-tuned on anonymized UIDAI data; the paper’s Aadhaar-design description states biometric images that enter the Central ID Repository (CIDR) do not leave it.
- Segmentation — YOLOv7-style face and fingerprint localization; EfficientNet U-Net iris masks with circle fitting for pupil and limbus.
- Quality — fingerprint UFIQ (NFIQ2-derived, FTIR-tuned); custom CNN estimators for face and iris. Quality feeds adaptive fusion weights later.
- PAD — spoof / liveness checks (prints, replays, contact lenses, non-distal phalanges, mixed packets). PAD does not score the 1:N match itself; it blocks fraudulent enrollments that would otherwise look “unique.”
- Embedding — DeepPrint (192-D per finger), ArcFace / InsightFace-style ResNet-50 face (512-D), ArcFace-style iris (512-D each eye). Matching uses inner product.
The concatenated multimodal template (3,456-D / 13.5 KB)
Embeddings are concatenated into one vector per identity:
t = [f1…f10 ; face ; iris_L ; iris_R] ∈ R^3456
10×192 + 512 + 2×512 = 3,456 dims → 13.5 KB (float32)
That fixed length is what makes exact GPU FAISS search practical: every gallery row is the same width, and similarity collapses to batched inner products plus top-k aggregation. Variable-length minutiae sets do not map as cleanly onto the same GPU matrix multiply path.
On the server cited in the paper, roughly 5M templates fit per H100, so eight GPUs hold about 40M identities in VRAM as an N × 3456 float32 matrix. Beyond that, additional shards are streamed or hosted on peer servers — a detail worth labeling on any multimodal ABIS architecture diagram so capacity planning stays honest.
Sharded GPU FAISS 1:N search and score fusion
The probe template is searched with FAISS IndexFlat on GPU (exact inner-product search in the paper’s setup). The gallery is partitioned across GPU servers; each shard returns candidates, then results are aggregated. On one server with 8×H100 and 2 TB RAM, about 40M residents fit; larger galleries stream additional shards.
The fused score is a weighted sum of per-modality inner products. The paper’s weight ratio is Face 12.5 : Iris 6.25 : Thumb+Index 2.3 : Other fingers 1, with support for age-specific and quality-adaptive weights. Reported throughput: 100 searches/sec on a 40M gallery on that single 8×H100 server — not “billions of searches per second.”
Thresholding, manual adjudication, and reported operating points
A threshold on the fused score labels duplicate vs unique. Candidates over the duplicate threshold — about 0.1% in the paper’s wording — are surfaced for manual adjudication.
Paper-reported metrics (hedge these as research results, not production SLAs):
- 220M demographically stratified gallery sampled from 1.55B Aadhaar records; adult probes: FNIR 0.3% at FPIR 0.5%.
- On a 20M gallery vs three unidentified COTS systems: Bharat ABIS FPIR 0.1% / FNIR 0.05% (Table 1); the paper describes performance as comparable or better than those COTS at their reported operating point. COTS vendors are not named in the paper.
What this paper is — and is not — claiming
Bharat ABIS is a prototype / research ABIS built from open-source architectures (DeepPrint, ArcFace/InsightFace-style face and iris, FAISS, YOLOv7 segmentation, and related stacks). It demonstrates competitive accuracy and search speed on sampled galleries up to 220M and a 20M COTS bake-off.
It is not a statement that Aadhaar’s production ABIS was replaced. Scaling to the full 1.55B gallery is described as underway and needing further model and engineering innovation — do not treat that as a finished billion-scale production deployment of Bharat ABIS. Gallery metrics are paper-reported; do not invent APIs, SLAs, or COTS brand identities the paper leaves blank.
Secondary press in late September 2026 summarized the same abstract numbers and noted historical commercial ABIS vendors used around Aadhaar. Treat those vendor names as press context only — the bake-off table in the paper leaves the three COTS unidentified, so architecture write-ups should keep them anonymous unless a primary source names the specific systems under test.
Diagram takeaways for system designers
- Split the diagram into enrollment 1:N (full multimodal template + sharded search + adjudication) vs auth 1:1 (lighter, often unimodal).
- Draw three modality lanes that converge only after embed — then a single concat box labeled 3,456-D / 13.5 KB.
- Show FAISS IndexFlat GPU shards and an aggregate top-k / fusion stage with explicit weight ratio callouts.
- Put PAD beside the pipeline as a fraud gate, not as the match score.
- Caption metrics with gallery size and “paper-reported / prototype” so readers do not confuse research OP with production cutover.
FAQ
What is ABIS 1:N vs 1:1?
1:N compares one probe to the whole gallery to catch duplicates at enrollment. 1:1 verifies a claimed identity, typically with one modality and an ID. Same ABIS ecosystem can use different templates and hardware budgets for each path.
Is Bharat ABIS production Aadhaar?
No. It is research prototyping and evaluation on sampled Aadhaar-derived galleries. The paper does not claim production ABIS replacement; full 1.55B scaling is still described as future work.
What is in the 13.5KB template?
Ten × 192-D DeepPrint fingerprint vectors, one 512-D face vector, and two 512-D iris vectors — 3,456 floats, 13.5 KB — fed to FAISS for exact GPU inner-product search.
Conclusion
A clear multimodal ABIS architecture diagram is a left-to-right enrollment story: capture three modalities, run seg/quality/PAD/embed per lane, concatenate a 13.5 KB template, shard FAISS 1:N search across GPUs, fuse with explicit weights, then threshold and adjudicate. Use Bharat ABIS as a concrete, cited reference design for biometric 1:N deduplication — and keep the prototype caveats in the same frame as the metrics. If you only remember one layout rule: modalities stay separate until the concat template; search, fusion, and adjudication sit after that single vector.
Diagram your multimodal ABIS pipeline
Map face, fingerprint, and iris branches through template concat, sharded GPU FAISS search, and adjudication in ByteDiagram — then animate enrollment 1:N vs auth 1:1 for your next architecture review.
Open Diagram Editor