基于视觉-语言基础模型的可泛化虹膜呈现攻击活体检测

Generalizable iris presentation attack liveness detection based on vision-language model

  • 摘要: 虹膜活体检测是保障虹膜识别系统可靠性的核心技术之一。针对现有方法依赖纯视觉特征建模,在复杂攻击场景下泛化能力不足的问题,本文提出一种结合文本驱动注意力机制与视觉基础模型的虹膜活体检测方法,旨在利用预训练语言模型的世界知识,显式引导模型学习过程,从而学习到更具判别性的虹膜特征表示。为了增强方法对虹膜特征的感知能力,引入基于文本提示的细粒度视觉文本对齐框架,实现虹膜图像特征与语义信息之间的精细匹配。同时,为了克服单一文本描述可能带来的偏见,采用可学习文本提示平均原型作为语义锚点。在此基础上,进一步通过基于槽注意力的文本条件约束对真假特征进行解耦,以提升方法在多样化攻击类型下的泛化性能。在虹膜活体检测比赛LivDet-Iris 2023数据集上的实验结果表明,所提方法相比于最先进的现有方法,综合平均分类错误率ACER1和ACER2分别降低了12.03%和6.46%,验证了该方法在复杂攻击场景下的有效性。除定量评估外,本文还开展了基于文本提示的定性实验,结果表明该方法能够捕捉到真假虹膜的特定特征,体现出相较于纯视觉方法更强的泛化能力。

     

    Abstract: Iris recognition, due to its uniqueness, stability, and high accuracy, has been widely applied in fields such as financial payments and public security. However, with continuous technological advancements, attackers are attempting to forge iris features in various ways to bypass identity verification systems. Iris liveness detection aims to distinguish bona fide iris samples from various types of attack iris images. Existing methods for iris liveness detection often suffer from inadequate robustness and generalization when faced with complex environmental conditions and diverse attack techniques, making them insufficient for practical applications. To address the limitation of existing iris liveness detection methods, which rely solely on visual feature modeling and exhibit insufficient generalization in scenarios involving both physical and digital attacks, this paper proposes an iris liveness detection method that integrates a text-driven attention mechanism with the vision foundation model, named Text-Guided Attention for IPAD (TGA-IPAD), which fuses information from both visual and textual modalities, leveraging the rich semantic knowledge of pre-trained language models to guide separation of bona fide and attack iris samples in the representation space. By utilizing textual semantics as a form of weak supervision, the model's learning process is explicitly constrained and guided, leading to more robust and discriminative iris feature representations. The architecture is designed to overcome the challenges of adapting generic vision-language models to the fine-grained domain of iris liveness detection. Specifically, to enhance the method’s ability to perceive iris features, a fine-grained visual-text alignment framework based on textual prompts is introduced, enabling precise matching between iris image features and semantic information, thereby directing the model's attention to subtle, attack-relevant regions such as texture irregularities, printing artifacts and so on. What’s more, multiple learnable textual prompts are employed to avoid bias from a single description, and their mean prototype features serve as the semantic anchor for alignment. Furthermore, the method employs text-conditioned constraints based on slot attention to decouple bona fide and spoof features in the latent space, thereby improving generalization across diverse attack types. Experimental results on the challenging LivDet-Iris 2023 competition datasets demonstrate the superiority of the proposed method over the winner methods from these competitions and several state-of-the-art (SOTA) approaches. Specifically, compared to the best SOTA method, the overall average classification error rate, ACER1 and ACER2, are reduced by 12.03% and 6.46%, respectively. Moreover, for synthetic iris attacks and physical attacks, the attack presentation classification error rate (APCER) are reduced by 27.98% and 1.43%, validating its effectiveness in open environment. Beyond quantitative evaluations, qualitative experiments utilizing textual prompts demonstrate that the model can effectively capture discriminative features specific to bona fide and attack irises, confirming that the integration of textual guidance leads to a more generalizable understanding of iris liveness detection. In conclusion, this study not only offers a high-performance solution for iris liveness detection but also pioneers a promising technical direction by effectively harnessing vision-language foundational models to significantly enhance the security and generalizability of biometric systems.

     

/

返回文章
返回