HashViT: Native Hash Learning with Dedicated HASH Token for Efficient Image Retrieval
HashViT presents a Vision Transformer framework that natively learns binary hash codes using a dedicated HASH token, addressing the feature-to-code discrepancy found in post-quantization methods. The HASH token is split into a Hash Register for binary code generation and a Semantic Workspace for continuous semantics, with a refinement adapter enabling progressive improvement across transformer layers. Experiments on three standard benchmarks show that HashViT achieves state-of-the-art or highly competitive image retrieval performance with compact Hamming codes.
Why it matters: By integrating binary code learning directly into the transformer backbone, HashViT could enable more efficient and accurate large-scale image retrieval systems.
Full story at: arXiv Information Retrieval ↗