Deep hashing encodes images into binary codes
Existing unsupervised methods rely on multi-view training
Most methods optimize retrieval, not collision resistance


Retrieval mAP (%) comparison under the unsupervised learning scenario. The number below each dataset is the retrieval performance of continuous features of pretrained VGG-16 and ViT-L/16 with dimensions 4096 and 1024, respectively using cosine similarity

Collision probability (‱). * denotes the hard-unsupervised.

Complexity across dataset sizes (N images). Numbers in parentheses indicate the complexity of the backbone itself.


Qualitative retrieval examples with 32-bit hash codes of CS3H with VGG16 backbone on Stanford Dogs dataset.

Qualitative retrieval examples with 12-bit hash codes of CS3H with ResNet backbone on Food-101 dataset
This work was supported by the ANR project ExcelLR (ANR-21-EXES0010) and the L3i Laboratory computing resources.
@inproceedings{duong2026collision,
  author={Duong, Anh-Kiet and Gomez-Krämer, Petra and Carozza, Jean-Michel},
  booktitle={2026 IEEE International Conference on Image Processing (ICIP)},Â
  title={Collision-Resistant Single-Pass Method for Unsupervised Fine-Grained Image Hashing},Â
  year={2026},
  pages={1-6},
  doi={10.1109/ICIP61757.2026.11630517}
}