This repository hosts the official implementation of our two-stage 3D object retrieval pipeline developed by team Stubborn_Strawberries. It demonstrates our approach for translating natural language queries and spatial context into accurate retrieval of 3D models, combining state-of-the-art vision–language embeddings, image captioning, and re-ranking techniques.
- ✨ Features
- 🗂️ Repository Structure
- 🚀 Pipeline Overview
- ⚙️ Installation
- 🎯 Usage
- 📈 Results
- 🎓 Citation
- 🤝 Contributing
- 📄 License
- Two-Stage Retrieval: Combines fast candidate retrieval with refined caption-based re-ranking.
- Multi-View Rendering: Generates 20 evenly spaced views per 3D object for comprehensive embedding.
- State-of-the-Art Models:
- SIGLIP for vision–language embeddings
- BLIP-2 (Flan-T5-xl) for image captioning
- BGE-M3 for text embedding and final ranking
- Scalable Storage: Uses Milvus vector database with a FLAT index for efficient similarity search.
├── model/
│ ├── __init__.py # Package initialization
│ ├── BGE_M3.py # BGE-M3 text embeddings interface
│ ├── BLIP_model.py # BLIP-2 image captioning implementation
│ └── SIGLIP_model.py # SIGLIP vision–language embedding interface
├── convert_file_obj_to_glb.py # OBJ → GLB converter
├── convert_glb_to_2d_imgs.py # 3D → 2D view renderer
├── embed_2d_imgs.py # Embed rendered views into Milvus
├── infer_without_caption_n_rerank.py # Stage 1 retrieval only
├── infer_with_caption_n_rerank.py # Full two-stage retrieval
├── make_dataset_for_blip.py # Prepare BLIP-2 training data
├── train_blip.py # BLIP-2 fine-tuning script
├── requirements.txt # Python dependencies
└── README.md # Project overview
A high-level view of the two-stage retrieval pipeline for 3D objects:
- Conversion:
convert_file_obj_to_glb.pytransforms.objfiles into binary.glb. - Rendering:
convert_glb_to_2d_imgs.pyproduces 20 uniformly spaced views (18° increments) around each model.
- Model: SIGLIP ViT-SO400M-16-SigLIP2-512
- Embedding:
embed_2d_imgs.pyencodes each view and stores vectors in Milvus with a FLAT index.
- Model: Fine-tuned BLIP-2 (Flan-T5-xl)
- Training:
- Epochs: 20
- Batch size: 32
- Learning rate: 2e-5
- Weight decay: 0.01
- Query Embedding: User text query → SIGLIP → vector
- Stage 1 Search: Milvus cosine similarity → top-k views
- Captioning: BLIP-2 generates captions for top candidates
- Re-ranking: BGE-M3 computes text embedding similarity between query and captions
- Output: Ranked list of 3D object IDs
-
Clone the repo:
git clone https://github.com/your-org/ROOMELSA.git cd ROOMELSA -
Setup environment:
python3 -m venv .venv source .venv/bin/activate pip install --upgrade pip pip install -r requirements.txt -
Milvus:
- Follow Milvus installation guide
- Create a collection with FLAT index type
python convert_file_obj_to_glb.py \
--root_dir /path/to/obj_files \
--output_dir /path/to/glb_files
python convert_glb_to_2d_imgs.py \
--root_dir /path/to/glb_files \
--output_dir /path/to/rendered_imagespython embed_2d_imgs.py \
--root_dir /path/to/rendered_images \
--collection_name your_milvus_collection-
Make Dataset:
python make_dataset_for_blip.py \ --src_root /path/to/public_data \ --dst_root /path/to/training_data
-
Train:
python train_blip.py \ --data_dir /path/to/training_data \ --output_dir /path/to/output \ --checkpoint_dir /path/to/checkpoints
-
Optional Hyperparams:
# Customize batch size, epochs, learning rate python scripts/train_blip.py \ --batch_size 16 \ --num_epochs 10 \ --learning_rate 1e-5
python infer_with_caption_n_rerank.py \
--query_path /path/to/queries.json \
--collection_name your_milvus_collection \
--output_csv_path /path/to/results.csv \
--visualize_path /path/to/visualizations \
--ckpt_path /path/to/best/fine_tuned_blip2For Stage 1 only (no reranking):
python infer_without_caption_n_rerank.py \
--query_path /path/to/queries.json \
--collection_name your_milvus_collection \
--output_csv_path /path/to/fast_results.csvIn the final leaderboard of the challenge, our team secured 1st place out of 18 participants with the following evaluation metrics:
- R@1 (Recall@1): 0.94
- R@5 (Recall@5): 1.00
- R@10 (Recall@10): 1.00
- MRR (Mean Reciprocal Rank): 0.97
This project is licensed under the MIT License. See the LICENSE file for details.

