Skip to content

About

This repository provides a complete, end-to-end system for retrieving precise 3D models by interpreting both natural language queries and spatial context.

Resources

Stars

9 stars

Watchers

1 watching

Forks

Repository files navigation

3D Object Retrieval Pipeline for ROOMELSA Grand Challenge 2025

ROOMELSA Grand Challenge

This repository hosts the official implementation of our two-stage 3D object retrieval pipeline developed by team Stubborn_Strawberries. It demonstrates our approach for translating natural language queries and spatial context into accurate retrieval of 3D models, combining state-of-the-art vision–language embeddings, image captioning, and re-ranking techniques.


📖 Table of Contents


✨ Features

  • Two-Stage Retrieval: Combines fast candidate retrieval with refined caption-based re-ranking.
  • Multi-View Rendering: Generates 20 evenly spaced views per 3D object for comprehensive embedding.
  • State-of-the-Art Models:
    • SIGLIP for vision–language embeddings
    • BLIP-2 (Flan-T5-xl) for image captioning
    • BGE-M3 for text embedding and final ranking
  • Scalable Storage: Uses Milvus vector database with a FLAT index for efficient similarity search.

🗂️ Repository Structure

├── model/
│   ├── __init__.py          # Package initialization
│   ├── BGE_M3.py            # BGE-M3 text embeddings interface
│   ├── BLIP_model.py        # BLIP-2 image captioning implementation
│   └── SIGLIP_model.py      # SIGLIP vision–language embedding interface
├── convert_file_obj_to_glb.py       # OBJ → GLB converter
├── convert_glb_to_2d_imgs.py        # 3D → 2D view renderer
├── embed_2d_imgs.py                 # Embed rendered views into Milvus
├── infer_without_caption_n_rerank.py   # Stage 1 retrieval only
├── infer_with_caption_n_rerank.py      # Full two-stage retrieval
├── make_dataset_for_blip.py          # Prepare BLIP-2 training data
├── train_blip.py                     # BLIP-2 fine-tuning script
├── requirements.txt                  # Python dependencies
└── README.md                         # Project overview


🚀 Pipeline Overview

A high-level view of the two-stage retrieval pipeline for 3D objects:

Pipeline image

1. 3D Model Rendering

  • Conversion: convert_file_obj_to_glb.py transforms .obj files into binary .glb.
  • Rendering: convert_glb_to_2d_imgs.py produces 20 uniformly spaced views (18° increments) around each model.

2. Vision–Language Embedding

  • Model: SIGLIP ViT-SO400M-16-SigLIP2-512
  • Embedding: embed_2d_imgs.py encodes each view and stores vectors in Milvus with a FLAT index.

3. Caption Generation

  • Model: Fine-tuned BLIP-2 (Flan-T5-xl)
  • Training:
    • Epochs: 20
    • Batch size: 32
    • Learning rate: 2e-5
    • Weight decay: 0.01

4. Retrieval Process

  1. Query Embedding: User text query → SIGLIP → vector
  2. Stage 1 Search: Milvus cosine similarity → top-k views
  3. Captioning: BLIP-2 generates captions for top candidates
  4. Re-ranking: BGE-M3 computes text embedding similarity between query and captions
  5. Output: Ranked list of 3D object IDs

⚙️ Installation

  1. Clone the repo:

    git clone https://github.com/your-org/ROOMELSA.git
    cd ROOMELSA
  2. Setup environment:

    python3 -m venv .venv
    source .venv/bin/activate
    pip install --upgrade pip
    pip install -r requirements.txt
  3. Milvus:


🎯 Usage

Data Preparation

python convert_file_obj_to_glb.py \
  --root_dir /path/to/obj_files \
  --output_dir /path/to/glb_files

python convert_glb_to_2d_imgs.py \
  --root_dir /path/to/glb_files \
  --output_dir /path/to/rendered_images

Generating Embeddings

python embed_2d_imgs.py \
  --root_dir /path/to/rendered_images \
  --collection_name your_milvus_collection

Training BLIP-2

  1. Make Dataset:

    python make_dataset_for_blip.py \
      --src_root /path/to/public_data \
      --dst_root /path/to/training_data
  2. Train:

    python train_blip.py \
      --data_dir /path/to/training_data \
      --output_dir /path/to/output \
      --checkpoint_dir /path/to/checkpoints
  3. Optional Hyperparams:

    # Customize batch size, epochs, learning rate
    python scripts/train_blip.py \
      --batch_size 16 \
      --num_epochs 10 \
      --learning_rate 1e-5

Inference

python infer_with_caption_n_rerank.py \
  --query_path /path/to/queries.json \
  --collection_name your_milvus_collection \
  --output_csv_path /path/to/results.csv \
  --visualize_path /path/to/visualizations \
  --ckpt_path /path/to/best/fine_tuned_blip2

For Stage 1 only (no reranking):

python infer_without_caption_n_rerank.py \
  --query_path /path/to/queries.json \
  --collection_name your_milvus_collection \
  --output_csv_path /path/to/fast_results.csv

📈 Results

In the final leaderboard of the challenge, our team secured 1st place out of 18 participants with the following evaluation metrics:

  • R@1 (Recall@1): 0.94
  • R@5 (Recall@5): 1.00
  • R@10 (Recall@10): 1.00
  • MRR (Mean Reciprocal Rank): 0.97

Leaderboard image


📄 License

This project is licensed under the MIT License. See the LICENSE file for details.

About

This repository provides a complete, end-to-end system for retrieving precise 3D models by interpreting both natural language queries and spatial context.

Resources

Stars

9 stars

Watchers

1 watching

Forks

Releases

Packages

Used by

Contributors

Languages