Skip to content

Repository files navigation

LG-CLIP 📚

Can Synthetic Images Serve as Effective and Efficient Class Prototypes?

ICASSP 2026 arXiv

🎉 Accepted by IEEE ICASSP 2026

LG-CLIP proposes a training-free zero-shot classification framework that replaces hand-crafted text prompts with synthetic image prototypes generated by Stable Diffusion — optionally enhanced with LLM-generated visual descriptions. This repository contains all code to reproduce the experiments from the paper.


Overview 🔍

The pipeline consists of five stages:

Real Images          Generated Images
     │                     │
[vanilla_clip.py]   [sd_gen / llm_sd_gen]→[gen_feat.py]
     │                     │
     └──────────┬──────────┘
            [mega.py]
                │
          Accuracy Report

Optional multi-scale extension:

[vanilla_clip_ms.py]  ←→  [muti_scale_gen_feat.py]
         └──────── [mega.py --ms1 _ms --ms2 _ms] ──────┘

Repository Structure 🌳

LG-CLIP/
├── vanilla_clip.py           # Extract CLIP features from real images; report zero-shot accuracy
├── vanilla_clip_ms.py        # Multi-scale variant of vanilla_clip.py
├── prompts_gen.py            # Use xAI Grok API to generate LLM-enriched SD prompts per class
├── sd_gen.py                 # Generate images via Stable Diffusion (simple class-name prompts)
├── sd_xl_gen.py              # Same as sd_gen.py but uses SD XL + accelerate
├── llm_sd_gen.py             # Generate images with LLM-enriched prompts (from prompts_gen.py)
├── text_gen_Ngen_made.py     # Filter top-N generated images using CLIP similarity score
├── gen_feat.py               # Extract CLIP features from generated images; report gen accuracy
├── muti_scale_gen_feat.py    # Multi-scale variant of gen_feat.py
├── muti_scale_merge.py       # Standalone merger for per-scale HDF5 files (real or gen)
├── mega.py                   # Main evaluation: prototype-based zero-shot classification
│
├── clip/                     # Local CLIP implementation (model.py, clip.py, tokenizer)
├── utils/
│   ├── myDataset.py          # Dataset loader classes (CUB, FLO, PET, FOOD, ImageNet, EUROSAT)
│   └── helper_func.py        # Utility functions (AverageMeter, numpy_to_pil, etc.)
│
├── scripts/                  # Bash scripts for each pipeline stage
│   ├── run_vanilla_clip.sh
│   ├── run_vanilla_clip_ms.sh
│   ├── run_prompts_gen.sh
│   ├── run_sd_gen.sh
│   ├── run_llm_sd_gen.sh
│   ├── run_gen_feat.sh
│   ├── run_muti_scale_gen_feat.sh
│   ├── run_text_gen_filter.sh
│   └── run_mega.sh
│
└── dataset/                  # (Not included) Dataset root directory
    ├── CUB/CUB_200_2011/
    ├── FLO/Flowers102/
    ├── PET/OxfordPets/
    ├── FOOD/
    ├── ImageNet/images/ILSVRC2012_img_val/
    ├── EUROSAT/
    ├── SD_gen/               # Plain SD generated images (auto-created)
    ├── LLM_SD_gen/           # LLM+SD generated images (auto-created)
    └── prompts/              # LLM-generated prompt JSON files (auto-created)

Datasets 📦

Dataset Classes Description
CUB 200 CUB-200-2011 fine-grained bird species
FLO 102 Oxford 102 Flowers
PET 37 Oxford-IIIT Pets
FOOD 101 Food-101
ImageNet 1000 ILSVRC2012 validation set (used as both train & test)
EUROSAT 10 EuroSAT satellite image classification

Expected Directory Layout

Each dataset follows a specific sub-structure under ./dataset/. Refer to the docstrings in utils/myDataset.py for the exact expected filenames (split JSONs, class name files, image directories).


Installation ⚙️

# Clone the repository
git clone https://github.com/<your-username>/LG-CLIP.git
cd LG-CLIP

# Create a conda environment (or use pip)
conda create -n lgclip python=3.9 -y
conda activate lgclip

# Install dependencies
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
pip install diffusers accelerate transformers
pip install h5py tqdm scipy pandas Pillow requests

Pipeline Usage 🚀

All scripts are run from the project root directory. Every Python script supports --help for full argument documentation. The bash scripts in scripts/ wrap the Python scripts with convenient CLI argument passing.

Stage 0 — Baseline: Zero-shot CLIP on Real Images

# Evaluate standard CLIP zero-shot accuracy on real test images
bash scripts/run_vanilla_clip.sh --dataset CUB --backbone ViT-L/14

# Multi-scale variant (10 crop scales, triangular aggregation)
bash scripts/run_vanilla_clip_ms.sh --dataset CUB --backbone ViT-L/14

Key arguments:

Argument Default Description
--dataset FLO Dataset: CUB | FLO | PET | FOOD | ImageNet | EUROSAT
--backbone RN50 CLIP backbone: RN50 | ViT-B/32 | ViT-B/16 | ViT-L/14
--image_root ./dataset Root directory of datasets
--device cuda:0 PyTorch device
--seed 2024 Random seed

Stage 1 — Generate Images

Option A: Plain Stable Diffusion

bash scripts/run_sd_gen.sh --dataset CUB --Ngen 10 --sd_version 2.1

Option B: LLM-enhanced SD (two-step)

Step 1 — Generate LLM prompts (requires xAI API key):

bash scripts/run_prompts_gen.sh \
    --dataset CUB \
    --api_key YOUR_XAI_API_KEY \
    --num_prompts 10

Step 2 — Generate images using the LLM prompts:

bash scripts/run_llm_sd_gen.sh \
    --dataset CUB \
    --json_file dataset/prompts/prompts_for_CUB.json \
    --Ngen 10

Key arguments for run_sd_gen.sh / run_llm_sd_gen.sh:

Argument Default Description
--dataset ImageNet Dataset name
--Ngen 10 Images per class
--sd_version 2.1 SD model: 2.1 | xl | 1.4
--gen_root_path ./dataset/SD_gen Output directory
--device cuda:0 GPU device

Stage 2 — Extract Generated Image Features

# Plain SD features
bash scripts/run_gen_feat.sh --dataset CUB --backbone ViT-L/14

# LLM+SD features
bash scripts/run_gen_feat.sh --dataset CUB --LLM LLM_ --backbone ViT-L/14

# Multi-scale variant
bash scripts/run_muti_scale_gen_feat.sh --dataset CUB --LLM LLM_ --backbone ViT-L/14

Key arguments:

Argument Default Description
--LLM "" "LLM_" for LLM-guided generation, "" for plain SD
--Ngen 10 Number of generated images per class
--backbone RN50 CLIP backbone for feature extraction

Stage 2b (Optional) — Filter Top-N Images

When you want to select fewer than 10 images per class (e.g., Ngen=5), first run the filter:

bash scripts/run_text_gen_filter.sh \
    --dataset CUB \
    --Ngen 5 \
    --LLM LLM_ \
    --backbone ViT-B/32

This creates a JSON index at dataset/CUB/json/LLM_SD_2.1_CUB_ViTB32_5.json. Then re-run run_gen_feat.sh with --Ngen 5.


Stage 3 — LG-CLIP Evaluation (mega.py)

# Standard evaluation (Ngen=10, LLM-enhanced, no multi-scale)
bash scripts/run_mega.sh \
    --dataset  CUB \
    --backbone ViT-L/14 \
    --Ngen     10 \
    --LLM      LLM_

# Multi-scale evaluation (requires ms features from previous steps)
bash scripts/run_mega.sh \
    --dataset  CUB \
    --backbone ViT-L/14 \
    --Ngen     10 \
    --LLM      LLM_ \
    --ms1      _ms \
    --ms2      _ms

Key arguments for mega.py:

Argument Default Description
--ms1 "" Suffix for real-image feature file ("" or "_ms")
--ms2 "" Suffix for gen-image feature file ("" or "_ms")
--LLM "" "LLM_" for LLM-guided prototypes
--Ngen 5 N generated images used in prototype averaging

Results are printed to stdout and appended to metalog.txt.


Feature File Naming Convention 📁

All features are cached as HDF5 files inside the dataset directory:

File Produced by Used by
CLIP_<backbone>_feature.hdf5 vanilla_clip.py mega.py
CLIP_<backbone>_feature_ms.hdf5 vanilla_clip_ms.py mega.py --ms1 _ms
<LLM>CLIP_<backbone>_feature_gen<N>.hdf5 gen_feat.py mega.py
<LLM>CLIP_<backbone>_feature_gen<N>_ms.hdf5 muti_scale_gen_feat.py mega.py --ms2 _ms

Where <backbone> = backbone name with - and / removed (e.g., ViTL14 for ViT-L/14).


Merging Multi-scale Features (Standalone) 🔧

If you need to re-merge per-scale features without re-extracting:

# Merge real-image features
python muti_scale_merge.py --mode real --dataset CUB --backbone ViT-L/14

# Merge generated-image features
python muti_scale_merge.py --mode gen --dataset CUB --backbone ViT-L/14 --LLM LLM_ --Ngen 10

Citation 📖

If you use this code in your research, please cite:

@inproceedings{shi2026lgclip,
  title     = {Can Synthetic Images Serve as Effective and Efficient Class Prototypes?},
  author    = {Shi, Dianxing and others},
  booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2026}
}

Paper: arXiv:2512.17160

About

[ICASSP 2026] Can Synthetic Images Serve as Effective and Efficient Class Prototypes?

Topics

Resources

Stars

2 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages