Can Synthetic Images Serve as Effective and Efficient Class Prototypes?
🎉 Accepted by IEEE ICASSP 2026
LG-CLIP proposes a training-free zero-shot classification framework that replaces hand-crafted text prompts with synthetic image prototypes generated by Stable Diffusion — optionally enhanced with LLM-generated visual descriptions. This repository contains all code to reproduce the experiments from the paper.
The pipeline consists of five stages:
Real Images Generated Images
│ │
[vanilla_clip.py] [sd_gen / llm_sd_gen]→[gen_feat.py]
│ │
└──────────┬──────────┘
[mega.py]
│
Accuracy Report
Optional multi-scale extension:
[vanilla_clip_ms.py] ←→ [muti_scale_gen_feat.py]
└──────── [mega.py --ms1 _ms --ms2 _ms] ──────┘
LG-CLIP/
├── vanilla_clip.py # Extract CLIP features from real images; report zero-shot accuracy
├── vanilla_clip_ms.py # Multi-scale variant of vanilla_clip.py
├── prompts_gen.py # Use xAI Grok API to generate LLM-enriched SD prompts per class
├── sd_gen.py # Generate images via Stable Diffusion (simple class-name prompts)
├── sd_xl_gen.py # Same as sd_gen.py but uses SD XL + accelerate
├── llm_sd_gen.py # Generate images with LLM-enriched prompts (from prompts_gen.py)
├── text_gen_Ngen_made.py # Filter top-N generated images using CLIP similarity score
├── gen_feat.py # Extract CLIP features from generated images; report gen accuracy
├── muti_scale_gen_feat.py # Multi-scale variant of gen_feat.py
├── muti_scale_merge.py # Standalone merger for per-scale HDF5 files (real or gen)
├── mega.py # Main evaluation: prototype-based zero-shot classification
│
├── clip/ # Local CLIP implementation (model.py, clip.py, tokenizer)
├── utils/
│ ├── myDataset.py # Dataset loader classes (CUB, FLO, PET, FOOD, ImageNet, EUROSAT)
│ └── helper_func.py # Utility functions (AverageMeter, numpy_to_pil, etc.)
│
├── scripts/ # Bash scripts for each pipeline stage
│ ├── run_vanilla_clip.sh
│ ├── run_vanilla_clip_ms.sh
│ ├── run_prompts_gen.sh
│ ├── run_sd_gen.sh
│ ├── run_llm_sd_gen.sh
│ ├── run_gen_feat.sh
│ ├── run_muti_scale_gen_feat.sh
│ ├── run_text_gen_filter.sh
│ └── run_mega.sh
│
└── dataset/ # (Not included) Dataset root directory
├── CUB/CUB_200_2011/
├── FLO/Flowers102/
├── PET/OxfordPets/
├── FOOD/
├── ImageNet/images/ILSVRC2012_img_val/
├── EUROSAT/
├── SD_gen/ # Plain SD generated images (auto-created)
├── LLM_SD_gen/ # LLM+SD generated images (auto-created)
└── prompts/ # LLM-generated prompt JSON files (auto-created)
| Dataset | Classes | Description |
|---|---|---|
| CUB | 200 | CUB-200-2011 fine-grained bird species |
| FLO | 102 | Oxford 102 Flowers |
| PET | 37 | Oxford-IIIT Pets |
| FOOD | 101 | Food-101 |
| ImageNet | 1000 | ILSVRC2012 validation set (used as both train & test) |
| EUROSAT | 10 | EuroSAT satellite image classification |
Each dataset follows a specific sub-structure under ./dataset/. Refer to the docstrings in utils/myDataset.py for the exact expected filenames (split JSONs, class name files, image directories).
# Clone the repository
git clone https://github.com/<your-username>/LG-CLIP.git
cd LG-CLIP
# Create a conda environment (or use pip)
conda create -n lgclip python=3.9 -y
conda activate lgclip
# Install dependencies
pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
pip install diffusers accelerate transformers
pip install h5py tqdm scipy pandas Pillow requestsAll scripts are run from the project root directory.
Every Python script supports --help for full argument documentation.
The bash scripts in scripts/ wrap the Python scripts with convenient CLI argument passing.
# Evaluate standard CLIP zero-shot accuracy on real test images
bash scripts/run_vanilla_clip.sh --dataset CUB --backbone ViT-L/14
# Multi-scale variant (10 crop scales, triangular aggregation)
bash scripts/run_vanilla_clip_ms.sh --dataset CUB --backbone ViT-L/14Key arguments:
| Argument | Default | Description |
|---|---|---|
--dataset |
FLO |
Dataset: CUB | FLO | PET | FOOD | ImageNet | EUROSAT |
--backbone |
RN50 |
CLIP backbone: RN50 | ViT-B/32 | ViT-B/16 | ViT-L/14 |
--image_root |
./dataset |
Root directory of datasets |
--device |
cuda:0 |
PyTorch device |
--seed |
2024 |
Random seed |
bash scripts/run_sd_gen.sh --dataset CUB --Ngen 10 --sd_version 2.1Step 1 — Generate LLM prompts (requires xAI API key):
bash scripts/run_prompts_gen.sh \
--dataset CUB \
--api_key YOUR_XAI_API_KEY \
--num_prompts 10Step 2 — Generate images using the LLM prompts:
bash scripts/run_llm_sd_gen.sh \
--dataset CUB \
--json_file dataset/prompts/prompts_for_CUB.json \
--Ngen 10Key arguments for run_sd_gen.sh / run_llm_sd_gen.sh:
| Argument | Default | Description |
|---|---|---|
--dataset |
ImageNet |
Dataset name |
--Ngen |
10 |
Images per class |
--sd_version |
2.1 |
SD model: 2.1 | xl | 1.4 |
--gen_root_path |
./dataset/SD_gen |
Output directory |
--device |
cuda:0 |
GPU device |
# Plain SD features
bash scripts/run_gen_feat.sh --dataset CUB --backbone ViT-L/14
# LLM+SD features
bash scripts/run_gen_feat.sh --dataset CUB --LLM LLM_ --backbone ViT-L/14
# Multi-scale variant
bash scripts/run_muti_scale_gen_feat.sh --dataset CUB --LLM LLM_ --backbone ViT-L/14Key arguments:
| Argument | Default | Description |
|---|---|---|
--LLM |
"" |
"LLM_" for LLM-guided generation, "" for plain SD |
--Ngen |
10 |
Number of generated images per class |
--backbone |
RN50 |
CLIP backbone for feature extraction |
When you want to select fewer than 10 images per class (e.g., Ngen=5), first run the filter:
bash scripts/run_text_gen_filter.sh \
--dataset CUB \
--Ngen 5 \
--LLM LLM_ \
--backbone ViT-B/32This creates a JSON index at dataset/CUB/json/LLM_SD_2.1_CUB_ViTB32_5.json.
Then re-run run_gen_feat.sh with --Ngen 5.
# Standard evaluation (Ngen=10, LLM-enhanced, no multi-scale)
bash scripts/run_mega.sh \
--dataset CUB \
--backbone ViT-L/14 \
--Ngen 10 \
--LLM LLM_
# Multi-scale evaluation (requires ms features from previous steps)
bash scripts/run_mega.sh \
--dataset CUB \
--backbone ViT-L/14 \
--Ngen 10 \
--LLM LLM_ \
--ms1 _ms \
--ms2 _msKey arguments for mega.py:
| Argument | Default | Description |
|---|---|---|
--ms1 |
"" |
Suffix for real-image feature file ("" or "_ms") |
--ms2 |
"" |
Suffix for gen-image feature file ("" or "_ms") |
--LLM |
"" |
"LLM_" for LLM-guided prototypes |
--Ngen |
5 |
N generated images used in prototype averaging |
Results are printed to stdout and appended to metalog.txt.
All features are cached as HDF5 files inside the dataset directory:
| File | Produced by | Used by |
|---|---|---|
CLIP_<backbone>_feature.hdf5 |
vanilla_clip.py |
mega.py |
CLIP_<backbone>_feature_ms.hdf5 |
vanilla_clip_ms.py |
mega.py --ms1 _ms |
<LLM>CLIP_<backbone>_feature_gen<N>.hdf5 |
gen_feat.py |
mega.py |
<LLM>CLIP_<backbone>_feature_gen<N>_ms.hdf5 |
muti_scale_gen_feat.py |
mega.py --ms2 _ms |
Where <backbone> = backbone name with - and / removed (e.g., ViTL14 for ViT-L/14).
If you need to re-merge per-scale features without re-extracting:
# Merge real-image features
python muti_scale_merge.py --mode real --dataset CUB --backbone ViT-L/14
# Merge generated-image features
python muti_scale_merge.py --mode gen --dataset CUB --backbone ViT-L/14 --LLM LLM_ --Ngen 10If you use this code in your research, please cite:
@inproceedings{shi2026lgclip,
title = {Can Synthetic Images Serve as Effective and Efficient Class Prototypes?},
author = {Shi, Dianxing and others},
booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
year = {2026}
}Paper: arXiv:2512.17160