A deep-learning pipeline for in-silico antibody design β sample candidate antibody sequences from a trained VAE, score them against a target antigen with a Siamese interaction classifier, and explore the results in an interactive dashboard.
Quickstart β’ Tutorial β’ How it works β’ Project structure β’ Data
Therapeutic (monoclonal) antibodies are among the most effective treatments available today for chronic inflammatory diseases (Crohn's disease, lupus, multiple sclerosis) and certain cancers, and can be rapidly adapted against fast-mutating pathogens such as SARS-CoV-2. Yet only a few dozen are on the market β antibody design is slow, expensive, and still relies heavily on costly in-vitro screening and binding-energy estimates that are themselves hard to compute.
Intelligent Antibodies explores a faster, in-silico alternative: a convolutional VAE learns to generate plausible antibody sequences, and a Siamese CNN+GRU classifier predicts whether a generated candidate would actually bind a given antigen β turning candidate discovery into rejection sampling in a learned latent space instead of a physics-based energy calculation.
- π§ͺ Generative pipeline β a 2-D-latent-space VAE samples candidate antibody sequences; a Siamese classifier scores each one against a target antigen (
intelligent_antibodies/modules/). - π Interactive dashboard (Streamlit) β generate candidates, explore the sampled latent space with box/lasso selection, inspect per-candidate stats (hydrophobicity, charge, decode confidence), and view colorized sequences.
- 𧬠Real structure prediction β fold any candidate with ESMFold (a practical, GPU-cluster-free stand-in for AlphaFold2) rendered as an interactive 3-D structure, colored by per-residue confidence.
- π§° Runs out of the box β no trained weights needed to try the dashboard: it detects whether real models exist under
run/models/and otherwise shows clearly-labeled simulated results. - ποΈ Full data pipeline β scripts to fetch, parse, and equilibrate the SAbDab antibody-antigen dataset from scratch.
Try the dashboard in under a minute β no data download or training required, it runs on simulated results out of the box.
git clone https://github.com/sebgra/intelligent_antibodies.git
cd intelligent_antibodies
uv sync # installs Python 3.12 + all dependencies
uv run streamlit run app/main.py # opens the dashboard in your browserEverything here is managed with uv β no
uvyet?curl -LsSf https://astral.sh/install.sh | sh(see the uv docs for other platforms).
Head to the Generate tab, check "Use bundled example antigen", and click Launch generation. You'll land on the Results tab with a full set of candidates, charts, and an interactive latent-space plot β all clearly labeled simulated, since no trained model is loaded yet. Follow the tutorial below to plug in real data and real trained models.
A step-by-step walkthrough from "just cloned it" to generating and folding real candidates.
uv run streamlit run app/main.pyThis is the Quickstart above. Worth doing first regardless of whether you plan to train real models β it's the fastest way to see what every tab does:
- Generate β pick an antigen (type a sequence, check the example box, or upload a FASTA file), and set the candidate count, interaction-score threshold, latent-space sampling temperature, and generated sequence length (the
vector_sizeencoding window, 200 aa by default). Weights are stored per sequence length, so the dashboard loads the models trained for the length you pick β if none exist for it, it says which lengths are trained and falls back to simulated results. - Results β Overview β KPIs, the score distribution across the sampled pool, and the top candidates.
- Results β Latent space β every sampled point in the VAE's 2-D latent space. Drag a box or lasso over a cluster of points to filter the other tabs down to just that selection.
- Results β Sequence analysis β length, hydrophobicity, and charge statistics, plus colorized sequences (by physicochemical residue class).
- Results β Candidate explorer β pick one candidate for a close-up: its colorized sequence, per-residue decode confidence, amino-acid composition, a schematic secondary-structure cartoon, and a button to fold it for real (see step 5).
- Dataset & Model β real statistics computed from the actual SAbDab data and training-curve figures (nothing simulated on this tab).
Four scripts, run in order from the repo root, turn the raw SAbDab reference files (data/SAbDab/All_PDB_files.txt, positive_samples.txt, already included) into the tables the models train on. Every script resolves its own paths relative to the repo root, so they work from anywhere.
uv run python scripts/download_pdbs.py # fetch FASTA files from RCSB
uv run python scripts/get_seq_table.py # -> data/SAbDab/sequences.csv
uv run python scripts/get_interaction_table.py # -> data/SAbDab/data.csv
uv run python scripts/filter_interaction_table.py # -> data/SAbDab/data_filtered.csv| Script | What it does |
|---|---|
download_pdbs.py |
Fetches the FASTA file for every structure referenced in All_PDB_files.txt into data/SAbDab/fasta/all_samples/. Skips ids it already has (safe to resume); --limit 20 for a quick smoke test, --force to re-fetch everything. |
get_seq_table.py |
Turns the downloaded FASTA files into sequences.csv: one row per chain, tagged |ab or |ag by whether its molecule description reads as an antibody chain. |
get_interaction_table.py |
Labels every pair of referenced structures 1 (interacting) or 0, using positive_samples.txt, into data.csv. |
filter_interaction_table.py |
Drops pairs whose antibody or antigen side has no known sequence (e.g. a download failed), into data_filtered.csv β the file training actually reads. |
What the generated files look like
data.csv:
ab;ag;interaction
5kel|ab;5kel|ag;1
5kel|ab;6cwt|ag;0
...
sequences.csv:
seq_id;specie;sequence
5kel|ag;Zaire ebolavirus (strain Mayinga-76) (128952);IPLGVIHNSTLQVSDVDKLVCRDKLSSTNQLRSVGLNLEGNGVATDVPSATKRWGFRSGVPPKVVNYEAGEWAENCYNLEIKKPDGSECLPAAPDGIRGFPRCRYVHKVSGTGPCAGDFAFHKEGAFFLYDRLASTVIYRGTTFAEGVVAFLILPQAKKDFFSSHPLREPVNATEDPSSGYYSTTIRYQATGFGTNETEYLFEVDNLTYVQLESRFTPQFLLQLNETIYTSGKRSNTTGKLIWKVNPEIDTTIGEWAFWETKKNLTRKIRSEELSFTVVSNGAKNISGQSPARTSSDPGTNTTTEDHKIMASENSSAMVQVHSQGREAAVSHLTTLATISTSPQSLTTKPGPDNSTHNTPVYKLDISEATQVEQHHRRTDNDSTASDTPSATTAAGPPKAENTNTSKSTDFLDPATTTSPQNHSETAGNNNTHHQDTGEESASSGKLGLITNTIAGVAGLITGGRRTRR
5kel|ag;Zaire ebolavirus (128952);EAIVNAQPKCNPNLHYWTTQDEGAAIGLAWIPYFGPAAEGIYTEGLMHNQDGLICGLRQLANETTQALQLFLRATTELRTFSILNRKAIDFLLQRWGGTCHILGPDCCIEPHDWTKNITDKIDQIIHDFVDKTLPDLEVDDDD
...
Two models, trained separately, both reading from the files step 2 produced:
uv run python scripts/train_vae.py # the generator (antibody VAE)
uv run python scripts/train_siamese.py # the discriminator (interaction classifier)Both default to the full dataset and the original hyperparameters (200-epoch VAE, 100-epoch Siamese with early stopping) β slow on CPU. Smoke-test the pipeline first with a tiny slice:
uv run python scripts/train_vae.py --epochs 5 --limit 300
uv run python scripts/train_siamese.py --epochs 3 --limit 2000A GPU is used automatically if TensorFlow can see one; otherwise it falls back to CPU. Run either script with --help for every option (filters, batch size, validation split, early-stopping patience, ...).
Each script saves its weights where the dashboard and scripts/generate_antibodies.py expect them, and refreshes the matching plot in plots/:
run/models/vae/vae-one-hot-200-encoder.keras
run/models/vae/vae-one-hot-200-decoder.keras
run/models/siamese/one-hot-200-model.h5
run/ is gitignored β these are local artifacts, not something to commit.
Once both models exist, generation switches from simulated to real automatically β no flag to flip:
uv run streamlit run app/main.pyThe Generate tab now shows "β Trained models found" and every subsequent run uses them for real; the Results banner says explicitly which one (simulated or real) produced what you're looking at.
Prefer the command line? scripts/generate_antibodies.py runs the same pipeline and writes a FASTA file of passing candidates:
uv run python scripts/generate_antibodies.py --antigen-id 6xe1 \
--n-candidates 30 --threshold 0.85 --temperature 1.2In Results β Candidate explorer, every candidate gets an instant schematic cartoon (a simplified secondary-structure heuristic β fast, offline, clearly labeled as illustrative). Click "π¬ Fold with ESMFold" underneath it to get a real predicted 3-D structure from the free public ESM Atlas API, rendered interactively and colored by per-residue confidence (pLDDT) using AlphaFold's own color convention. This needs an internet connection and can take up to about a minute, which is why it's on demand rather than automatic β see How it works for why ESMFold rather than AlphaFold2 itself.
flowchart LR
AG[Target antigen] --> ENC[One-hot encode]
Z["Sample z ~ N(0, T)<br/>(2-D latent space)"] --> VAE[VAE decoder]
VAE --> CAND[Candidate antibody sequence]
ENC --> SIAM[Siamese CNN + GRU classifier]
CAND --> SIAM
SIAM -->|score β₯ threshold| RANK[Ranked candidates]
RANK --> FOLD["ESMFold / schematic cartoon"]
- Generator β convolutional VAE (
modules/models/VAEFull.py): trained on antibody sequences only (one-hot encoded, Conv2D encoder / Conv2DTranspose decoder), with a 2-D latent space. Samplingz = temperature Γ N(0, 1)and decoding produces a new candidate sequence;temperaturetrades sampling diversity against decode fidelity. - Discriminator β Siamese classifier (
modules/models/SiameseInteractionClassifier.py): a shared Conv1D + bidirectional-GRU tower embeds both the candidate antibody and the target antigen; the embeddings are combined and passed through a sigmoid to predict an interaction probability. - Rejection sampling (
modules/utils/inference.py): candidates are sampled in batches and kept only if their predicted score clears a threshold β the same loop the dashboard andscripts/generate_antibodies.pyboth drive. - Dataset balancing (
modules/dataset.py): real antibody-antigen pairs are overwhelmingly non-interacting, so training downsamples the negative class to match the positive one before fitting the classifier.
intelligent_antibodies/
βββ app/ Streamlit dashboard (app/main.py) + mock/real generation bridges
βββ intelligent_antibodies/ The installable package: models, data loading, encoding, inference
β βββ modules/
β βββ models/ VAE + Siamese network definitions
β βββ layers/ Custom Keras layers (sampling, variational loss)
β βββ utils/ Encoding schemes, inference loop, shared paths
βββ scripts/ Data pipeline + training + generation CLIs (see the tutorial above)
βββ data/ SAbDab / CoV-AbDab reference files and derived tables
βββ notebooks/ Original research notebooks (EDA, encoding experiments, prototyping)
βββ plots/ Training-curve figures, refreshed by the training scripts
βββ web_interface/ Earlier Flask + Vue.js prototype, superseded by app/ but kept for reference
βββ run/ Trained weights and generation outputs (gitignored, created locally)
Two datasets, both free to access:
- SAbDab β antibody-antigen immune complexes characterized by X-ray crystallography, across many species. This is the dataset the current pipeline trains on.
- CoV-AbDab β SARS-CoV antibody/antigen pairs (
data/CoV-AbDab/). Included for future work; not yet wired into the training scripts above.
Two reference files drive the data pipeline (step 2 of the tutorial):
All_PDB_files.txtβ one id per constitutive antigen-antibody structure. The first four characters are the RCSB PDB id; the rest identify the chains within that structure.positive_samples.txtβ pairs of structures known to form immune complexes, e.g.4gms_J_N_E\t2vir_B_A_C. The relation is reciprocal and unordered β either partner can be the antibody or the antigen side, sofilter_interaction_table.pychecks both orderings. Any pair not listed here is treated as a negative (non-interacting) sample.
FASTA files are fetched directly from SAbDab/RCSB, e.g.:
>1A2Y_1|Chain A|IGG1-KAPPA D1.3 FV (LIGHT CHAIN)|Mus musculus (10090)
DIVLTQSPASLSASVGETVTITCRASGNIHNYLAWYQQKQGKSPQLLVYYTTTLADGVPSRFSGSGSGTQYSLKINSLQPEDFGSYYCQHFWSTPRTFGGGTKLEIK
>1A2Y_2|Chain B|IGG1-KAPPA D1.3 FV (HEAVY CHAIN)|Mus musculus (10090)
QVQLQESGPGLVAPSQSLSITCTVSGFSLTGYGVNWVRQPPGKGLEWLGMIWGDGNTDYNSALKSRLSISKDNSKSQVFLKMNSLHTDDTARYYCARERDYRLDYWGQGTTLTVSS
>1A2Y_3|Chain C|LYSOZYME|Gallus gallus (9031)
KVFGRCELAAAMKRHGLANYRGYSLGNWVCAAKFESNFNTQATNRNTDGSTDYGILQINSRWWCNDGRTPGSRNLCNIPCSALLSSDITASVNCAKKIVSDGNGMNAWVAWRNRCKGTDVQAWIRGCRL
More structures can be pulled from SAbDab's own search tool, e.g. antibody + protein antigen, with affinity or without; the full search page has more filters. A backup copy of a similar dataset is available at mit-ll/AlphaSeq_Antibody_Dataset.
Three tab-separated files under data/CoV-AbDab/, each row a SARS-CoV identifier plus a pair of sequence columns: positive dataset.txt and negative dataset.txt (interacting / non-interacting), and independent test.txt as a held-out set.
- Deep learning benchmark for antibody-antigen binding β motivating article.
- piercelab/antibody_benchmark β a standard antibody-antigen docking benchmark.
- emersON106/AbAgIntPre (paper) β a closely related Siamese-network approach to antibody-antigen interaction prediction, including its own curated SAbDab subset.
- Sequence encodings explored during prototyping: one-hot, k-mer / Prot-Vec style encoders, and Chaos Game Representation.
- ESM Atlas β the public ESMFold API this project's structure-prediction feature calls.
This is an active research / learning project, not a production tool:
- No pretrained weights are shipped β train your own (see the tutorial) or use the dashboard's simulated mode to explore the interface first.
- The Siamese classifier's custom metrics (
accuracy/f1/mcc) are simple, approximate formulas kept for continuity with earlier experiments, not calibrated, production-grade metrics. - CoV-AbDab is included but not yet wired into training.
- No license has been chosen yet β please open an issue if you'd like to use this project and licensing matters to you.
Issues and pull requests are welcome.