Skip to content

Repository files navigation

🧬 Intelligent Antibodies

A deep-learning pipeline for in-silico antibody design β€” sample candidate antibody sequences from a trained VAE, score them against a target antigen with a Siamese interaction classifier, and explore the results in an interactive dashboard.

Python 3.12 uv Streamlit Status

Quickstart β€’ Tutorial β€’ How it works β€’ Project structure β€’ Data


Therapeutic (monoclonal) antibodies are among the most effective treatments available today for chronic inflammatory diseases (Crohn's disease, lupus, multiple sclerosis) and certain cancers, and can be rapidly adapted against fast-mutating pathogens such as SARS-CoV-2. Yet only a few dozen are on the market β€” antibody design is slow, expensive, and still relies heavily on costly in-vitro screening and binding-energy estimates that are themselves hard to compute.

Intelligent Antibodies explores a faster, in-silico alternative: a convolutional VAE learns to generate plausible antibody sequences, and a Siamese CNN+GRU classifier predicts whether a generated candidate would actually bind a given antigen β€” turning candidate discovery into rejection sampling in a learned latent space instead of a physics-based energy calculation.

Features

  • πŸ§ͺ Generative pipeline β€” a 2-D-latent-space VAE samples candidate antibody sequences; a Siamese classifier scores each one against a target antigen (intelligent_antibodies/modules/).
  • πŸ“Š Interactive dashboard (Streamlit) β€” generate candidates, explore the sampled latent space with box/lasso selection, inspect per-candidate stats (hydrophobicity, charge, decode confidence), and view colorized sequences.
  • 🧬 Real structure prediction β€” fold any candidate with ESMFold (a practical, GPU-cluster-free stand-in for AlphaFold2) rendered as an interactive 3-D structure, colored by per-residue confidence.
  • 🧰 Runs out of the box β€” no trained weights needed to try the dashboard: it detects whether real models exist under run/models/ and otherwise shows clearly-labeled simulated results.
  • πŸ—‚οΈ Full data pipeline β€” scripts to fetch, parse, and equilibrate the SAbDab antibody-antigen dataset from scratch.

Quickstart

Try the dashboard in under a minute β€” no data download or training required, it runs on simulated results out of the box.

git clone https://github.com/sebgra/intelligent_antibodies.git
cd intelligent_antibodies

uv sync                           # installs Python 3.12 + all dependencies
uv run streamlit run app/main.py  # opens the dashboard in your browser

Everything here is managed with uv β€” no uv yet? curl -LsSf https://astral.sh/install.sh | sh (see the uv docs for other platforms).

Head to the Generate tab, check "Use bundled example antigen", and click Launch generation. You'll land on the Results tab with a full set of candidates, charts, and an interactive latent-space plot β€” all clearly labeled simulated, since no trained model is loaded yet. Follow the tutorial below to plug in real data and real trained models.

Tutorial

A step-by-step walkthrough from "just cloned it" to generating and folding real candidates.

1. Explore the dashboard with simulated data

uv run streamlit run app/main.py

This is the Quickstart above. Worth doing first regardless of whether you plan to train real models β€” it's the fastest way to see what every tab does:

  • Generate β€” pick an antigen (type a sequence, check the example box, or upload a FASTA file), and set the candidate count, interaction-score threshold, latent-space sampling temperature, and generated sequence length (the vector_size encoding window, 200 aa by default). Weights are stored per sequence length, so the dashboard loads the models trained for the length you pick β€” if none exist for it, it says which lengths are trained and falls back to simulated results.
  • Results β†’ Overview β€” KPIs, the score distribution across the sampled pool, and the top candidates.
  • Results β†’ Latent space β€” every sampled point in the VAE's 2-D latent space. Drag a box or lasso over a cluster of points to filter the other tabs down to just that selection.
  • Results β†’ Sequence analysis β€” length, hydrophobicity, and charge statistics, plus colorized sequences (by physicochemical residue class).
  • Results β†’ Candidate explorer β€” pick one candidate for a close-up: its colorized sequence, per-residue decode confidence, amino-acid composition, a schematic secondary-structure cartoon, and a button to fold it for real (see step 5).
  • Dataset & Model β€” real statistics computed from the actual SAbDab data and training-curve figures (nothing simulated on this tab).

2. Get the real data

Four scripts, run in order from the repo root, turn the raw SAbDab reference files (data/SAbDab/All_PDB_files.txt, positive_samples.txt, already included) into the tables the models train on. Every script resolves its own paths relative to the repo root, so they work from anywhere.

uv run python scripts/download_pdbs.py           # fetch FASTA files from RCSB
uv run python scripts/get_seq_table.py            # -> data/SAbDab/sequences.csv
uv run python scripts/get_interaction_table.py    # -> data/SAbDab/data.csv
uv run python scripts/filter_interaction_table.py # -> data/SAbDab/data_filtered.csv
Script What it does
download_pdbs.py Fetches the FASTA file for every structure referenced in All_PDB_files.txt into data/SAbDab/fasta/all_samples/. Skips ids it already has (safe to resume); --limit 20 for a quick smoke test, --force to re-fetch everything.
get_seq_table.py Turns the downloaded FASTA files into sequences.csv: one row per chain, tagged |ab or |ag by whether its molecule description reads as an antibody chain.
get_interaction_table.py Labels every pair of referenced structures 1 (interacting) or 0, using positive_samples.txt, into data.csv.
filter_interaction_table.py Drops pairs whose antibody or antigen side has no known sequence (e.g. a download failed), into data_filtered.csv β€” the file training actually reads.
What the generated files look like

data.csv:

ab;ag;interaction
5kel|ab;5kel|ag;1
5kel|ab;6cwt|ag;0
...

sequences.csv:

seq_id;specie;sequence
5kel|ag;Zaire ebolavirus (strain Mayinga-76) (128952);IPLGVIHNSTLQVSDVDKLVCRDKLSSTNQLRSVGLNLEGNGVATDVPSATKRWGFRSGVPPKVVNYEAGEWAENCYNLEIKKPDGSECLPAAPDGIRGFPRCRYVHKVSGTGPCAGDFAFHKEGAFFLYDRLASTVIYRGTTFAEGVVAFLILPQAKKDFFSSHPLREPVNATEDPSSGYYSTTIRYQATGFGTNETEYLFEVDNLTYVQLESRFTPQFLLQLNETIYTSGKRSNTTGKLIWKVNPEIDTTIGEWAFWETKKNLTRKIRSEELSFTVVSNGAKNISGQSPARTSSDPGTNTTTEDHKIMASENSSAMVQVHSQGREAAVSHLTTLATISTSPQSLTTKPGPDNSTHNTPVYKLDISEATQVEQHHRRTDNDSTASDTPSATTAAGPPKAENTNTSKSTDFLDPATTTSPQNHSETAGNNNTHHQDTGEESASSGKLGLITNTIAGVAGLITGGRRTRR
5kel|ag;Zaire ebolavirus (128952);EAIVNAQPKCNPNLHYWTTQDEGAAIGLAWIPYFGPAAEGIYTEGLMHNQDGLICGLRQLANETTQALQLFLRATTELRTFSILNRKAIDFLLQRWGGTCHILGPDCCIEPHDWTKNITDKIDQIIHDFVDKTLPDLEVDDDD
...

3. Train the models

Two models, trained separately, both reading from the files step 2 produced:

uv run python scripts/train_vae.py        # the generator (antibody VAE)
uv run python scripts/train_siamese.py    # the discriminator (interaction classifier)

Both default to the full dataset and the original hyperparameters (200-epoch VAE, 100-epoch Siamese with early stopping) β€” slow on CPU. Smoke-test the pipeline first with a tiny slice:

uv run python scripts/train_vae.py --epochs 5 --limit 300
uv run python scripts/train_siamese.py --epochs 3 --limit 2000

A GPU is used automatically if TensorFlow can see one; otherwise it falls back to CPU. Run either script with --help for every option (filters, batch size, validation split, early-stopping patience, ...).

Each script saves its weights where the dashboard and scripts/generate_antibodies.py expect them, and refreshes the matching plot in plots/:

run/models/vae/vae-one-hot-200-encoder.keras
run/models/vae/vae-one-hot-200-decoder.keras
run/models/siamese/one-hot-200-model.h5

run/ is gitignored β€” these are local artifacts, not something to commit.

4. Generate real candidates

Once both models exist, generation switches from simulated to real automatically β€” no flag to flip:

uv run streamlit run app/main.py

The Generate tab now shows "βœ… Trained models found" and every subsequent run uses them for real; the Results banner says explicitly which one (simulated or real) produced what you're looking at.

Prefer the command line? scripts/generate_antibodies.py runs the same pipeline and writes a FASTA file of passing candidates:

uv run python scripts/generate_antibodies.py --antigen-id 6xe1 \
    --n-candidates 30 --threshold 0.85 --temperature 1.2

5. Predict a real structure

In Results β†’ Candidate explorer, every candidate gets an instant schematic cartoon (a simplified secondary-structure heuristic β€” fast, offline, clearly labeled as illustrative). Click "πŸ”¬ Fold with ESMFold" underneath it to get a real predicted 3-D structure from the free public ESM Atlas API, rendered interactively and colored by per-residue confidence (pLDDT) using AlphaFold's own color convention. This needs an internet connection and can take up to about a minute, which is why it's on demand rather than automatic β€” see How it works for why ESMFold rather than AlphaFold2 itself.

How it works

flowchart LR
    AG[Target antigen] --> ENC[One-hot encode]
    Z["Sample z ~ N(0, T)<br/>(2-D latent space)"] --> VAE[VAE decoder]
    VAE --> CAND[Candidate antibody sequence]
    ENC --> SIAM[Siamese CNN + GRU classifier]
    CAND --> SIAM
    SIAM -->|score β‰₯ threshold| RANK[Ranked candidates]
    RANK --> FOLD["ESMFold / schematic cartoon"]
Loading
  • Generator β€” convolutional VAE (modules/models/VAEFull.py): trained on antibody sequences only (one-hot encoded, Conv2D encoder / Conv2DTranspose decoder), with a 2-D latent space. Sampling z = temperature Γ— N(0, 1) and decoding produces a new candidate sequence; temperature trades sampling diversity against decode fidelity.
  • Discriminator β€” Siamese classifier (modules/models/SiameseInteractionClassifier.py): a shared Conv1D + bidirectional-GRU tower embeds both the candidate antibody and the target antigen; the embeddings are combined and passed through a sigmoid to predict an interaction probability.
  • Rejection sampling (modules/utils/inference.py): candidates are sampled in batches and kept only if their predicted score clears a threshold β€” the same loop the dashboard and scripts/generate_antibodies.py both drive.
  • Dataset balancing (modules/dataset.py): real antibody-antigen pairs are overwhelmingly non-interacting, so training downsamples the negative class to match the positive one before fitting the classifier.

Project structure

intelligent_antibodies/
β”œβ”€β”€ app/                      Streamlit dashboard (app/main.py) + mock/real generation bridges
β”œβ”€β”€ intelligent_antibodies/   The installable package: models, data loading, encoding, inference
β”‚   └── modules/
β”‚       β”œβ”€β”€ models/           VAE + Siamese network definitions
β”‚       β”œβ”€β”€ layers/           Custom Keras layers (sampling, variational loss)
β”‚       └── utils/            Encoding schemes, inference loop, shared paths
β”œβ”€β”€ scripts/                  Data pipeline + training + generation CLIs (see the tutorial above)
β”œβ”€β”€ data/                     SAbDab / CoV-AbDab reference files and derived tables
β”œβ”€β”€ notebooks/                Original research notebooks (EDA, encoding experiments, prototyping)
β”œβ”€β”€ plots/                    Training-curve figures, refreshed by the training scripts
β”œβ”€β”€ web_interface/            Earlier Flask + Vue.js prototype, superseded by app/ but kept for reference
└── run/                      Trained weights and generation outputs (gitignored, created locally)

Data

Two datasets, both free to access:

  • SAbDab β€” antibody-antigen immune complexes characterized by X-ray crystallography, across many species. This is the dataset the current pipeline trains on.
  • CoV-AbDab β€” SARS-CoV antibody/antigen pairs (data/CoV-AbDab/). Included for future work; not yet wired into the training scripts above.

SAbDab

Two reference files drive the data pipeline (step 2 of the tutorial):

  • All_PDB_files.txt β€” one id per constitutive antigen-antibody structure. The first four characters are the RCSB PDB id; the rest identify the chains within that structure.
  • positive_samples.txt β€” pairs of structures known to form immune complexes, e.g. 4gms_J_N_E\t2vir_B_A_C. The relation is reciprocal and unordered β€” either partner can be the antibody or the antigen side, so filter_interaction_table.py checks both orderings. Any pair not listed here is treated as a negative (non-interacting) sample.

FASTA files are fetched directly from SAbDab/RCSB, e.g.:

>1A2Y_1|Chain A|IGG1-KAPPA D1.3 FV (LIGHT CHAIN)|Mus musculus (10090)
DIVLTQSPASLSASVGETVTITCRASGNIHNYLAWYQQKQGKSPQLLVYYTTTLADGVPSRFSGSGSGTQYSLKINSLQPEDFGSYYCQHFWSTPRTFGGGTKLEIK
>1A2Y_2|Chain B|IGG1-KAPPA D1.3 FV (HEAVY CHAIN)|Mus musculus (10090)
QVQLQESGPGLVAPSQSLSITCTVSGFSLTGYGVNWVRQPPGKGLEWLGMIWGDGNTDYNSALKSRLSISKDNSKSQVFLKMNSLHTDDTARYYCARERDYRLDYWGQGTTLTVSS
>1A2Y_3|Chain C|LYSOZYME|Gallus gallus (9031)
KVFGRCELAAAMKRHGLANYRGYSLGNWVCAAKFESNFNTQATNRNTDGSTDYGILQINSRWWCNDGRTPGSRNLCNIPCSALLSSDITASVNCAKKIVSDGNGMNAWVAWRNRCKGTDVQAWIRGCRL

More structures can be pulled from SAbDab's own search tool, e.g. antibody + protein antigen, with affinity or without; the full search page has more filters. A backup copy of a similar dataset is available at mit-ll/AlphaSeq_Antibody_Dataset.

CoV-AbDab

Three tab-separated files under data/CoV-AbDab/, each row a SARS-CoV identifier plus a pair of sequence columns: positive dataset.txt and negative dataset.txt (interacting / non-interacting), and independent test.txt as a held-out set.

Resources & references

Project status

This is an active research / learning project, not a production tool:

  • No pretrained weights are shipped β€” train your own (see the tutorial) or use the dashboard's simulated mode to explore the interface first.
  • The Siamese classifier's custom metrics (accuracy/f1/mcc) are simple, approximate formulas kept for continuity with earlier experiments, not calibrated, production-grade metrics.
  • CoV-AbDab is included but not yet wired into training.
  • No license has been chosen yet β€” please open an issue if you'd like to use this project and licensing matters to you.

Issues and pull requests are welcome.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

3 watching

Forks

Releases

Packages

Used by

Contributors

Languages