Repository navigation
Conversation
Update project metadata, uv environment, CI workflows, and setup docs for Python 3.13. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Index explicit called genotypes against a validated SNP panel, expose KING-robust pair and cohort matching, and add VCFtools acceptance and performance benchmarks. Document implementation decisions, data requirements, integration tests, and benchmark interpretation. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Generate deterministic synthetic cohorts and separately measure index construction, tiled genotype retrieval, and pair scoring. Record workload dimensions and phase metrics in the benchmark report. Document reproducible benchmark usage and results; add tests for generated inputs and phase accounting. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
|
Performance comparison with vcftools 🐌 KING benchmark reportMode: smoke. Cohort: 128 observations, 512 markers, 8128 unique pairs. Input preparation and VRS annotation are excluded; this is not end-to-end timing. Workloads are reported separately:
Isolated index, retrieval, and scoring phasesThese phase durations are measured inside the worker, excluding subprocess startup.
Workload comparisonsRelative performance is
The first comparison measures matching against an already-built VRS-Matcher index; the VCFtools figure still includes reading the VCF, so it is not an algorithm-only comparison. The second includes VRS-Matcher index construction and is the one-shot pipeline view. VCFtools reads the prepared VCF directly. The tools also emit different pair sets as noted above. The Identity-only DB bytes: 9519104. See manifest.json, raw logs, pairs.tsv, and timings.tsv for reproducibility. |
|
Title: feature/vcftools-benchmark Purpose: Add KING-robust kinship matching and VCFtools benchmarks to enable bioinformaticists to estimate relationships between samples and search indexed cohorts for duplicates and relatives using KING (Manichaikul) kinship estimates. This PR introduces a second matching algorithm alongside the existing Key Changes:
Architecture & ImplementationCore additions:
Test Coverage
Validation & Reproducibility
Merge Readiness and Risk AssessmentStatus: Clean and ready to merge with attention to the notes below. Strengths:
Observations:
|
Add KING-robust kinship matching and VCFtools benchmarks
Summary and use case
Enable a bioinformaticist to compare two observations or search an indexed cohort for potential duplicates and relatives using Manichaikul (KING) kinship estimates. The existing carried-allele index supports allele similarity but cannot distinguish a called reference genotype from missing or filtered data. This branch adds explicit genotype indexing and a
king-robustplugin while preserving the existingidentityworkflow.It also gives engineers a reproducible way to compare numerical results and computational costs with VCFtools
--relatedness2, separating initial indexing from repeated queries. See the KING user story and acceptance criteria and bioinformatics comparison with VCFtools.Intended base:
development(9382c73).Architecture changes
load-samples --index-genotypes --panel PANEL.tsv. Additional SQLite tables store the declared panel, marker identities, observation provenance/completion status, and passing ALT dosages 0, 1, and 2. Validate reference/annotation consistency, observation uniqueness, and QC policy; commit genotype and allele indexing atomically. An absent genotype row means unknown, never reference. The carried-allele table retains its existing identity semantics.PluginContextwith panel metadata and explicit genotype access. Add dedicatedKinshipResultandKinshipMatchestypes, CLI rendering, and full-precision JSON. Score jointly called markers, preserve negative estimates, and report zero-denominator cases as unscorable. Existing identity-style plugins retain API version 1 and their result types. See the plugin contract.Usage and compatibility
The KING usage guide documents panel/header requirements, filtering, output fields, and benchmark commands. The initial scope is human autosomal, diploid, biallelic SNVs. Existing allele-only databases still support
identity; KING requires re-ingestion from source VCFs because reference calls cannot be reconstructed.shared-variantsrejects kinship results.The default between-family score is distinct from the within-family diagnostic used for VCFtools comparison. On the eight-marker oracle, these are 0.0 and 0.1, respectively; VCFtools reports 0.1. Missing-data tests explicitly verify the difference between the plugin's pairwise callable denominator and VCFtools' individual heterozygote totals.
Validation
Verified at branch head with Python 3.13.5:
ruff check .andruff format --check .— passed.pytest -q --no-cov— 137 passed, 3 opt-in tests skipped.Tests cover VCF ingestion, QC boundaries, missing/reference/partial calls, panel validation, rollback, statistical oracles, candidate restrictions, ranking, batching, streaming, CLI/JSON compatibility, synthetic generation, and phase accounting. The existing 1000 Genomes network/SeqRepo integration test was not run. The validation report records the earlier numerical validation and smoke measurements; the performance ADR records subsequent instrumentation experiments.
Limits
Biological acceptance remains pending a genome-wide panel with independently verified replicate and pedigree labels. Synthetic results establish implementation correctness and instrumentation, not biological accuracy or representative throughput. Timings exclude upstream preparation and VRS annotation. VCFtools emits ordered pairs and diagonals, while the plugin all-pairs path emits distinct unordered pairs; reported workload ratios do not establish universal performance superiority. Unscorable search results remain retained in memory.