Skip to content

[ENH]: Addition of initial ASV benchmarking suite with synthetic data - #503

Closed
KaranSinghDev wants to merge 5 commits into
dswah:mainfrom
KaranSinghDev:feature/add-asv-benchmarks
Closed

KaranSinghDev wants to merge 5 commits into
dswah:mainfrom
KaranSinghDev:feature/add-asv-benchmarks

Conversation

@KaranSinghDev

@KaranSinghDev KaranSinghDev commented Mar 8, 2026 •

Copy link
Copy Markdown

Fixes #471

This PR implements a formal benchmarking infrastructure using Airspeed Velocity (asv). This fulfills the long-standing request mentioned in #99 to provide a mechanism for tracking performance history and preventing speed regressions.

Implementation

  1. Infrastructure: Added asv.conf.json configured for isolated virtualenv builds and direct pip installation for cross-platform stability.
  2. Synthetic Data Strategy: Benchmarks generate data in-memory using numpy to avoid dependencies on pandas or external CSV files and ensures we measure algorithmic throughput rather than disk I/O speed.
  3. Thread Isolation: Explicitly set OMP_NUM_THREADS=1 and MKL_NUM_THREADS=1 to ensure timing remains consistent across different hardware configurations.
  4. Initial Suites:
    • LinearGAMFit: Measures fitting, prediction, and gridsearch latency.
    • PoissonGAMFit: Benchmarks the performance of the iterative PIRLS loop using a stable mathematical signal to ensure rapid convergence and prevent timeouts.

Verification

I have thoroughly verified in an isolated environment. All benchmarks execute in under 300ms, ensuring the suite is fast enough for CI/CD integration.

Sample Local Output:

· Discovering benchmarks
· Running 4 total benchmarks (1 commits * 1 environments * 4 benchmarks)
[ 62.50%] ··· bench_fit.LinearGAMFit.time_fit           83.3±10ms
[ 75.00%] ··· bench_fit.LinearGAMFit.time_gridsearch    280±200ms
[ 87.50%] ··· bench_fit.LinearGAMFit.time_predict       17.8±4ms
[100.00%] ··· bench_fit.PoissonGAMFit.time_fit         202±100ms

The tests have been done on Python 3.10/3/13. Linux (Ubuntu), with CPU: i7 12700h.

@KaranSinghDev
KaranSinghDev force-pushed the feature/add-asv-benchmarks branch from 780ec18 to 6ef3213 Compare March 8, 2026 10:52
@hritikkumarpradhan

Copy link
Copy Markdown

Hi @KaranSinghDev, thanks for putting this ASV infrastructure together!

I’m focusing on heavily optimizing pyGAM's matrix operations, and having this automated suite in place is exactly what the project needs to safely track mathematical performance regressions.

I was looking through your benchmarks/ directory setup, and it looks incredibly clean. I've been locally profiling a severe O(N^3) memory bottleneck in the Effective Degrees of Freedom (EDoF) calculation (np.diagonal(U1.dot(U1.T))) for overparameterized datasets (p > n). I have a vectorized O(N) fix ready, but we desperately need a baseline in ASV to permanently prove the memory reduction.

Here is a snapshot of the memory profile I ran locally, showing the ~760 MiB spike on the legacy calculation versus the near-zero allocation of the vectorized approach:

Generating synthetic U1 matrix with shape (10000, 500)...

--- Running Legacy Calculation ---
Legacy Execution Time: 1.2450 seconds
Filename: benchmark_edof.py

Line #    Mem usage    Increment  Occurrences   Line Contents
=============================================================
    22     76.4 MiB      0.0 MiB           1       start_time = time.time()
    23                                             
    24                                             # The bottleneck:
    25    839.5 MiB    763.1 MiB           1       edof = np.diagonal(u1_matrix.dot(u1_matrix.T))
    26                                             
    27    839.5 MiB      0.0 MiB           1       end_time = time.time()

--- Running Optimized Calculation ---
Optimized Execution Time: 0.0120 seconds
Filename: benchmark_edof.py

Line #    Mem usage    Increment  Occurrences   Line Contents
=============================================================
    38     76.5 MiB      0.0 MiB           1       start_time = time.time()
    39                                             
    40                                             # The fix:
    41    114.6 MiB     38.1 MiB           1       edof = (u1_matrix**2).sum(axis=1)
    42                                             
    43    114.6 MiB      0.0 MiB           1       end_time = time.time()
    

I’d love to translate my profiling scripts into an EDoFBenchmark ASV class and contribute them directly to your suite.

Would you be open to me opening a PR directly against your karan-asv-branch to add this? That way, we can get your infrastructure and the first major memory stress-test merged together as a complete package! Let me know what you think.

@dswah

dswah commented Apr 16, 2026

Copy link
Copy Markdown
Owner

This contribution looks very cool!
I think that for it to be completely useful it needs to be added to the git workflows!
Can we add that to this PR?

@KaranSinghDev
KaranSinghDev force-pushed the feature/add-asv-benchmarks branch from f15deb8 to b058c70 Compare April 21, 2026 04:26
@KaranSinghDev

KaranSinghDev commented Apr 21, 2026 •

Copy link
Copy Markdown
Author

@dswah I've updated the PR with the GitHub Actions workflow. I also merged and refactored @hritikkumarpradhan's EDoF benchmark.

I ensured the matrix dimensions are locked as constants and added a random seed for stability. Profiling locally confirmed it safely isolates the O(N^3) bottleneck (~800MB spike) without exceeding the 7GB memory limit of the GitHub runners.

As implementation, the files made and what they do :
1.asv.conf.json: Global configuration that establishes the benchmarking environment and restricts history tracking to the main branch to ensure CI stability.

2.benchmarks/bench_fit.py: Macro-benchmarks measuring fit(), predict(), and gridsearch() latency.

  1. benchmarks/bench_edof.py: A micro-benchmark isolating the matrix bottleneck in the EDoF calculation.

4..github/workflows/benchmark.yml: The CI automation engine. It uses the uv package manager and asv continuous to automatically compare PR performance against the main branch baseline.

Let me know your thoughts on this .

P.S. Regarding the changes such as individual profiling of components, imporving dashbaord etc. I was thinking of making a new PR and not adding them in it to keep it easy to understand this PR and its work

@hritikkumarpradhan

Copy link
Copy Markdown

Thanks @KaranSinghDev for the clean integration!
Just to add context from my side — the bench_edof.py currently captures the legacy O(N³) bottleneck (~800 MiB spike) as the baseline. I'll be opening a follow-up PR to add the optimized (U1**2).sum(axis=1) counterpart benchmarks once this is merged, so the ASV suite will have both sides of the comparison.

Agreed on keeping this PR focused — the CI workflow + benchmarks are already a solid, self-contained addition.
Looking forward to @dswah's review!

@KaranSinghDev KaranSinghDev closed this by deleting the head repository Jul 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[ENH] Implement automated benchmarking suite

3 participants