Skip to content

Write timeouts, and faster reads on SQLite and PostgreSQL - #50

Merged
fungiboletus merged 7 commits into
mainfrom
timeouts-and-read-performance
Oct 5, 2026
Merged

fungiboletus merged 7 commits into
mainfrom
timeouts-and-read-performance

Conversation

@fungiboletus

Copy link
Copy Markdown
Member

What

Found while loading a real 1.33 M-sample series (2.6 years, one request) into the three main backends. Four problems, each fixed with a measured before and after.

1. Write timeouts, and a duplicated import

One 30 s timeout covered every request. A single 1.33 M-sample Arrow request into TimescaleDB hit it: a request is stored batch by batch (8,192 samples, one transaction each), so the 504 left 27 batches committed, the SDK retried, and the series held 1,351,680 rows for 688,128 distinct timestamps.

  • New SENSAPP_HTTP_WRITE_TIMEOUT_SECONDS (300 s) for /publish, the InfluxDB write and the Prometheus remote write. A full 64 MiB body took about 30 s on the slowest backend measured, so this leaves ten times that.
  • SENSAPP_HTTP_SERVER_TIMEOUT_SECONDS (reads and the rest): 30 s to 120 s. A raw read is capped, an aggregation scans its window (0.17 to 0.98 s per 1.33 M samples), so 120 s covers about 100 M samples on the slowest query measured.
  • Python SDK: timeout=125 for reads and write_timeout=330 for writes (was 10 s for everything), a little above the server's so its 504 reaches the client.
  • Replayed against TimescaleDB: one request, 51.6 s, 1,331,266 distinct rows, no 504.
  • A timed-out write still leaves its first batches stored (the SDK documents that retries can duplicate). Making a request atomic is in ideas/atomic-write-requests.md.

2. SQLite: a time window or the last sample scanned the whole series

(? IS NULL OR timestamp_us >= ?) keeps SQLite from using the time column of its index. Now timestamp_us >= COALESCE(?, -i64::MAX). /last 54 ms to 0.5 ms, 168 hourly buckets over a week 65 to 5 ms, a raw week 78 to 26 ms.

3. PostgreSQL: aggregations and plan flips

  • Buckets are computed on the integer timestamp instead of date_bin(to_timestamp(ts / 1e6)): identical buckets on 1.4 M random timestamps (negative origins included), 6x faster in psql.
  • Same COALESCE time bounds as SQLite.
  • After five executions PostgreSQL uses a generic plan for a prepared statement, which made windows or reads without bounds slow whichever way the bounds are written. The connections now plan with their parameters (force_custom_plan, as the TimescaleDB backend already does). That made the load 2.7x slower (the bulk INSERT .. unnest re-planned for every batch), so write transactions go back to the default plan cache (SET LOCAL plan_cache_mode = auto): load unchanged at 19.7 s.
  • BRIN stays.

4. Notes

done/ (three finished tasks), ideas/ (atomic write requests, TimescaleDB write plan cache, a first numbers note with what to cover in the planned benchmark).

Measurements

1.33 M-sample series, release build, one laptop, reads 60 s after the load with no manual ANALYZE (a BRIN index has no usable summary right after a bulk load). Median ms, PostgreSQL before to after:

query before after
step=1d avg, whole history 561 135
step=1h avg, whole history 579 153
step=1h max, whole history 419 146
step=1d last, whole history 561 312
step=1h avg, one week 21 9.6
/last 65 47 (BRIN reads the table)
load, 1.33 M samples 18 to 19 s 19.7 s

This is a quick single-run comparison, not the planned benchmark; ideas/read-and-write-performance-first-numbers.md has the full tables, the BRIN against B-tree measurements and the caveats.

Behaviour changes to know

  • Default timeouts change: reads 30 s to 120 s, writes 30 s to 300 s, SDK 10 s to 125 s / 330 s.
  • ClickHouse already buckets on the integer (intDiv) and has no IS NULL OR bounds. TimescaleDB is untouched: it prunes chunks at run time even with a forced generic plan.
  • For samples before the origin (before 1970 with no start), PostgreSQL now floors the bucket like date_bin did, while SQLite and ClickHouse truncate toward zero. Not reachable with real data.
  • When merging with frontend-good-enough: its drift test compares frontend/openapi.json with the server's OpenAPI document, and the 504 text of the three writes changed here. UPDATE_OPENAPI=1 cargo test frontend_openapi_document fixes it.

Tests

  • New backend-generic time_window_reads (all eight sample types; no bound, a start, an end, both, between two samples, a limit, outside the data; reads of one series, the latest sample, and by labels), plus a guard test against the IS NULL OR form returning in sqlite/ and postgresql/. Swapping two binds makes it fail. The only "time range" test before it asserted is_some() || is_none().
  • real_router::writes_have_a_timeout_of_their_own, SDK tests on the two timeouts.
  • Run locally: full suite on SQLite (246 + 301) and on PostgreSQL (246 + 291), cargo clippy --all-targets -D warnings with postgres,sqlite,timescaledb,clickhouse,rrdcached,bigquery and with the default features, cargo fmt --check, SDK ruff and pytest (69 passed). The new window test also ran on TimescaleDB before the PostgreSQL changes. Not run locally: the full TimescaleDB, ClickHouse, DuckDB and BigQuery suites, left to CI.

🤖 Generated with Claude Code

fungiboletus and others added 7 commits October 4, 2026 20:14
…operation

One 30 s timeout covered every request. A 1.33 M-sample Arrow request into TimescaleDB hit it:
504 after 27 batches were committed, the SDK retried, and the series held 1,351,680 rows for
688,128 distinct timestamps.

- SENSAPP_HTTP_WRITE_TIMEOUT_SECONDS (300 s) for /publish, the InfluxDB write and the Prometheus
  remote write; the reads keep 30 s, the vacuum its own.
- SDK: timeout=35 for reads, write_timeout=330 for writes, a little above the server's.
- Docs: why 300 s (measured), a write is stored batch by batch, prefer requests of 100k-200k samples.
- Replayed against TimescaleDB: one request, 51.6 s, 1,331,266 distinct rows, no 504.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…aleDB

A quick single-run comparison (one series of 1.33 M samples), and one finding from it: SQLite scans
the whole series for a time window or the last sample, because of the `? IS NULL OR` form of its
queries (133 ms for a 1-hour window, 0.06 ms as a plain range).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…he last sample

`(? IS NULL OR timestamp_us >= ?)` keeps the planner from using the time column: it read the whole
series for any window. `timestamp_us >= COALESCE(?, -i64::MAX)` means the same and is plannable.
On a series of 1.33 M samples: /last 54 ms to 0.5 ms, a week of 1-hour buckets 65 ms to 5 ms,
a raw week 78 ms to 26 ms. In the sqlite3 shell, one hour: 175 ms to 0.1 ms.

New backend-generic test (all sample types, every combination of bounds); it fails when a bind is
swapped. Passes on SQLite, PostgreSQL and TimescaleDB.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…he SDK

A raw read is bounded, an aggregation scans its window: 0.17 to 0.98 s per 1.33 M samples, so
120 s covers about 100 M samples on the slowest query measured. Writes stay at 300 s / 330 s.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…PostgreSQL ones

SQLite after the index fix. PostgreSQL measured settled: its BRIN index is not the cause of slow
windows (a week in 22 ms, like TimescaleDB). What hurt: BRIN summaries lag after a bulk load (91 ms
until autovacuum catches up), the `$2 IS NULL OR ...` form of the time bounds turns into a full-series
scan in a generic plan (115 ms against 1.8 ms for the COALESCE form), and the bucket expression is 6
times slower than integer arithmetic (1.25 s against 0.21 s in psql). A B-tree wins on the latest
sample and on interleaved series, and cost nothing visible on the load.

New idea: ideas/postgresql-windows-and-buckets.md.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…ans for reads only

- Buckets on the integer instead of date_bin over to_timestamp(timestamp_us / 1e6): the same buckets
  (1.4 M random timestamps, 0 differences), 6 times faster in psql.
- Time bounds as `timestamp_us >= COALESCE($2, -i64::MAX)`; a guard test keeps the `IS NULL OR` form
  out of the SQLite and PostgreSQL sources.
- The connections force custom plans (a generic plan after five executions made windows or reads
  without bounds slow, whichever way the bounds are written), and write transactions go back to the
  default (`SET LOCAL plan_cache_mode = auto`): custom plans of the bulk unnest inserts cost 3x on the load.

1.33 M samples, 60 s after the load: 1d avg 561 to 135 ms, 1h avg 579 to 153, 1h max 419 to 146,
1d last 561 to 312, 1h buckets over a week 21 to 9.6, load 19.7 s (unchanged). BRIN is kept.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
…DB's write plans

PostgreSQL after its fixes next to TimescaleDB and SQLite. What explained PostgreSQL's reads: the
generic plans of prepared statements, not BRIN. New idea: TimescaleDB's slow load is mostly its forced
custom plans (23.5 s instead of 58 s in a throwaway experiment).

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
@fungiboletus
fungiboletus merged commit 7ad27e2 into main Oct 5, 2026
17 checks passed
@fungiboletus
fungiboletus deleted the timeouts-and-read-performance branch October 5, 2026 07:57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant