Repository navigation
Rebalancer permanently misses an on-chain receive when its Confirmed update lands one sync cycle late #98
Description
Activity
The current way I do on-chain events is pretty hacky. Once this PR is in, I can refactor them to be much cleaner: lightningdevkit/ldk-node#448 and should be able to prevent this.
Makes sense — a state-based refactor on top of ldk-node#448's onchain-transactions-in-events is a cleaner shape than patching the timestamp window. We'll keep our regtest mitigation (which just keeps confirmations flowing) in the meantime and are happy to re-run our reproduction against the refactor when it lands if that's useful for validation.
A second shape of the same trigger gap, seen in the field this week — adding it here rather than as a new issue since it bears on the state-based refactor you mentioned, but happy to split it out if you'd prefer.
What we hit. A staging wallet (Mutinynet, Cashu trusted tier — though nothing below is provider-specific) had 323,575 sats sitting on-chain for several days. Its channel had been cooperatively closed by the user earlier, rebalancing was later re-enabled, and every sync since logged
Detected onchain sync with balance updates, but no new onchain payments found(rebalancer.rs @ed29c6c, theNonearm after thelist_payments()filter). The balance was aboverebalance_minand far under the channel cap; it simply had no new inbound on-chain payment behind it, soneeds_onchain_rebalancenever selected a trigger.Workaround that worked. We deposited 10,000 sats on-chain from another wallet. That receive tripped the trigger and the rebalancer spliced the full standing balance into the existing LSP channel (splice funding confirmed ~4 min after the deposit). Datapoint for this issue's race: the deposit confirmed at 05:05:26, the syncs at 05:05:57 and 05:07:17 still logged the warning, and the 05:08:37 sync caught it — continuous signet blocks gave it the second chance you'd expect.
The question. #2 says rebalancing is disabled after a close until re-enabled, which makes sense. What we didn't anticipate is that after re-enabling, balance with no new receive behind it — a close output, or anything older than the sync cursor after a restart — is never evaluated again. Is that intended (only future receives graduate, and the wallet should move close outputs itself), or would you consider re-enable / startup doing a one-time evaluation of standing spendable balance, with a synthetic trigger (e.g. the close txid) where no payment exists to promote? If it's the former, that's fine — we'll handle it on our side and you can close this part out. We can re-run this against the refactor on our staging wallet whenever it's useful.
Checked #107 (
70dca69) against this while reviewing #108, so the thread has current facts.The on-chain trigger window is unchanged:
needs_onchain_rebalancestill selects onlatest_update_timestamp > onchain_sync_time(orange-sdk/src/rebalancer.rs:205at70dca69;:159ated29c6c). What #107 does change there: the sync-time store now happens after the event loop (:239), so a failedOnchainPaymentReceivedadd is retried on the next tick; and a new early return when the total/spendable on-chain balances haven't moved since the last scan (:191). For the wedged state described above, that early return stops theno new onchain payments foundline (:303) from repeating without re-evaluating the standing balance — quieter, not fixed.ldk-node#448 is still open, so the state-based refactor planned on top of it remains blocked upstream. We'll keep our regtest mitigation in place and re-run the repro against whichever lands first. Nothing to do here — just keeping the record straight.
Ah yeah I see ldk-node#448 got delayed until next release but #107 was not targeting to fix this issue
Summary
needs_onchain_rebalance(orange-sdk/src/rebalancer.rs @5a540f66) selects sweep-trigger payments with a timestamp window:latest_update_timestamp > onchain_sync_time(lines 163 and 211), whileonchain_sync_timeis unconditionally advanced to the node'slatest_onchain_wallet_sync_timestampon every tick that observes a new sync (line 173). If a payment'sConfirmedupdate is first visible on a tick where the storedonchain_sync_timehas already advanced past that payment'slatest_update_timestamp, the payment fails the window on that tick — and every later tick, because the stored time only grows. The wallet then loops"Detected onchain sync with balance updates, but no new onchain payments found"(line 261) with a spendable balance aboverebalance_minthat is never swept.Reproduction / evidence (regtest, bitcoind v29 chain source)
While hardening our integration battery we hit this as a ~1-in-4 flake on "fund on-chain → auto-sweep opens LSP channel", and isolated it:
latest_update_timestamp, giving the payment repeated chances to pass the window. That's why it looks like a rare flake on regtest and would look like sporadic "stuck on-chain funds" in the field — e.g. a mobile app backgrounded/offline across the confirmation, syncing after the update timestamp is already behind the stored sync time.Why this matters
The on-chain→Lightning sweep is the load-bearing behavior of the two-tier model: a missed sweep leaves user funds resting on-chain indefinitely with no user-visible error and no recovery path short of another deposit.
Suggested fix shape
Eligibility should be state-based rather than timestamp-window-based: a payment that is
Inbound + Succeeded + Onchain Confirmedand not already promoted is a valid trigger regardless of when its update landed. The dedupe needed for that already exists in this function — thecan_mark_as_triggercheck againsttx_metadata(lines 222-226) — so thelatest_update_timestamp > onchain_sync_timeclause at line 211 could be dropped (or kept only as a fast-path) with the metadata check preventing re-promotion. The event-emission filter at line 163 needs the same treatment or a parallel dedupe to avoid duplicateOnchainPaymentReceivedevents.Happy to submit a PR along those lines from our fork if that's welcome.