Skip to content

Rebalancer permanently misses an on-chain receive when its Confirmed update lands one sync cycle late #98

Description

@hash-money

Summary

needs_onchain_rebalance (orange-sdk/src/rebalancer.rs @ 5a540f66) selects sweep-trigger payments with a timestamp window: latest_update_timestamp > onchain_sync_time (lines 163 and 211), while onchain_sync_time is unconditionally advanced to the node's latest_onchain_wallet_sync_timestamp on every tick that observes a new sync (line 173). If a payment's Confirmed update is first visible on a tick where the stored onchain_sync_time has already advanced past that payment's latest_update_timestamp, the payment fails the window on that tick — and every later tick, because the stored time only grows. The wallet then loops "Detected onchain sync with balance updates, but no new onchain payments found" (line 261) with a spendable balance above rebalance_min that is never swept.

Reproduction / evidence (regtest, bitcoind v29 chain source)

While hardening our integration battery we hit this as a ~1-in-4 flake on "fund on-chain → auto-sweep opens LSP channel", and isolated it:

  • Fund a fresh wallet on-chain, confirm the deposit, then stop producing blocks. If the race is lost, the wallet logs the line-261 warning on every subsequent sync and never rebalances.
  • The miss is permanent: neither restarting the wallet nor sending a fresh unrelated on-chain nudge recovers the original payment (the new payment can trigger its own sweep of the whole spendable balance, but only if its update wins the race).
  • Continuous block production masks the bug: each new confirmation refreshes latest_update_timestamp, giving the payment repeated chances to pass the window. That's why it looks like a rare flake on regtest and would look like sporadic "stuck on-chain funds" in the field — e.g. a mobile app backgrounded/offline across the confirmation, syncing after the update timestamp is already behind the stored sync time.

Why this matters

The on-chain→Lightning sweep is the load-bearing behavior of the two-tier model: a missed sweep leaves user funds resting on-chain indefinitely with no user-visible error and no recovery path short of another deposit.

Suggested fix shape

Eligibility should be state-based rather than timestamp-window-based: a payment that is Inbound + Succeeded + Onchain Confirmed and not already promoted is a valid trigger regardless of when its update landed. The dedupe needed for that already exists in this function — the can_mark_as_trigger check against tx_metadata (lines 222-226) — so the latest_update_timestamp > onchain_sync_time clause at line 211 could be dropped (or kept only as a fast-path) with the metadata check preventing re-promotion. The event-emission filter at line 163 needs the same treatment or a parallel dedupe to avoid duplicate OnchainPaymentReceived events.

Happy to submit a PR along those lines from our fork if that's welcome.

Activity

  1. benthecarman commented on Jul 22, 2026

    @benthecarman
    Collaborator

    The current way I do on-chain events is pretty hacky. Once this PR is in, I can refactor them to be much cleaner: lightningdevkit/ldk-node#448 and should be able to prevent this.

  2. hash-money commented on Jul 22, 2026

    @hash-money
    ContributorAuthor

    Makes sense — a state-based refactor on top of ldk-node#448's onchain-transactions-in-events is a cleaner shape than patching the timestamp window. We'll keep our regtest mitigation (which just keeps confirmations flowing) in the meantime and are happy to re-run our reproduction against the refactor when it lands if that's useful for validation.

  3. hash-money commented on Sep 10, 2026

    @hash-money
    ContributorAuthor

    A second shape of the same trigger gap, seen in the field this week — adding it here rather than as a new issue since it bears on the state-based refactor you mentioned, but happy to split it out if you'd prefer.

    What we hit. A staging wallet (Mutinynet, Cashu trusted tier — though nothing below is provider-specific) had 323,575 sats sitting on-chain for several days. Its channel had been cooperatively closed by the user earlier, rebalancing was later re-enabled, and every sync since logged Detected onchain sync with balance updates, but no new onchain payments found (rebalancer.rs @ ed29c6c, the None arm after the list_payments() filter). The balance was above rebalance_min and far under the channel cap; it simply had no new inbound on-chain payment behind it, so needs_onchain_rebalance never selected a trigger.

    Workaround that worked. We deposited 10,000 sats on-chain from another wallet. That receive tripped the trigger and the rebalancer spliced the full standing balance into the existing LSP channel (splice funding confirmed ~4 min after the deposit). Datapoint for this issue's race: the deposit confirmed at 05:05:26, the syncs at 05:05:57 and 05:07:17 still logged the warning, and the 05:08:37 sync caught it — continuous signet blocks gave it the second chance you'd expect.

    The question. #2 says rebalancing is disabled after a close until re-enabled, which makes sense. What we didn't anticipate is that after re-enabling, balance with no new receive behind it — a close output, or anything older than the sync cursor after a restart — is never evaluated again. Is that intended (only future receives graduate, and the wallet should move close outputs itself), or would you consider re-enable / startup doing a one-time evaluation of standing spendable balance, with a synthetic trigger (e.g. the close txid) where no payment exists to promote? If it's the former, that's fine — we'll handle it on our side and you can close this part out. We can re-run this against the refactor on our staging wallet whenever it's useful.

  4. hash-money commented on Sep 15, 2026

    @hash-money
    ContributorAuthor

    Checked #107 (70dca69) against this while reviewing #108, so the thread has current facts.

    The on-chain trigger window is unchanged: needs_onchain_rebalance still selects on latest_update_timestamp > onchain_sync_time (orange-sdk/src/rebalancer.rs:205 at 70dca69; :159 at ed29c6c). What #107 does change there: the sync-time store now happens after the event loop (:239), so a failed OnchainPaymentReceived add is retried on the next tick; and a new early return when the total/spendable on-chain balances haven't moved since the last scan (:191). For the wedged state described above, that early return stops the no new onchain payments found line (:303) from repeating without re-evaluating the standing balance — quieter, not fixed.

    ldk-node#448 is still open, so the state-based refactor planned on top of it remains blocked upstream. We'll keep our regtest mitigation in place and re-run the repro against whichever lands first. Nothing to do here — just keeping the record straight.

  5. benthecarman commented on Sep 15, 2026

    @benthecarman
    Collaborator

    Ah yeah I see ldk-node#448 got delayed until next release but #107 was not targeting to fix this issue

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions