Status: proposal — design agreed, not yet implemented (2026-07-03).
Issue #877 asks for a way for clients to detect that their streams have gone stale after a game-state change (quickload etc.), proposing a global protocol-level "all streams invalidated" notification driven by a server-side state generation counter.
We prefer a granular design: when an individual stream is no longer valid — detected by its underlying RPC throwing an exception — the client is notified and the stream is removed, per stream. A stream is not necessarily tied to a single game object, so "object destroyed" is not the right trigger; "the underlying call threw" is.
Key finding: the server already does this for procedure-call streams.
ProcedureCallStream.UpdateInternal (core/src/Service/ProcedureCallStream.cs:45-53)
catches exceptions, converts them to an Error via Services.Instance.HandleException,
and sets Changed; Core.StreamServerUpdate (core/src/Core.cs:550-551, 566-572)
sends that error to the client exactly once, then removes the stream. Existing tests
cover it (client/python/krpc/test/test_stream.py:201-237). The quickload scenario
genuinely triggers this path: e.g. Vessel.InternalVessel
(service/SpaceCenter/src/Services/Vessel.cs:67-69) re-resolves by Guid on each access
and FlightGlobalsExtensions.GetVesselById throws when the vessel is gone.
The actual gaps this design fixes:
EventStreamhas no exception handling (core/src/Service/EventStream.cs:35-40): a throwing event predicate propagates out ofstream.Update()(Core.cs:537, also unguarded) and abortsStreamServerUpdatefor all clients, every frame, forever — no error sent, stream never removed.- Clients don't treat an error result as invalidation: the stream stays in their
local manager map forever;
wait()blocks forever after an error; C# and C++ have crash bugs when a callback is registered on an erroring stream (C# update thread dies from an InvalidCastException; C++ callsstd::terminate). - The "error result ⇒ stream removed" convention is undocumented in the protocol docs.
Decisions:
- Protocol: implicit convention — no
.protochange; document that aStreamResultcarrying an error means the server has removed the stream. The server always removes on error, so this is unambiguous;bool removedcan be added proto3-compatibly later if transient (non-fatal) stream errors are ever wanted. - Client scope: all four stream-capable clients — Python, C#, Java, C++. (lua/cnano have no stream support; websockets/serialio are conformance-test packages.)
- Typed
ObjectDestroyedException: deferred to a follow-up tied to issue #771.
Mirror ProcedureCallStream.UpdateInternal:
public override void UpdateInternal() {
if (continuation != null) {
try {
if (continuation())
Trigger();
} catch (System.Exception e) {
var result = StreamResult.Result;
result.Reset();
result.Error = Services.Instance.HandleException(e);
Changed = true;
continuation = null; // poison, like StreamContinuation does
}
}
if (shouldRemove)
Core.Instance.RemoveStream (Id);
}Reuses Services.Instance.HandleException (core/src/Service/Services.cs:367-387) —
already maps [KRPCException] types and honors VerboseErrors. Core's existing
HasError → removeStreams logic then notifies + removes with no further changes.
Sent() (lines 51-56) unboxes (bool)result.Value; after the error path calls
result.Reset(), Value is null → NRE in Core's Sent() loop (Core.cs:560-562).
Change to:
if (result.HasValue && (bool)result.Value)
result.Value = false;Wrap stream.Update() in try/catch: on System.Exception, log and set an error
ProcedureResult on the stream (the Stream.Result setter in
core/src/Service/Stream.cs sets Changed when HasError). Guarantees no stream
implementation can ever abort the update loop for all clients.
doc/src/communication-protocols/messages.rst — in the Streams / StreamResult
section (~lines 179-226), document:
- When a stream's RPC throws, the
resultcontains anerror; the server sends it exactly once and then removes the stream. - Clients must treat such a stream as invalid: no further updates will arrive for its
id; calling
KRPC.RemoveStreamis unnecessary but harmless (server tolerates unknown ids). - Note this is how clients detect streams invalidated by game-state changes (quickload/revert) — the #877 scenario.
For each client, on receiving an error StreamResult:
- Build/store the exception as the stream value (existing behavior — keep).
- Mark the StreamImpl removed and delete it from the manager map (no leak; later updates for the id are already skipped; server ids are never reused).
- Reading the value keeps raising the stored error forever — NOT "Stream does not exist".
wait()/start(wait=True)on an invalidated stream raises the stored error immediately instead of blocking forever.remove()on an invalidated stream is an idempotent no-op that preserves the stored error (no RPC, no overwrite with "Stream does not exist").- Callbacks: deliver the exception where the callback type permits (Python/Java);
where it can't (C# typed callbacks, C++
std::function<void(std::string)>), skip callbacks for the terminal error and document detection viaGet()/wait()— fixing the current crashes.
streammanager.pyStreamImpl: add_removedflag +removedproperty.StreamManager.update()(lines 160-180): in theHasField("error")branch, after_update_stream(...), mark removed anddel self._streams[result.id].StreamImpl.remove()(lines 96-99): early-return if_removed; set_removedfor user-initiated removal too.stream.pyStream.wait()/start(): if impl removed and stored value is an Exception, raise it before waiting.event.pyEvent.wait()(line 38): don't reset value toFalsewhen removed/errored — currently clobbers the stored error then blocks forever.
StreamImpl.cs:Removedproperty; idempotent, error-preservingRemove().StreamManager.csUpdate()(lines 90-109): on error, mark removed +streams.Remove(id); wrap callback invocations in try/catch. FixStream.cstyped-callback wrapper (~line 137): skip when valueis Exception(currently InvalidCastException kills the update thread).Stream.csWait()/Start(wait): if removed, rethrow stored exception instead ofMonitor.Wait.
StreamImpl.java:removedflag; idempotentremove().StreamManager.javaupdate()(lines 68-89): onhasError(), mark removed +streams.remove(id); per-callback try/catch so a throwing callback can't kill the update thread.Stream.javawaitForUpdate*()/startAndWait(): if removed, throw stored error immediately (reuseget()'s unwrap logic, lines 100-107).
stream_impl.hpp/.cpp:removedflag +has_error()accessor.stream_manager.cppupdate()(lines 70-90): on error, mark removed +streams.erase(id); guard the callback dispatch withif (!stream->has_error())— currentlyget_data()rethrows inside the update thread →std::terminate.stream.hppwait()/start(true): if impl removed, callimpl->get_data()(rethrows) instead of waiting.
Add a throwing event next to OnTimer (~line 452), e.g.
ThrowingEvent(uint updatesBeforeThrow = 0) whose predicate throws
InvalidOperationException after N updates — covers both immediate and later variants.
Client bindings regenerate via the existing clientgen / services-testservice
targets in each client BUILD file.
Extend the existing exception tests (lines 201-237):
- after the error raises, the id is gone from
conn._stream_manager._streams; - repeated reads keep raising the same error;
wait()raises immediately (guard with timeout — must not hang);add_callbackcallback receives the exception object;remove()after invalidation is a no-op, error still raised;test_event.py: throwing event —event.wait()raises, doesn't hang.
Mirror the Python cases in client/csharp/test/StreamTest.cs + EventTest.cs,
client/java/test/krpc/client/StreamTest.java + EventTest.java,
client/cpp/test/test_stream.cpp + test_event.cpp. Critical regressions to cover:
error arriving while a callback is registered must not kill the update thread (C#) or
terminate the process (C++); wait-after-error must not block.
Small unit test for EventStream.UpdateInternal + Sent after a throwing continuation
under core/test/Service/ (the exception path doesn't need Core.Instance).
bazel test //core:testbazel test //client/python:test //client/csharp:test //client/java:test //client/cpp:test— each spins up TestServer; new tests exercise the full path: server exception → error StreamUpdate → client invalidation.- Manual #877 scenario in KSP: stream on a vessel, quicksave, destroy vessel,
quickload → client gets
ArgumentException("No such vessel ...")raised once from the stream;wait()doesn't hang; server log shows "Removing stream as it returned an error"; no per-frame exception spam / frame-rate degradation. - Regression: existing event tests (
OnTimer,OnTimerUsingLambda) still pass — they exercise theshouldRemovepath through the modifiedUpdateInternal.
- Typed
ObjectDestroyedExceptionin SpaceCenter lookups (FlightGlobalsExtensions.GetVesselByIdetc.) so clients can distinguish "object gone — recreate streams" from bugs; tie to issue #771. - Error-less server-side removal:
Event.remove()(EventStream.cs:47-49) removes the stream with no client notification — the implicit convention doesn't cover it. Abool removedStreamResult field would; revisit if it becomes a problem. - Respond on issue #877 explaining the granular approach vs. the proposed global generation counter.
- Python
StreamManager.updateholds_update_lockwhile invoking user callbacks; deleting from_streamsunder the same RLock is safe, and a re-entrantstream.remove()from a callback becomes a no-op via the membership check — add a test.