Skip to content

Roadmap: diagnostic evidence integrity, agent evaluation, and operational coverage #919

Description

@pedrosakuma

Mission and outcome

Follow-up to #551, based on the mission/completeness reassessment of 2026-09-11.

The mission is on-demand performance diagnosis of running .NET applications, driven by an LLM through MCP or by humans/automation through the Core-only CLI and BenchmarkDotNet diagnoser. Standard diagnostics require no target code changes or prior instrumentation; method-parameter capture remains an explicit, opt-in instrumentation exception.

The next priority is diagnostic trust, not a larger tool catalog: preserve the distinction between observed facts, inferred explanations, and unavailable evidence from collection through the final diagnosis.

This issue formalizes the agreed roadmap. It does not authorize enabling experimental collectors, changing runtime support, removing protections without evidence, or introducing new MCP tools. The delivery phases below are internal to this roadmap; they do not renumber the repository's existing Phase 16.

Evidence and current limits

Finding Evidence Classification
Unknown ThreadPool adjustment reasons can become Starvation, including through synthetic worker counts; downstream summaries count them as starvation adjustments. EventPipeThreadPoolCollector.NormalizeHillClimbing, CollectionQueryDispatcher, ThreadPoolComparableProjector. The collector emits an inference note, but the reason is still promoted. Concrete local semantic issue.
Playbooks still equate growing connection queues or generic hill-climbing activity with confirmed starvation. investigation-playbooks.md, sections 1 and 1a, versus the bounded-hypothesis contract in the README. Concrete documentation/interpretation inconsistency.
Scenario evidence evaluation is not end-to-end evaluation of the agent. Four scenario manifests; live tests call ScenarioEvaluator.CreateReport without an interpretation, resulting in Interpretation.NotRun. #736 delivered an explicitly advisory keyword-based mapper. Known evaluation scope gap, not absence of all evaluation.
A Linux scenario quarantine remains after the ten Core quarantines were removed. ScenarioLiveTests.cs: culture-lookup remains WindowsOnlyFact citing #147; the isolated workflow omits Linux culture-lookup. Concrete remaining coverage restriction; requires its own reassessment.
Compatibility claims are broader than the mechanisms continuously exercised. runtime-version-compat-matrix.md lists live ClrMD, deeper drilldowns and method-parameter cross-version gaps. perf-compat-matrix.md documents manual rather than automated perf sidecar coverage. Explicit coverage gaps, not demonstrated runtime incompatibility.
Concurrent aggregate reliability remains unresolved. #879 and draft #916; independently reproduced false process-exit behavior reported in dotnet/runtime#133736. Existing investigation; do not duplicate or overstate causal closure.
Upstream heap capture capabilities are evolving. dotnet/runtime#129457 introduces EventPipe blocking mode for .NET 11; dotnet/runtime#131678 fixes a NativeAOT ETW root-walk crash and silently missing stack roots. Version-gated research candidates, not demonstrated local defects or permission to enable them.

Execution index and readiness

The four phases were detailed independently and consolidated into 15 new
bounded subissues
, plus the existing aggregate investigation #879. All are
native GitHub subissues; completion dependencies are also registered natively.
Phase 1 is integrated; other lanes remain open as indicated below.

Serial execution checkpoint (2026-09-12)

Phase 1 was delivered one issue at a time, with separate independent review,
corrections, required CI approval and merge before starting the next child:
#925 -> #935, #926 -> #936, #927 -> #937, #928 -> #938.
The integrated baseline is 096eef8702e28ddc0cc85f20339ba5f9c578e072.

The known-graph report
preserves measured outcomes, failed development assumptions, protocol deviations
and remaining loss/root-fidelity limits. The four child outcomes do not imply
universal graph completeness or authorization to enable experimental collectors.
The unrelated Windows startup failures on the first #936 CI attempt remain
documented in that PR; a single explicit diagnostic rerun passed, without claiming
the underlying startup issue was resolved.

#920 is integrated through #939, including an isolated local Copilot CLI
transport that does not require extracting credentials or supplying a separate
harness API key. After preserved launch/protocol failures, uninstrumented smoke
07 passed on code revision 6e84c07cb847e1b097444acbfa8ebaf1adeae62a: two real
model turns, one model-chosen counters capture, a structured diagnosis with
resolving citations and uncertainty, and independently observed target cleanup
within the unchanged bounds. Independent review of the fixes found no significant
issues. All four required CI gates passed before merge, following one explicitly
documented Windows failed-job rerun. The original gcdump closed-pipe shutdown
failure remains unresolved and is tracked in #940. Resolving citations
did not establish semantic correctness; mismatched counter citations remain
evidence for #921's human calibration. #921 is in progress: #941 integrates the
review-packet/import infrastructure, and one actual primary human review of the
known development warmup is recorded. All three claims were judged supported
with appropriate certainty, with an explicit caveat that the case is evident/easy.
Independent human review, a frozen protocol, and broader development/heldout
evaluation remain pending; calibration is not complete;
#922-#924 remain explicitly deferred.
Phases 3-4 have not started in this batch. Structural independence of C/D lanes
does not override the requested serial execution policy.

Lane Deliverable Issue Prerequisite / readiness
A1 ThreadPool provenance and starvation semantics #925 Completed via #935
A2.1 ThreadPool quality pilot and minimal shared contract #926 Completed via #936
A2.2 gcdump quality propagation #927 Completed via #937
A2.3 Known-graph pressure validation #928 Completed with explicit fidelity limits via #938
B1.1 Blinded, bounded real-agent harness #920 Completed via #939; real smoke passed with explicit interpretation limits
B1.2 Human-calibrated advisory pilot #921 Infrastructure via #941 and one primary development review; independent/frozen empirical pilot pending
B1.3 Network-wait feasibility #922 Deferred: #921 plus explicit prioritization
B1.4 One database-wait mechanism #923 Deferred: #921 plus explicit prioritization
B1.5 Retained growth versus allocation churn #924 Deferred: #921 (including A2 outcomes) plus prioritization
C1 Linux culture-lookup reassessment #929 Independent; isolated current-baseline trials
C2a Live ClrMD and layout-sensitive version coverage #931 Independent; versioned targets and scoped ptrace
C2b Method-parameter version/lifecycle coverage #933 Independent; supported runtimes/payloads and approval
C2c Advisory perf sidecar automation #934 Independent; safely provisioned Linux topology
C3 Existing aggregate reliability investigation #879 / draft #916 Existing evidence/upstream triage; no duplicate
D1 Blocking-mode safety/fidelity research #930 Exact runtime/client API/protocol proof
D2 Fixed-build NativeAOT root-walk research #932 Exact embedded-runtime fix and fixture proof

Scheduling: Phase 1 and #920 are merged. #921 is the next
semantic/evaluation item. The primary warmup review does not replace independent
human review or a frozen, broader pilot. C1/C2 lanes are structurally independent, but the
current execution policy is serial. Coordinate shared fixture/workflow edits
rather than treating them as isolated files. Harness/protocol/fixture design may overlap prerequisites,
but native blockers describe acceptance readiness. D1/D2 are independent
research studies, not implementation authorization. Keep catalog expansion
explicitly deferred and model-quality/topology evaluation advisory.

Refined upstream availability: #930 records a source-verified .NET 11
blocking runtime candidate and diagnostics main API, separately from
unverified released-client binary availability. #932 identifies fixed
NativeAOT RC1 source/distribution metadata, while the inspected stable tags
remain unfixed. Neither is evidence of local safety or fidelity. Source and
version links, remaining triggers, and GO/NO-GO boundaries live in those
children; do not wait vaguely for a fix that already has a candidate build.

Phase 1 - Evidence integrity and interpretation boundaries

A1. Correct ThreadPool evidence semantics - first implementation priority

  • Create a focused implementation issue linked here before starting code changes: A1: Correct ThreadPool adjustment provenance and starvation attribution end to end #925, implemented in fix(threadpool): preserve evidence provenance and causal boundaries (#925) #935.
  • Keep observed, inferred and unknown adjustment reasons distinguishable. Worker growth alone must not become confirmed starvation; synthetic counts must not masquerade as observed measurements.
  • Trace the complete path: collector -> artifact -> query/summary -> comparison -> BenchmarkDotNet attribution -> MCP/CLI-facing guidance and playbooks.
  • Prefer additive provenance fields and explicit deprecation where compatibility requires it, but do not preserve incorrect causal meaning merely to preserve historical values.
  • Cover normal growth, warmup, explicit runtime starvation, unknown/missing payloads and synthetic-count paths.

Exit: every consumer can distinguish explicit starvation evidence from inference or missing information; documentation no longer promotes queue growth or arbitrary hill-climbing events to confirmation.

A2. Pilot a structured evidence-quality contract

Depends on A1's semantic decisions. Start with ThreadPool and gcdump, not a speculative rewrite of every collector.

Execution issues: #926 (ThreadPool/minimal contract) -> #927 (gcdump
propagation) -> #928 (bounded known-graph validation). A2 requires both
implementations and the empirical outcome; no universal completeness claim
is implied by closing them.

  • Inventory and reuse existing warnings, notes, truncation flags and loss metadata before adding new contracts.
  • Distinguish detected loss/truncation/timeout, inferred data, capture-window limitations, and uncertainty the collection mechanism cannot exclude.
  • Define which conclusions remain supportable when evidence is degraded, and which must become inconclusive.
  • Preserve limitations through summaries, drilldowns, exports and baseline comparisons across applicable product surfaces.
  • Exercise known-shape heaps with controlled capture pressure and a known expected graph. Distinguish collector retention caps from transport loss and target-side limitations.

Exit: degraded captures remain useful without claiming unsupported completeness. No detected loss is not automatically proof of a complete graph. Do not claim a universal completeness detector.

Phase 2 - End-to-end diagnostic evaluation

B1. Advisory agent-evaluation pilot

Build on A1/A2 and the existing deterministic scenario suite, #646, #681 and #736.

Execution issues: #920 (blinded, bounded harness) -> #921 (human-calibrated
pilot). Design can begin earlier; meaningful live evaluation requires usable
A1/A2 semantics. The following are explicitly deferred feasibility spikes,
not additional prerequisites for the initial four-shape pilot: #922
(network waiting), #923 (one database-wait mechanism), and #924
(retained growth versus allocation churn). Each requires the calibrated
pilot and explicit prioritization; only a GO decision opens its separately
scoped catalog-integration follow-up.

  • Retain deterministic evidence checks as the first layer.
  • Add a separate advisory layer in which an agent starts from a symptom, chooses tools and produces a diagnosis with evidence references and uncertainty.
  • Start with the four existing evidence shapes, plus healthy controls, competing explanations and deliberately insufficient/degraded captures. Do not expose scenario answers or source code to the diagnosing agent.
  • Calibrate against human-reviewed interpretations and held-out cases; the keyword mapper is an aid, not the correctness oracle.
  • Record model, prompt, product version, workload and capture provenance for repeatability.
  • Report supported attribution, citation correctness, unsupported certainty, correct abstention, next-step usefulness, approval compliance and investigation cost separately. A weighted average must not hide unsafe or unsupported conclusions.
  • Expand the catalog only for new evidence shapes, initially network/database waiting and retention versus allocation churn, each as a separately scoped follow-up.

Exit: the pilot can distinguish plausible but incorrect diagnoses from supported diagnoses and justified inconclusion. It remains advisory; PR gating requires a separate readiness decision backed by calibration evidence.

Phase 3 - Operational coverage and reliability

These items can proceed independently in parallel with Phase 1. Avoid a full collectors x runtimes x operating systems Cartesian matrix.

C1. Reassess Linux culture-lookup coverage

Execution issue: #929. Restoration and an evidence-backed continued
restriction are both valid bounded outcomes.

Exit: the remaining restriction is either safely removed or accurately justified. Do not infer success from removal of the separate ten Core quarantines.

C2. Extend mechanism-based compatibility coverage

Independent execution issues: #931 (live ClrMD/layout-sensitive
drilldowns), #933 (method parameters), #934 (perf sidecar). Their child
criteria specify representative cells, not a Cartesian matrix.

  • Prioritize live ClrMD and representative runtime-layout-sensitive drilldowns against .NET 8/9/10.
  • Add representative method-parameter coverage across supported target versions, keeping its explicit approval and instrumentation boundary.
  • Exercise perf-backed capture in one production-representative Linux sidecar topology with matching UID, PID/socket visibility and narrowly scoped capabilities.
  • Make unsupported environments and missing prerequisites explicit outcomes, not success-shaped no-ops. Start new topology jobs as advisory until their stability supports gating.
  • Keep support documentation tied to the precise mechanism, runtime, OS and topology actually exercised.

Exit: supported, partially covered and unavailable paths are clearly distinguished, with representative live evidence for the newly covered mechanisms.

C3. Continue the existing aggregate investigation

Track #879 and #916 rather than opening a duplicate bug.

#879 is reused as the native child. Its existing qualified evidence remains
authoritative; this decomposition does not change #916's draft status or
authorize a mitigation/merge. Include the original concurrent aggregate
topology in acceptance, and keep proxy-start timeout evidence separate.

  • Follow runtime maintainer triage of Linux: Process reports its ptrace-stopped child as exited when the parent is the tracer dotnet/runtime#133736.
  • Preserve the distinction between the independently reproduced false-exit defect and the still-qualified causal link to the historical aggregate hang; the independent unmodified ClrMD controls produced 0 hangs in 40 attaches.
  • Evaluate a local mitigation only if supported by evidence and consistent with process ownership and cleanup semantics.
  • Assess any correction against both the minimal reproduction and the original aggregate topology. Keep the independent proxy-start timeout separate.

Exit: closure is based on evidence for the actual failing topology, not extrapolation from clean sequential runs. This investigation does not block unrelated semantic and evaluation work.

Phase 4 - Version-gated upstream research

D1. EventPipe blocking mode and heap integrity

Research issue: #930. Source-pinned experiments and released-package
recommendations have distinct prerequisites.

  • Track [EventPipe] Add Blocking EventPipe Mode dotnet/runtime#129457 and the corresponding diagnostics-client release/API availability (CollectTracing6, EventPipeBufferingMode.Block).
  • When a suitable .NET 11 build is available, compare known-graph completeness under capture pressure with the current path.
  • Measure target pause/throughput impact, cancellation, reader failure, memory-pressure behavior and old-runtime negotiation.

Exit: a sourced GO/NO-GO decision and, only for GO, a separately scoped implementation issue. Blocking producers is not a default safety improvement and does not eliminate all loss modes.

D2. NativeAOT root-walk revalidation

Research issue: #932. A source-verified fixed RC1 candidate is available;
the actual published AOT executable's embedded runtime must still be proven
before isolated trials. No production guard is removed by this research.

  • Track the actual release/backport availability of Fix NativeAOT crash in Server GC with ETW GCHeapDump tracing dotnet/runtime#131678. At this assessment, the inspected stable 8/9/10 tags still contain the faulty condition.
  • On a build containing the fix, exercise Server GC and Workstation GC with known stack/static roots and expected graph structure.
  • Check graph/root completeness as well as absence of a crash; identify remaining type-name, static-root and transport limitations independently.

Exit: a version-specific GO/NO-GO decision. Keep the current NativeAOT gcdump prohibition unless the complete local path is demonstrated safe and its fidelity limits are explicit. This fix alone does not supply ClrMD/DAC parity or typed retention analysis.

Dependencies and sequencing

Work item Prerequisite Execution lane
A1 ThreadPool semantics None First small implementation PR
A2 Evidence-quality pilot A1 semantic decisions Core contract and applicable consumers
B1 Agent-evaluation pilot A1/A2 usable evidence semantics Advisory evaluation; harness design may start earlier
C1 Linux culture-lookup None; current dependency baseline Independent coverage reassessment
C2 Compatibility coverage Appropriate runtime/topology prerequisites Independent representative coverage
C3 Aggregate reliability Existing #879 evidence and upstream progress Existing investigation, not a global blocker
D1/D2 Upstream spikes Builds containing the relevant changes Research/watch; no main-runtime upgrade implied

Scope boundaries and tracking rules

  • No new MCP tools, no replacement of dotnet-monitor/APM, no mandatory prior target instrumentation.
  • Keep the CLI Core-only; reuse Core semantics without importing server dependencies.
  • Preserve intentional NativeAOT degradation rather than claiming universal parity.
  • MCP SDK migration is already shipped; do not reopen it without new evidence.
  • OTel Profiles remains Alpha/watch P5 (watch): OTel Profiles (pprof/OTLP) export for CPU/allocation samples #550, not a dependency of these phases.
  • Native mutex paired-probe research remains separately tracked in Spike paired uprobe/uretprobe latency for confirmed native mutex blocking #852.
  • Collector hotpath measurements are not proof of negligible impact on the target. Measure target impact where a new capture mode changes producer behavior.
  • Create small linked implementation/spike issues as work is picked up; link their PRs and durable evidence here. Do not bundle independent trails into one implementation PR.
  • Keep Phase 16 -- Roadmap (MCP protocol evolution + external capability gaps) #551 as the repository-wide roadmap entry point and link this follow-up from it.
  • A child issue closes when its bounded acceptance criteria are met, not when the entire capability is declared complete.
  • This roadmap closes when immediate deliverables have explicit outcomes and deferred work has versioned triggers plus linked follow-ups. External blockers must not disappear behind a checked umbrella item.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    diagnosticsAdvanced diagnostic featuresmetaTracking / coordination issueresearch-backedHas a documented research artifact

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions