You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up to #551, based on the mission/completeness reassessment of 2026-09-11.
The mission is on-demand performance diagnosis of running .NET applications, driven by an LLM through MCP or by humans/automation through the Core-only CLI and BenchmarkDotNet diagnoser. Standard diagnostics require no target code changes or prior instrumentation; method-parameter capture remains an explicit, opt-in instrumentation exception.
The next priority is diagnostic trust, not a larger tool catalog: preserve the distinction between observed facts, inferred explanations, and unavailable evidence from collection through the final diagnosis.
This issue formalizes the agreed roadmap. It does not authorize enabling experimental collectors, changing runtime support, removing protections without evidence, or introducing new MCP tools. The delivery phases below are internal to this roadmap; they do not renumber the repository's existing Phase 16.
Evidence and current limits
Finding
Evidence
Classification
Unknown ThreadPool adjustment reasons can become Starvation, including through synthetic worker counts; downstream summaries count them as starvation adjustments.
Scenario evidence evaluation is not end-to-end evaluation of the agent.
Four scenario manifests; live tests call ScenarioEvaluator.CreateReport without an interpretation, resulting in Interpretation.NotRun. #736 delivered an explicitly advisory keyword-based mapper.
Known evaluation scope gap, not absence of all evaluation.
A Linux scenario quarantine remains after the ten Core quarantines were removed.
ScenarioLiveTests.cs: culture-lookup remains WindowsOnlyFact citing #147; the isolated workflow omits Linux culture-lookup.
Concrete remaining coverage restriction; requires its own reassessment.
Compatibility claims are broader than the mechanisms continuously exercised.
Version-gated research candidates, not demonstrated local defects or permission to enable them.
Execution index and readiness
The four phases were detailed independently and consolidated into 15 new
bounded subissues, plus the existing aggregate investigation #879. All are
native GitHub subissues; completion dependencies are also registered natively.
Phase 1 is integrated; other lanes remain open as indicated below.
Serial execution checkpoint (2026-09-12)
Phase 1 was delivered one issue at a time, with separate independent review,
corrections, required CI approval and merge before starting the next child: #925 -> #935, #926 -> #936, #927 -> #937, #928 -> #938.
The integrated baseline is 096eef8702e28ddc0cc85f20339ba5f9c578e072.
The known-graph report
preserves measured outcomes, failed development assumptions, protocol deviations
and remaining loss/root-fidelity limits. The four child outcomes do not imply
universal graph completeness or authorization to enable experimental collectors.
The unrelated Windows startup failures on the first #936 CI attempt remain
documented in that PR; a single explicit diagnostic rerun passed, without claiming
the underlying startup issue was resolved.
#920 is integrated through #939, including an isolated local Copilot CLI
transport that does not require extracting credentials or supplying a separate
harness API key. After preserved launch/protocol failures, uninstrumented smoke
07 passed on code revision 6e84c07cb847e1b097444acbfa8ebaf1adeae62a: two real
model turns, one model-chosen counters capture, a structured diagnosis with
resolving citations and uncertainty, and independently observed target cleanup
within the unchanged bounds. Independent review of the fixes found no significant
issues. All four required CI gates passed before merge, following one explicitly
documented Windows failed-job rerun. The original gcdump closed-pipe shutdown
failure remains unresolved and is tracked in #940. Resolving citations
did not establish semantic correctness; mismatched counter citations remain
evidence for #921's human calibration. #921 is in progress: #941 integrates the
review-packet/import infrastructure, and one actual primary human review of the
known development warmup is recorded. All three claims were judged supported
with appropriate certainty, with an explicit caveat that the case is evident/easy.
Independent human review, a frozen protocol, and broader development/heldout
evaluation remain pending; calibration is not complete; #922-#924 remain explicitly deferred.
Phases 3-4 have not started in this batch. Structural independence of C/D lanes
does not override the requested serial execution policy.
Scheduling: Phase 1 and #920 are merged. #921 is the next
semantic/evaluation item. The primary warmup review does not replace independent
human review or a frozen, broader pilot. C1/C2 lanes are structurally independent, but the
current execution policy is serial. Coordinate shared fixture/workflow edits
rather than treating them as isolated files. Harness/protocol/fixture design may overlap prerequisites,
but native blockers describe acceptance readiness. D1/D2 are independent
research studies, not implementation authorization. Keep catalog expansion
explicitly deferred and model-quality/topology evaluation advisory.
Refined upstream availability:#930 records a source-verified .NET 11
blocking runtime candidate and diagnostics main API, separately from
unverified released-client binary availability. #932 identifies fixed
NativeAOT RC1 source/distribution metadata, while the inspected stable tags
remain unfixed. Neither is evidence of local safety or fidelity. Source and
version links, remaining triggers, and GO/NO-GO boundaries live in those
children; do not wait vaguely for a fix that already has a candidate build.
Phase 1 - Evidence integrity and interpretation boundaries
A1. Correct ThreadPool evidence semantics - first implementation priority
Keep observed, inferred and unknown adjustment reasons distinguishable. Worker growth alone must not become confirmed starvation; synthetic counts must not masquerade as observed measurements.
Trace the complete path: collector -> artifact -> query/summary -> comparison -> BenchmarkDotNet attribution -> MCP/CLI-facing guidance and playbooks.
Prefer additive provenance fields and explicit deprecation where compatibility requires it, but do not preserve incorrect causal meaning merely to preserve historical values.
Cover normal growth, warmup, explicit runtime starvation, unknown/missing payloads and synthetic-count paths.
Exit: every consumer can distinguish explicit starvation evidence from inference or missing information; documentation no longer promotes queue growth or arbitrary hill-climbing events to confirmation.
A2. Pilot a structured evidence-quality contract
Depends on A1's semantic decisions. Start with ThreadPool and gcdump, not a speculative rewrite of every collector.
Execution issues:#926 (ThreadPool/minimal contract) -> #927 (gcdump
propagation) -> #928 (bounded known-graph validation). A2 requires both
implementations and the empirical outcome; no universal completeness claim
is implied by closing them.
Inventory and reuse existing warnings, notes, truncation flags and loss metadata before adding new contracts.
Distinguish detected loss/truncation/timeout, inferred data, capture-window limitations, and uncertainty the collection mechanism cannot exclude.
Define which conclusions remain supportable when evidence is degraded, and which must become inconclusive.
Preserve limitations through summaries, drilldowns, exports and baseline comparisons across applicable product surfaces.
Exercise known-shape heaps with controlled capture pressure and a known expected graph. Distinguish collector retention caps from transport loss and target-side limitations.
Exit: degraded captures remain useful without claiming unsupported completeness. No detected loss is not automatically proof of a complete graph. Do not claim a universal completeness detector.
Phase 2 - End-to-end diagnostic evaluation
B1. Advisory agent-evaluation pilot
Build on A1/A2 and the existing deterministic scenario suite, #646, #681 and #736.
Execution issues:#920 (blinded, bounded harness) -> #921 (human-calibrated
pilot). Design can begin earlier; meaningful live evaluation requires usable
A1/A2 semantics. The following are explicitly deferred feasibility spikes,
not additional prerequisites for the initial four-shape pilot: #922
(network waiting), #923 (one database-wait mechanism), and #924
(retained growth versus allocation churn). Each requires the calibrated
pilot and explicit prioritization; only a GO decision opens its separately
scoped catalog-integration follow-up.
Retain deterministic evidence checks as the first layer.
Add a separate advisory layer in which an agent starts from a symptom, chooses tools and produces a diagnosis with evidence references and uncertainty.
Start with the four existing evidence shapes, plus healthy controls, competing explanations and deliberately insufficient/degraded captures. Do not expose scenario answers or source code to the diagnosing agent.
Calibrate against human-reviewed interpretations and held-out cases; the keyword mapper is an aid, not the correctness oracle.
Record model, prompt, product version, workload and capture provenance for repeatability.
Report supported attribution, citation correctness, unsupported certainty, correct abstention, next-step usefulness, approval compliance and investigation cost separately. A weighted average must not hide unsafe or unsupported conclusions.
Expand the catalog only for new evidence shapes, initially network/database waiting and retention versus allocation churn, each as a separately scoped follow-up.
Exit: the pilot can distinguish plausible but incorrect diagnoses from supported diagnoses and justified inconclusion. It remains advisory; PR gating requires a separate readiness decision backed by calibration evidence.
Phase 3 - Operational coverage and reliability
These items can proceed independently in parallel with Phase 1. Avoid a full collectors x runtimes x operating systems Cartesian matrix.
C1. Reassess Linux culture-lookup coverage
Execution issue:#929. Restoration and an evidence-backed continued
restriction are both valid bounded outcomes.
Run bounded, isolated trials with the current dependency versions and preserve failed/aborted evidence.
Restore Linux scenario coverage only if supported by the results.
Exit: the remaining restriction is either safely removed or accurately justified. Do not infer success from removal of the separate ten Core quarantines.
C2. Extend mechanism-based compatibility coverage
Independent execution issues:#931 (live ClrMD/layout-sensitive
drilldowns), #933 (method parameters), #934 (perf sidecar). Their child
criteria specify representative cells, not a Cartesian matrix.
Prioritize live ClrMD and representative runtime-layout-sensitive drilldowns against .NET 8/9/10.
Add representative method-parameter coverage across supported target versions, keeping its explicit approval and instrumentation boundary.
Exercise perf-backed capture in one production-representative Linux sidecar topology with matching UID, PID/socket visibility and narrowly scoped capabilities.
Make unsupported environments and missing prerequisites explicit outcomes, not success-shaped no-ops. Start new topology jobs as advisory until their stability supports gating.
Keep support documentation tied to the precise mechanism, runtime, OS and topology actually exercised.
Exit: supported, partially covered and unavailable paths are clearly distinguished, with representative live evidence for the newly covered mechanisms.
C3. Continue the existing aggregate investigation
Track #879 and #916 rather than opening a duplicate bug.
#879 is reused as the native child. Its existing qualified evidence remains
authoritative; this decomposition does not change #916's draft status or
authorize a mitigation/merge. Include the original concurrent aggregate
topology in acceptance, and keep proxy-start timeout evidence separate.
Preserve the distinction between the independently reproduced false-exit defect and the still-qualified causal link to the historical aggregate hang; the independent unmodified ClrMD controls produced 0 hangs in 40 attaches.
Evaluate a local mitigation only if supported by evidence and consistent with process ownership and cleanup semantics.
Assess any correction against both the minimal reproduction and the original aggregate topology. Keep the independent proxy-start timeout separate.
Exit: closure is based on evidence for the actual failing topology, not extrapolation from clean sequential runs. This investigation does not block unrelated semantic and evaluation work.
Phase 4 - Version-gated upstream research
D1. EventPipe blocking mode and heap integrity
Research issue:#930. Source-pinned experiments and released-package
recommendations have distinct prerequisites.
Exit: a sourced GO/NO-GO decision and, only for GO, a separately scoped implementation issue. Blocking producers is not a default safety improvement and does not eliminate all loss modes.
D2. NativeAOT root-walk revalidation
Research issue:#932. A source-verified fixed RC1 candidate is available;
the actual published AOT executable's embedded runtime must still be proven
before isolated trials. No production guard is removed by this research.
On a build containing the fix, exercise Server GC and Workstation GC with known stack/static roots and expected graph structure.
Check graph/root completeness as well as absence of a crash; identify remaining type-name, static-root and transport limitations independently.
Exit: a version-specific GO/NO-GO decision. Keep the current NativeAOT gcdump prohibition unless the complete local path is demonstrated safe and its fidelity limits are explicit. This fix alone does not supply ClrMD/DAC parity or typed retention analysis.
Dependencies and sequencing
Work item
Prerequisite
Execution lane
A1 ThreadPool semantics
None
First small implementation PR
A2 Evidence-quality pilot
A1 semantic decisions
Core contract and applicable consumers
B1 Agent-evaluation pilot
A1/A2 usable evidence semantics
Advisory evaluation; harness design may start earlier
Collector hotpath measurements are not proof of negligible impact on the target. Measure target impact where a new capture mode changes producer behavior.
Create small linked implementation/spike issues as work is picked up; link their PRs and durable evidence here. Do not bundle independent trails into one implementation PR.
A child issue closes when its bounded acceptance criteria are met, not when the entire capability is declared complete.
This roadmap closes when immediate deliverables have explicit outcomes and deferred work has versioned triggers plus linked follow-ups. External blockers must not disappear behind a checked umbrella item.
Mission and outcome
Follow-up to #551, based on the mission/completeness reassessment of 2026-09-11.
The mission is on-demand performance diagnosis of running .NET applications, driven by an LLM through MCP or by humans/automation through the Core-only CLI and BenchmarkDotNet diagnoser. Standard diagnostics require no target code changes or prior instrumentation; method-parameter capture remains an explicit, opt-in instrumentation exception.
The next priority is diagnostic trust, not a larger tool catalog: preserve the distinction between observed facts, inferred explanations, and unavailable evidence from collection through the final diagnosis.
This issue formalizes the agreed roadmap. It does not authorize enabling experimental collectors, changing runtime support, removing protections without evidence, or introducing new MCP tools. The delivery phases below are internal to this roadmap; they do not renumber the repository's existing Phase 16.
Evidence and current limits
Starvation, including through synthetic worker counts; downstream summaries count them as starvation adjustments.EventPipeThreadPoolCollector.NormalizeHillClimbing,CollectionQueryDispatcher,ThreadPoolComparableProjector. The collector emits an inference note, but the reason is still promoted.investigation-playbooks.md, sections 1 and 1a, versus the bounded-hypothesis contract in the README.ScenarioEvaluator.CreateReportwithout an interpretation, resulting inInterpretation.NotRun. #736 delivered an explicitly advisory keyword-based mapper.ScenarioLiveTests.cs:culture-lookupremainsWindowsOnlyFactciting #147; the isolated workflow omits Linux culture-lookup.runtime-version-compat-matrix.mdlists live ClrMD, deeper drilldowns and method-parameter cross-version gaps.perf-compat-matrix.mddocuments manual rather than automated perf sidecar coverage.Execution index and readiness
The four phases were detailed independently and consolidated into 15 new
bounded subissues, plus the existing aggregate investigation #879. All are
native GitHub subissues; completion dependencies are also registered natively.
Phase 1 is integrated; other lanes remain open as indicated below.
Serial execution checkpoint (2026-09-12)
Phase 1 was delivered one issue at a time, with separate independent review,
corrections, required CI approval and merge before starting the next child:
#925 -> #935, #926 -> #936, #927 -> #937, #928 -> #938.
The integrated baseline is
096eef8702e28ddc0cc85f20339ba5f9c578e072.The known-graph report
preserves measured outcomes, failed development assumptions, protocol deviations
and remaining loss/root-fidelity limits. The four child outcomes do not imply
universal graph completeness or authorization to enable experimental collectors.
The unrelated Windows startup failures on the first #936 CI attempt remain
documented in that PR; a single explicit diagnostic rerun passed, without claiming
the underlying startup issue was resolved.
#920 is integrated through #939, including an isolated local Copilot CLI
transport that does not require extracting credentials or supplying a separate
harness API key. After preserved launch/protocol failures, uninstrumented smoke
07 passed on code revision
6e84c07cb847e1b097444acbfa8ebaf1adeae62a: two realmodel turns, one model-chosen counters capture, a structured diagnosis with
resolving citations and uncertainty, and independently observed target cleanup
within the unchanged bounds. Independent review of the fixes found no significant
issues. All four required CI gates passed before merge, following one explicitly
documented Windows failed-job rerun. The original gcdump closed-pipe shutdown
failure remains unresolved and is tracked in #940. Resolving citations
did not establish semantic correctness; mismatched counter citations remain
evidence for #921's human calibration. #921 is in progress: #941 integrates the
review-packet/import infrastructure, and one actual primary human review of the
known development warmup is recorded. All three claims were judged supported
with appropriate certainty, with an explicit caveat that the case is evident/easy.
Independent human review, a frozen protocol, and broader development/heldout
evaluation remain pending; calibration is not complete;
#922-#924 remain explicitly deferred.
Phases 3-4 have not started in this batch. Structural independence of C/D lanes
does not override the requested serial execution policy.
Scheduling: Phase 1 and #920 are merged. #921 is the next
semantic/evaluation item. The primary warmup review does not replace independent
human review or a frozen, broader pilot. C1/C2 lanes are structurally independent, but the
current execution policy is serial. Coordinate shared fixture/workflow edits
rather than treating them as isolated files. Harness/protocol/fixture design may overlap prerequisites,
but native blockers describe acceptance readiness. D1/D2 are independent
research studies, not implementation authorization. Keep catalog expansion
explicitly deferred and model-quality/topology evaluation advisory.
Refined upstream availability: #930 records a source-verified .NET 11
blocking runtime candidate and diagnostics main API, separately from
unverified released-client binary availability. #932 identifies fixed
NativeAOT RC1 source/distribution metadata, while the inspected stable tags
remain unfixed. Neither is evidence of local safety or fidelity. Source and
version links, remaining triggers, and GO/NO-GO boundaries live in those
children; do not wait vaguely for a fix that already has a candidate build.
Phase 1 - Evidence integrity and interpretation boundaries
A1. Correct ThreadPool evidence semantics - first implementation priority
Exit: every consumer can distinguish explicit starvation evidence from inference or missing information; documentation no longer promotes queue growth or arbitrary hill-climbing events to confirmation.
A2. Pilot a structured evidence-quality contract
Depends on A1's semantic decisions. Start with ThreadPool and gcdump, not a speculative rewrite of every collector.
Execution issues: #926 (ThreadPool/minimal contract) -> #927 (gcdump
propagation) -> #928 (bounded known-graph validation). A2 requires both
implementations and the empirical outcome; no universal completeness claim
is implied by closing them.
Exit: degraded captures remain useful without claiming unsupported completeness. No detected loss is not automatically proof of a complete graph. Do not claim a universal completeness detector.
Phase 2 - End-to-end diagnostic evaluation
B1. Advisory agent-evaluation pilot
Build on A1/A2 and the existing deterministic scenario suite, #646, #681 and #736.
Execution issues: #920 (blinded, bounded harness) -> #921 (human-calibrated
pilot). Design can begin earlier; meaningful live evaluation requires usable
A1/A2 semantics. The following are explicitly deferred feasibility spikes,
not additional prerequisites for the initial four-shape pilot: #922
(network waiting), #923 (one database-wait mechanism), and #924
(retained growth versus allocation churn). Each requires the calibrated
pilot and explicit prioritization; only a GO decision opens its separately
scoped catalog-integration follow-up.
Exit: the pilot can distinguish plausible but incorrect diagnoses from supported diagnoses and justified inconclusion. It remains advisory; PR gating requires a separate readiness decision backed by calibration evidence.
Phase 3 - Operational coverage and reliability
These items can proceed independently in parallel with Phase 1. Avoid a full collectors x runtimes x operating systems Cartesian matrix.
C1. Reassess Linux culture-lookup coverage
Execution issue: #929. Restoration and an evidence-backed continued
restriction are both valid bounded outcomes.
Exit: the remaining restriction is either safely removed or accurately justified. Do not infer success from removal of the separate ten Core quarantines.
C2. Extend mechanism-based compatibility coverage
Independent execution issues: #931 (live ClrMD/layout-sensitive
drilldowns), #933 (method parameters), #934 (perf sidecar). Their child
criteria specify representative cells, not a Cartesian matrix.
Exit: supported, partially covered and unavailable paths are clearly distinguished, with representative live evidence for the newly covered mechanisms.
C3. Continue the existing aggregate investigation
Track #879 and #916 rather than opening a duplicate bug.
#879 is reused as the native child. Its existing qualified evidence remains
authoritative; this decomposition does not change #916's draft status or
authorize a mitigation/merge. Include the original concurrent aggregate
topology in acceptance, and keep proxy-start timeout evidence separate.
Exit: closure is based on evidence for the actual failing topology, not extrapolation from clean sequential runs. This investigation does not block unrelated semantic and evaluation work.
Phase 4 - Version-gated upstream research
D1. EventPipe blocking mode and heap integrity
Research issue: #930. Source-pinned experiments and released-package
recommendations have distinct prerequisites.
CollectTracing6,EventPipeBufferingMode.Block).Exit: a sourced GO/NO-GO decision and, only for GO, a separately scoped implementation issue. Blocking producers is not a default safety improvement and does not eliminate all loss modes.
D2. NativeAOT root-walk revalidation
Research issue: #932. A source-verified fixed RC1 candidate is available;
the actual published AOT executable's embedded runtime must still be proven
before isolated trials. No production guard is removed by this research.
Exit: a version-specific GO/NO-GO decision. Keep the current NativeAOT gcdump prohibition unless the complete local path is demonstrated safe and its fidelity limits are explicit. This fix alone does not supply ClrMD/DAC parity or typed retention analysis.
Dependencies and sequencing
Scope boundaries and tracking rules