You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Re-run the same five Pi diagrams from one pinned revision.
Compare repair rounds, aggregate agent work, artifact-ready time, and handoff-ready time.
Confirm node/relationship/label/source-evidence counts do not regress.
Confirm final visual receipts and manual review pass unchanged.
Benchmark target
Treat this as an A/B target, not a promise: reduce aggregate work from ~44 minutes to approximately 31-35 minutes on the same five Pi inputs, with no quality-gate reductions.
Problem
The five-diagram Pi benchmark spent 2,637.592s (43m57.592s) of aggregate agent work even though the renderer/checker commands themselves are fast.
Measured evidence:
The controllable cost is diagnosis, late viewport feedback, repeated orchestration, and report transcription—not renderer CPU.
Goal
Reduce the same five-diagram aggregate benchmark by 20-30% without weakening semantic, geometric, visual, evidence, or delivery quality.
Scope
P0 — actionable diagnostics
P0 — viewport feedback before delivery
P0 — orchestration and measurement
P1 — safe reuse
authoring-kit <type>returning the exact type schema, common schema, and one matching example with hashes.P2 — only after profiling
Non-negotiable quality gates
Verification
Benchmark target
Treat this as an A/B target, not a promise: reduce aggregate work from ~44 minutes to approximately 31-35 minutes on the same five Pi inputs, with no quality-gate reductions.