Repository navigation
🤖 tests: bug-bash scenario for the bash AI proxy - #5775
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: e520c795ae
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
…l calls the proxy itself
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: d3a6a857c7
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 31e3a046c9
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 84885d749b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 074dfd8f29
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Summary
make bug-bashgets aBUGBASH_SCENARIO=bash-ai-proxyscenario. A loopback fake provider is the app's chat model and the proxy's upstream. On each agent turn the fake model itself calls Xum's proxy, using the key Xum derives for that workspace's bash commands. A keyword in the message picks the calls, and the reply lists the HTTP results. Explorers can test the proxy, the Cost tab and Analytics end to end with no real model spend and with agent tools off.PR 4 of 4. Test harness only.
Implementation
tests/bugbash/fakeProvider.ts: fixed call plans (one call, streamed, OpenAI chat and responses, five calls, twelve background calls, wrong key). The explorer picks only the plan. Probe answers carry fixed token counts, so explorers can check the Cost tab against them. The fake finds the workspace from the worktree path in the system prompt.startApp.tsstarts the fake and points both providers at it. It needsBUGBASH_AI=mockand setsXUM_DISABLE_AGENT_TOOLS=1andXUM_DISABLE_TERMINALS=1(AGENTS.md: no agent tools or terminals until the app runs in a sandbox). The bash env injection itself is covered by unit tests in 🤖 feat: count AI calls from bash commands in workspace costs (opt-in) #5772.e2e.config.tsgives explorers a scenario context that lists the known, tracked findings.charters.bash-ai-proxy.txtholds the charters.Run:
BUGBASH_AI=mock BUGBASH_SCENARIO=bash-ai-proxy BUGBASH_EFFORT=high \ make bug-bash BUGBASH_ARGS="--charters tests/bugbash/charters.bash-ai-proxy.txt"Validation
Four bug-bash rounds ran with an earlier version of this scenario, in which the fake model asked the agent to run fixed bash commands. Each round used Opus 5.5 and Sonnet 5.5 explorers at effort high. The first round covered six charters, about 16.6M explorer tokens in total.
Generated with
xum• Model:anthropic:claude-opus-5-5• Thinking:high• Cost:$32.63