Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions tests/bugbash/charters.bash-ai-proxy.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
# Charters for the bash AI proxy. Run with BUGBASH_AI=mock BUGBASH_SCENARIO=bash-ai-proxy and
# BUGBASH_ARGS="--charters tests/bugbash/charters.bash-ai-proxy.txt" (see e2e.config.ts).
# The scenario has no agent tools and no terminals: the fake model calls the proxy itself.
proxy-costs|web|skeptic|Turn on Settings > Providers > 'Count AI calls from bash commands', then in the 'Bug bash playground' workspace send 'probe', '[proxy:many]' and '[proxy:openai]'; after each, cross-check the documented probe token counts against the Stats > Cost tab Session rows (per-model rows, components, session total) and Analytics; check that Last Request stays the chat request; report every number that does not add up
proxy-switch|web|state|In Settings > Providers, find the 'Bash commands' switch and toggle it on and off, reloading after each change; after each state send 'probe' in the 'Bug bash playground' workspace and check that calls succeed (HTTP 200) only while the switch is on and are refused (HTTP 503) while it is off, that the switch state persists, and that the Cost tab never loses or double-counts earlier spend
proxy-live|web|skeptic|Turn on the bash AI proxy switch, open the Stats > Cost tab of the 'Bug bash playground' workspace, then send '[proxy:background]' and watch the Cost tab for about 20 seconds without reloading (it should grow with each call); then turn the switch off while the calls still run and watch for 30 more seconds; report totals that do not grow as calls happen, grow too much, change after a reload, or keep growing after the switch is off
proxy-inputs|web|fuzzer|Turn on the bash AI proxy switch, then in the 'Bug bash playground' workspace send '[proxy:stream]', '[proxy:bad-key]', '[proxy:openai]' and messages that mix several keywords, unicode and 300 characters; report wrong or missing usage, errors shown as success, a stuck chat, or a wrong key that still gets counted
proxy-workspaces|web|state|Turn on the bash AI proxy switch, create a second workspace in demo-app, and send 'probe' and '[proxy:many]' in each workspace; check that each workspace's Cost tab counts only its own proxied calls, that Analytics (which totals the whole demo-app project) shows the sum of both workspaces, and that archiving and restoring a workspace keeps its numbers
proxy-trust|web|state|Turn on the bash AI proxy switch and send 'probe' in the 'Bug bash playground' workspace (expect HTTP 200); then send '[proxy:background]' and, while it runs, revoke trust for demo-app in Settings > Security; send '[proxy:status]' until it says finished: calls after the revoke must show HTTP 403 and must not be counted in the Stats > Cost tab; only then trust the project again and send 'probe' (HTTP 200, counted once); report anything that leaks, sticks or double-counts
proxy-phone|phone|newcomer|On this phone-sized screen, open Settings > Providers and find the 'Bash commands' section; read its title and description and report text that is confusing, wrong, cut off or overlapping; turn it on, send 'probe' in the 'Bug bash playground' workspace and check the reply and Analytics (the Cost tab is known to be unreachable on phones, do not try it)
46 changes: 45 additions & 1 deletion tests/bugbash/e2e.config.ts
Original file line number Diff line number Diff line change
Expand Up @@ -86,8 +86,49 @@ const mockNonBugs = realAi
"the mock echoes your text, and after a retry it may echo [CONTINUE].",
];

// BUGBASH_SCENARIO=bash-ai-proxy swaps mock AI for a loopback fake provider (startApp.ts,
// fakeProvider.ts) so explorers can drive the bash AI proxy end to end.
const scenario = process.env.BUGBASH_SCENARIO ?? "";

// What the explorer must know about the bash AI proxy scenario instead of the mock-AI notes.
const bashAiProxyContext = [
"The app is Xum, a desktop and browser app for running parallel AI coding agents.",
"It starts with one project, demo-app, and one workspace, 'Bug bash playground', in the left sidebar.",
"The feature under test is the bash AI proxy. When Settings > Providers > 'Bash commands' >",
"'Count AI calls from bash commands' is on (it is off by default), bash commands get a",
"per-workspace 'xum-proxy-...' key and endpoints that point at Xum. Xum forwards those calls with",
"the provider key from its settings and adds their tokens and cost to that workspace's",
"right-sidebar Stats > Cost tab (Session rows per model and total) and to Analytics (the bar-chart",
"button). The Cost tab's Last Request stays the chat request: proxied calls never replace it.",
"Turning the switch off makes the proxy refuse calls (HTTP 503); revoking the project's trust",
"makes it refuse that workspace's calls (HTTP 403).",
"AI is a local fake provider, nothing is billed, and agent tools are off, so the agent runs no",
"command. Instead the fake model itself calls the proxy with the key Xum gives that workspace's",
"bash commands, and replies 'Results of ...' with one HTTP result per call (an expected refusal is a result, not a failure of the reply). A keyword in your message",
"picks the calls: none = one Anthropic call; '[proxy:stream]' = one streamed Anthropic call;",
"'[proxy:openai]' = one OpenAI chat and one OpenAI responses call; '[proxy:many]' = five Anthropic",
"calls; '[proxy:background]' = twelve Anthropic calls, one every 5 seconds, after the reply (about",
"a minute); '[proxy:bad-key]' = a call with a wrong key (expect HTTP 401); '[proxy:status]' = no call, only the background results so far (it says 'finished' when all ran). Every reply also lists the background results. One plan per message:",
"when a message names several keywords, only the first one in this list runs.",
"Each Anthropic probe reports input_tokens 1234 plus cache_read_input_tokens 100 (Anthropic counts",
"cache reads separately, so 1334 input in total) and 56 output tokens, on claude-opus-5-5. Each OpenAI",
"probe reports 1234 prompt tokens of which 100 are cached (1134 uncached) and 56 output tokens, on",
"gpt-6.1-sol. The chat turns themselves cost a little on the workspace model (claude-sonnet-5-5 by",
"default), so they get their own Cost tab row. Analytics totals the whole project, not one workspace.",
"Terminals are turned off for this session, so an error when opening one is expected.",
"Do not sign in to any provider, MCP server or external service, and do not enter real secrets.",
"Known and not bugs: the per-model cost column rounds amounts under $0.01 to ~$0.00; this test",
"browser denies clipboard writes; workspace names must be lowercase branch names.",
"Known and tracked, do not report again: at phone width the right sidebar (Stats > Cost) cannot be",
"opened (#5767); Analytics response counts leave out bash proxy rows and label them agent 'unknown'",
"(#5766); Analytics shows timestamps as '1.8T', repeats y-axis ticks, says '1 responses' and",
"overlaps at 390px (#5768); at 390px the workspace footer overflows and a long bash Script block is",
"cut off (#5769); the Archived Workspaces list shows a stale cost until reload (#5786). By design:",
"the 'Workspace created' row follows the first message.",
].join(" ");

// What the local app cannot do, and the explorer's own blind spots (see `e2e guide bug-bash`).
const context = [
const mockOrRealContext = [
"The app is Xum, a desktop and browser app for running parallel AI coding agents.",
"It starts with one project, demo-app, and one workspace, 'Bug bash playground', in the left sidebar.",
...aiContext,
Expand Down Expand Up @@ -123,6 +164,8 @@ const context = [
"after about one second; Ctrl+/ cycles to the next model; Fast mode is unavailable in this setup.",
].join(" ");

const context = scenario === "bash-ai-proxy" ? bashAiProxyContext : mockOrRealContext;

// Exploration steps need large per-step budgets (`e2e guide bug-bash`, step 1).
// Only the active provider reads its key, so both can be set at once.
const effort = explorerEffort();
Expand All @@ -141,6 +184,7 @@ const APP_ENV_VARS = [
"BUGBASH_AI_RESOLVED",
"BUGBASH_AI_REASON",
"BUGBASH_APP_MODEL",
"BUGBASH_SCENARIO",
"ANTHROPIC_API_KEY",
"ANTHROPIC_BASE_URL",
"OPENAI_API_KEY",
Expand Down
Loading
Loading