Skip to content

Fix agent leaderboard model comparison - #109

Merged
Julia-Lex merged 1 commit into
mainfrom
codex/fix-agent-leaderboard-comparison
Jul 19, 2026
Merged

Julia-Lex merged 1 commit into
mainfrom
codex/fix-agent-leaderboard-comparison

Conversation

@Julia-Lex

Copy link
Copy Markdown
Contributor

Summary

  • rename the Agent Leaderboard model column from Model / runtime to Model
  • normalize gpt-5.5 model labels by removing the medium runtime suffix
  • add the comparable browser-use + gpt-5.5 row from the model leaderboard into the agent leaderboard

Tests

  • exact leaderboard table parser check
  • git diff --check

@Julia-Lex
Julia-Lex merged commit 05a8aed into main Jul 19, 2026
1 of 3 checks passed
@Julia-Lex
Julia-Lex deleted the codex/fix-agent-leaderboard-comparison branch July 19, 2026 00:42

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0bc546d738

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread landing/index.html
<td class="num">254.4</td>
</tr>
<tr data-rank="3" data-agent="Claude Code" data-model="gpt-5.5 medium" data-browser="Lexmount" data-pass="156" data-fail="54" data-total="210" data-success="74.29" data-steps="" data-e2e="228.4">
<tr data-rank="3" data-agent="browser-use" data-model="gpt-5.5" data-browser="Lexmount" data-pass="156" data-fail="54" data-total="210" data-success="74.3" data-steps="14.00" data-e2e="212.0">

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep results from different judges in separate leaderboards

Do not add this browser-use result to the Agent Leaderboard as a directly comparable row: it was copied from the Model Leaderboard result originally published as a 2026-06-16 snapshot evaluated with the gpt-5.4 judge (see commit 5b94115), whereas this section explicitly states that its main-agent results use the gpt-4.1 stepwise judge. Because changing the judge can change which tasks pass, ranking the 156/210 result alongside the gpt-4.1-judged agent runs produces an invalid agent comparison; re-evaluate this run with the same judge or disclose and separate it.

Useful? React with 👍 / 👎.

Comment thread landing/index.html
<td class="num">254.4</td>
</tr>
<tr data-rank="3" data-agent="Claude Code" data-model="gpt-5.5 medium" data-browser="Lexmount" data-pass="156" data-fail="54" data-total="210" data-success="74.29" data-steps="" data-e2e="228.4">
<tr data-rank="3" data-agent="browser-use" data-model="gpt-5.5" data-browser="Lexmount" data-pass="156" data-fail="54" data-total="210" data-success="74.3" data-steps="14.00" data-e2e="212.0">

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Update the advertised agent count

Adding this sixth agent leaves the leaderboard introduction saying that the agent view compares five agents, even though the table now contains Hermes, Codex, browser-use, Claude Code, Cursor, and OpenClaw. Update the summary count so the public page accurately describes the table.

Useful? React with 👍 / 👎.

Comment thread landing/index.html
Comment on lines +557 to +560
<tr data-rank="1" data-agent="Hermes" data-model="gpt-5.5" data-browser="Lexmount" data-pass="174" data-fail="36" data-total="210" data-success="82.86" data-steps="" data-e2e="184.4">
<td class="num">1</td>
<td><span class="cell-strong">Hermes</span></td>
<td><code class="cell-mono">gpt-5.5 medium</code></td>
<td><code class="cell-mono">gpt-5.5</code></td>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Preserve the medium runtime qualifier

Keep the medium qualifier in the agent rows rather than relabeling every configuration as bare gpt-5.5. The previous Model / runtime values recorded a material inference/runtime setting (and Cursor's reported model variant), while the newly imported browser-use result only had the bare model label; erasing that distinction makes differently configured runs appear identical and removes information needed to interpret or reproduce their performance.

Useful? React with 👍 / 👎.

Comment thread landing/index.html
<td class="num">254.4</td>
</tr>
<tr data-rank="3" data-agent="Claude Code" data-model="gpt-5.5 medium" data-browser="Lexmount" data-pass="156" data-fail="54" data-total="210" data-success="74.29" data-steps="" data-e2e="228.4">
<tr data-rank="3" data-agent="browser-use" data-model="gpt-5.5" data-browser="Lexmount" data-pass="156" data-fail="54" data-total="210" data-success="74.3" data-steps="14.00" data-e2e="212.0">

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use consistent precision for success-rate sorting

Store this success rate at the same precision as the other agent rows. This row and Claude Code both have exactly 156 passes out of 210, but their data-success values are 74.3 and 74.29; when a user clicks the Success % header, landing/assets/script.js compares those numeric attributes and incorrectly treats the equal rates as different. Normalize the underlying values and use an explicit tie-breaker if these agents must have distinct ranks.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant