Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions evals/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,10 @@ Passing requires every functional browser/API, TypeScript, data-preservation/res

See [the full protocol](PROTOCOL.md), [task definitions](yoloeval/catalog.py), [grading code](browser/checks.mjs), and [acceptance/early-stop rules](yoloeval/unseen.py). Jev routes the candidate; deterministic checks grade the results. There is no LLM judge in this protocol.

## Recorded evidence

The [recorded result](results/unseen-002/README.md) includes all 120 outcomes, 393 model-call records, source diffs, provenance, and an offline audit. Run `python3 evals/audit_results.py` to recalculate its acceptance and timings without API calls.

## Interpretation

The recorded unseen result was **45/60 → 53/60**, mean **70.58s → 47.37s**, P50 **54.30s → 30.92s**. This is a complete-bundle result, not an isolated Jev ablation. Campaign 001 was invalidated for a native-search-input grader bug; campaign 002 restarted all 120 attempts after correcting it. A provider outage was held between attempts; all scored service errors remained in the final data.
Expand Down
70 changes: 70 additions & 0 deletions evals/audit_results.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
#!/usr/bin/env python3
"""Audit the published 120-attempt historical result without API calls."""
import hashlib
import json
import math
import statistics
from pathlib import Path

from yoloeval.catalog import tasks
from yoloeval.scoring import quantile
from yoloeval.unseen import decision

ROOT = Path(__file__).resolve().parent/'results/unseen-002'


def read_rows(name):
rows = [json.loads(line) for line in (ROOT/name).read_text().splitlines()]
keys = [(row['role'], row['task_id'], row['trial']) for row in rows]
expected = {(role, task.id, trial) for role in ('baseline', 'candidate')
for task in tasks() for trial in (1, 2, 3)}
if len(keys) != len(expected) or set(keys) != expected:
raise ValueError('Missing, duplicate or unexpected attempts in '+name)
return {key: row for key, row in zip(keys, rows)}


def main():
for line in (ROOT/'SHA256SUMS').read_text().splitlines():
expected, name = line.split(' ', 1)
if hashlib.sha256((ROOT/name).read_bytes()).hexdigest() != expected:
raise ValueError('Published evidence changed: '+name)
results, calls, changes = (read_rows(name) for name in ('results.jsonl', 'calls.jsonl', 'changes.jsonl'))
rows = {role: [] for role in ('baseline', 'candidate')}
for key, envelope in results.items():
role, task, trial = key
row = envelope['result']
if (row['task_id'], row['trial']) != (task, trial) or row['status'] != 'completed':
raise ValueError('Incomplete or misidentified result: '+str(key))
if row['passed'] != (not row['agent_error'] and all(c['passed'] for c in row['checks'])):
raise ValueError('Pass flag differs from checks/error: '+str(key))
trace = calls[key]['calls']
if row['usage']['model_calls'] != len(trace) or not all(c['complete'] for c in trace):
raise ValueError('Missing or undrained calls: '+str(key))
allowed = {'muse-spark-1-3', 'jev-1.13.0'} if role == 'candidate' else {'muse-spark-1-3'}
for call in trace:
if call['requested_model'] not in allowed:
raise ValueError('Unexpected model: '+str(key))
if call['status'] == 200 and call.get('returned_model') != call['requested_model']:
raise ValueError('Successful call returned a different model: '+str(key))
rows[role].append(row)
computed = decision(rows)
recorded = json.loads((ROOT/'decision.json').read_text())
for field in ('status', 'accepted', 'complete', 'counts', 'critical_regressions'):
if computed[field] != recorded[field]:
raise ValueError('Acceptance result differs: '+field)
for field in ('candidate_mean_s', 'candidate_p50_s'):
if not math.isclose(computed[field], recorded[field], abs_tol=1e-9):
raise ValueError('Timing differs: '+field)
summary = {}
for role, samples in rows.items():
durations = [r['timing']['agent_wall_s'] for r in samples]
summary[role] = {'attempts': len(samples), 'passed': sum(r['passed'] for r in samples),
'mean_s': statistics.mean(durations), 'p50_s': quantile(durations, .5),
'p95_s': quantile(durations, .95)}
print(json.dumps({'audit_passed': True, 'acceptance': computed['accepted'], 'summary': summary,
'scored_calls': sum(len(c['calls']) for c in calls.values()),
'preserved_patches': len(changes)}, indent=2))


if __name__ == '__main__':
main()
32 changes: 32 additions & 0 deletions evals/results/unseen-002/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,32 @@
# Recorded unseen validation: 26 September 2026

| Metric | Original vanilla | Frozen winner |
|---|---:|---:|
| Passed | 45/60 | 53/60 |
| Mean agent time | 70.58s | 47.37s |
| Nearest-rank P50 | 54.30s | 30.92s |
| P95 | 180.06s | 130.77s |

The winner met the preregistered gate: at least vanilla's passes, no paired storage-check regressions, and lower mean and P50. Mean time fell **32.9%** and P50 **43.1%**, including failures. This comparison used the original vanilla revision `6a00e38` and the frozen Round 032 winner (`b1d5a3f`, published with identical product source at `review/frozen-r032`). Both used Muse `muse-spark-1-3`; the winner also used Jev `jev-1.13.0`.

## Audit without running models

From the repository root:

```sh
python3 evals/audit_results.py
```

The audit checks evidence hashes, all 120 task/repetition identities, every pass flag against its checks/error, all 393 drained model calls and model identities, and recalculates the acceptance decision and timings. `results.jsonl`, `calls.jsonl`, and `changes.jsonl` each retain one entry per attempt, including failures. `preregistration.json` contains the original schedule and hashes. `provenance.json` documents source hashes and replacement of local filesystem prefixes.

The full 109 MB working archive includes screenshots, traces, app trees and offline controls. It is retained locally and **not** uploaded here. The compact published data supports numeric and check-outcome auditing; a fresh [replication](../../README.md) regenerates browser and application evidence. It is not a byte-for-byte substitute for that archive.

## Failures and limitations

Vanilla had five attempts ending in upstream service errors, four timeouts and six requirement failures. The winner had four upstream service errors, one timeout and two requirement failures. Its three non-service failures were quiz cases: creation used the wrong option labels, a feature repair exhausted the edit budget, and one bug fix retained the old grading loop alongside its replacement.

A service outage caused seven consecutive upstream failures. The campaign paused for approximately 561 seconds between completed attempts, then resumed after recovery. Ten tiny operational probes are logged separately and excluded; every scored service error remained in the 120 results. Removing both sides of any pair with a non-200 call leaves 54 pairs: vanilla 44/54, winner 51/54, with approximately 35.2% lower winner mean time. That is a post-hoc sensitivity check, not the acceptance rule.

Campaign 001 was invalidated after a native search-input grader false negative. Campaign 002 reran all 120 attempts with the corrected grader and unchanged actors. None of campaign 001's rows is substituted here. The previously tuned, different suite reached 60/60; that exposed result is not this unseen validation.

Five app families and correlated repetitions do not establish universal reliability. These tasks are now published and exposed. The later integration with current main has offline test coverage but has **not** been assigned these frozen binary's performance numbers.
9 changes: 9 additions & 0 deletions evals/results/unseen-002/SHA256SUMS
Original file line number Diff line number Diff line change
@@ -0,0 +1,9 @@
339370d90d44f810b93555c2a88f297933f38e17a307a2eabbf88f5a4707f99d README.md
5e13560e4c23037dc9294720953a1d16bd956fa8ac5917ea6be132ef5ac21492 calls.jsonl
81ab76dc51abc26f258c71b30d53fd64db9dfdedc704b278cc822b66b0e76b7e changes.jsonl
9231973b7574c7f2da3b18950e8d176069580b0717a4f9f8a582fa5c60f4058c decision.json
b2efac2d65b7eb95c9b501ab96c2a3440a48a3be10c0ff8e73c7ff6fb1caeaf8 preregistration.json
d891753fe6d9616eb1e1c8f9bde81066f0d4110dc50cd629ff71c6a66b166461 provenance.json
3949da99f46048850d95ccc54ea7e957c6548ce7c60ff7b1598ff27a846f306a provider-health-probes.jsonl
2d006eb27e50e6126740a0337559d7d966bc98e804ddfd2e8e427ada9ffe6334 provider-outage-hold.json
73d8f9fa5cac0b06f62e2cf7a698a4b11985fb77ff7f8617097fa4e616a8d3a9 results.jsonl
Loading
Loading