Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 15 additions & 0 deletions .env.example
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
# TypeSafe's Jev decides which action to take.
JEV_API_KEY=
JEV_MODEL=jev-latest
# JEV_URL=https://api.typesafe.ai/v1/systemone

# A small text model writes strings, only when the operation is TYPE_TEXT.
TEXT_MODEL_API_KEY=
TEXT_MODEL_BASE_URL=https://api.openai.com/v1
TEXT_MODEL=gpt-4o-mini
TEXT_MODEL_REASONING=none

# Moli, via Lexmount.
LEXMOUNT_API_KEY=
LEXMOUNT_PROJECT_ID=
LEXMOUNT_BASE_URL=https://api.lexmount.com
114 changes: 113 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
@@ -1 +1,113 @@
# jev-nolayout
# Jev NoLayout

**A browser agent that never asks the layout engine anything.**

Give it one goal. [TypeSafe's Jev](https://docs.typesafe.ai/introduction) picks an operation and an element. A small model writes text only when the operation is `TYPE_TEXT`.

Built for [Moli](https://browser.lexmount.com), which keeps page structure and interaction state in memory and renders only when a picture is actually needed. Geometry there is a snapshot from the last render, so this agent reads **structure** instead: semantics for what is live, `textContent` for what it says, and element dispatch for what it does.

## Why no layout

Same page, same selectors. The only difference is what the extractor asks:

| Asking | Controls found | Text available |
| --- | --- | --- |
| Geometry — `getBoundingClientRect`, `innerText` | **5** | **25 chars** |
| Structure — semantics, `textContent` | **134** | **89,091 chars** |

*Google Flights on Moli. The DOM is identical in both cases — 146 interactive elements — but 140 of them report a zero-sized box, because the box was measured before the page finished changing.*

Nothing in `snapshot.js` calls `getBoundingClientRect`, `checkVisibility`, `elementFromPoint` or `innerText`. Actions are dispatched on the element, never at a coordinate, so a stale layout cannot misdirect a click.

## The action space

Every observation produces a fresh element table:

```text
[1] combobox Where from? · Zürich
[2] combobox Where to? · empty
[3] textbox Departure · date picker · empty
[4] button Done · date picker
[5] button Done · 2 of 3
...
```

Operations are `CLICK`, `TYPE_TEXT`, `SELECT`, `WAIT`, `DONE` and `BLOCKED`. Only observed elements are ever offered, so the model cannot name one that does not exist.

```text
one decision request
┌───────────────────────────┐
page → element table → operation │
│ click_target │
│ type_text_target │
│ select_target, if present │
└─────────────┬─────────────┘
use the matching target
│
CLICK [7] ─────┤──→ browser
TYPE_TEXT [1] ─────┘
↓
small model → text → browser
```

Target questions are speculative: if the operation is `CLICK`, only `click_target` can execute. Two decisions, **one network round trip**.

### Labels carry location

Dropping the viewport cull surfaces every control with a given name, not just the one on screen. Google's date picker has four buttons that all read `Done`, and only one commits the date. So same-named controls are labelled by where they live:

```text
Done · date picker ← the one that confirms
Done · 2 of 3
Done · 3 of 3
```

## Try it

```bash
git clone https://github.com/lexmount/jev-nolayout.git
cd jev-nolayout
uv sync
cp .env.example .env
# Add JEV_API_KEY, TEXT_MODEL_API_KEY and your Lexmount credentials.

uv run jev-nolayout https://en.wikipedia.org/wiki/Espresso "Open the article about Latte"
```

```text
goal Open the article about Latte
from https://en.wikipedia.org/wiki/Espresso
browser Moli

1. CLICK caffè latte · 1 of 2
5412 ms decision 1397 ms 703 actions offered
clicked
2. DONE
3885 ms decision 1247 ms 282 actions offered

done · 2 steps · 9.4s
https://en.wikipedia.org/wiki/Latte
```

Pass `--browser normal` to run the same agent against standard Chrome. It works there too — reading structure is not a workaround, it is simply a better question.

## In code

```python
from jev_nolayout import Agent, moli_session

with moli_session() as browser:
browser.navigate("https://docs.python.org/3/")
for state in Agent(browser, "Go to the Standard Library reference").run():
print(state.steps[-1])
```

`examples/flights.py` runs a live Google Flights search and verifies the result against the page itself, not against the model's claim of success.

## Status

Multi-step navigation is solid. Heavy single-page applications that swap a field for a popup mid-interaction are not yet reliable — see `examples/flights.py`. Progress and open problems are tracked in the issues.

## License

Apache 2.0
60 changes: 60 additions & 0 deletions examples/flights.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
"""A live Google Flights search on Moli, verified independently of the model."""
import base64
import os
import sys
from urllib.parse import parse_qs, urlparse

sys.path.insert(0, os.path.dirname(os.path.dirname(os.path.abspath(__file__))))

from jev_nolayout import Agent, moli_session # noqa: E402
from jev_nolayout.cli import load_env # noqa: E402

URL = "https://www.google.com/travel/flights?hl=en"
DEPART = os.environ.get("DEPART_ON", "September 27, 2026")
GOAL = (f"Find one-way flights from Zurich to London on {DEPART}, for one adult in economy. "
"Stop when matching flight options are visible. Do not select or book a flight.")


def verify(snapshot):
"""Check the page itself, not the model's claim of success."""
parsed = urlparse(snapshot.url)
encoded = parse_qs(parsed.query).get("tfs", [""])[0]
try:
month, day, year = DEPART.replace(",", "").split()
months = "JanFebMarAprMayJunJulAugSepOctNovDec"
stamp = f"{year}-{months.index(month[:3]) // 3 + 1:02d}-{int(day):02d}"
decoded = base64.urlsafe_b64decode(encoded + "=" * (-len(encoded) % 4))
date_in_url = stamp.encode() in decoded
except Exception:
date_in_url = False
values = {a.label.split(" · ")[0].strip(): (a.current_value or a.value)
for a in snapshot.actions}
flights = [a.label for a in snapshot.actions if "Select flight" in a.label]
return {
"search_page": parsed.path == "/travel/flights/search",
"origin": values.get("Where from?") == "Zürich",
"destination": values.get("Where to?") == "London",
"date_in_url": date_in_url,
"has_results": bool(flights),
}


def main():
load_env()
with moli_session(os.environ.get("BROWSER_MODE", "light")) as browser:
browser.navigate(URL)
agent = Agent(browser, GOAL, on_step=lambda s: print(
f" {s.elapsed_ms:>6} ms {s.operation:<10} {s.label[:46]}", flush=True))
state = None
for state in agent.run(): # noqa: B007 - we want the last yield
pass
checks = verify(browser.observe())

print(f"\n {state.status} {len(state.steps)} steps {state.elapsed_ms / 1000:.1f}s")
for name, ok in checks.items():
print(f" {'PASS' if ok else 'FAIL'} {name}")
sys.exit(0 if all(checks.values()) else 1)


if __name__ == "__main__":
main()
7 changes: 7 additions & 0 deletions jev_nolayout/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
"""A browser agent that never asks the layout engine anything."""
from .agent import Agent, Run, Step
from .browser import Action, Browser, PageChanged, Snapshot
from .session import moli_session

__all__ = ["Agent", "Run", "Step", "Action", "Browser", "PageChanged",
"Snapshot", "moli_session"]
157 changes: 157 additions & 0 deletions jev_nolayout/agent.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,157 @@
"""The loop: observe structurally, decide once, act on the element, repeat."""
from __future__ import annotations

import contextlib
import time
from dataclasses import dataclass, field

from .browser import Browser, PageChanged
from .model import choose, field_text

MAX_STEPS = 40


@dataclass
class Step:
n: int
operation: str
label: str = ""
text: str = ""
outcome: str = ""
actions_offered: int = 0
decision_ms: int = 0
elapsed_ms: int = 0


@dataclass
class Run:
goal: str
status: str = "ready"
steps: list[Step] = field(default_factory=list)
history: list[dict] = field(default_factory=list)
elapsed_ms: int = 0
url: str = ""


class Agent:
"""Drive one goal to completion on an already-open page."""

def __init__(self, browser: Browser, goal: str, on_step=None):
self.browser = browser
self.goal = goal
self.on_step = on_step
self.run_state = Run(goal=goal)

def run(self):
started = time.perf_counter()
blank_reads = 0

for n in range(1, MAX_STEPS + 1):
step_started = time.perf_counter()
try:
snapshot = self.browser.observe()
except PageChanged:
blank_reads += 1
if blank_reads >= 3:
self.run_state.status = "blocked"
break
time.sleep(0.4)
continue
blank_reads = 0

if not snapshot.actions:
self.run_state.status = "blocked"
break

decision = choose(snapshot, self.goal, self.run_state.history)
operation = decision["operation"]
action = decision["action"]

step = Step(n=n, operation=operation, decision_ms=decision["ms"],
actions_offered=len(snapshot.actions),
label=action.label if action else "")

if operation in {"DONE", "BLOCKED"}:
step.outcome = operation.lower()
step.elapsed_ms = round((time.perf_counter() - step_started) * 1000)
self._record(step)
self.run_state.status = "done" if operation == "DONE" else "blocked"
break

if operation == "WAIT":
self.browser.settle()
step.outcome = "waited"
step.elapsed_ms = round((time.perf_counter() - step_started) * 1000)
self._record(step)
yield self.run_state
continue

text = ""
if operation == "TYPE_TEXT":
text = field_text(self.goal, action, self.run_state.history)
step.text = text
if not text:
# Better to lose a step than to submit an empty field.
step.outcome = "no value to type"
step.elapsed_ms = round((time.perf_counter() - step_started) * 1000)
self._record(step)
yield self.run_state
continue

before = snapshot.marker
before_actions = snapshot.actions
try:
step.outcome = self.browser.act(action, text)
except PageChanged as error:
step.outcome = f"stale: {error}"
step.elapsed_ms = round((time.perf_counter() - step_started) * 1000)
self._record(step)
yield self.run_state
continue

try:
after = self.browser.observe()
except PageChanged:
# act() swallows PageChanged inside settle(), so a navigation
# still in flight when the settle timer expires surfaces here.
# Record the action as taken -- it was -- and let the next
# iteration read the page it landed on.
self.run_state.history.append(
{"operation": operation, "label": action.label,
"text": text or None, "page_changed": True})
step.outcome += " (navigating)"
step.elapsed_ms = round((time.perf_counter() - step_started) * 1000)
self._record(step)
yield self.run_state
continue

changed = after.marker != before
# Say what the action produced, not just that something moved. A
# fill that opens an autocomplete list replaces the field it was
# typed into, so the next observation no longer shows the value --
# without this note the model reads that as "the text did not take"
# and types it again, forever. Naming the new options tells it the
# next move is to pick one.
opened = [a.label for a in after.actions
if a.role in {"option", "gridcell", "menuitem"}
and a.node not in {b.node for b in before_actions}][:6] if changed else []
self.run_state.history.append(
{"operation": operation, "label": action.label,
"text": text or None, "page_changed": changed,
"now_offered": opened or None})
step.outcome += "" if changed else " (no change)"
step.elapsed_ms = round((time.perf_counter() - step_started) * 1000)
self._record(step)
yield self.run_state
else:
self.run_state.status = "max_steps"

self.run_state.elapsed_ms = round((time.perf_counter() - started) * 1000)
with contextlib.suppress(PageChanged):
self.run_state.url = self.browser.observe().url
yield self.run_state

def _record(self, step: Step) -> None:
self.run_state.steps.append(step)
if self.on_step:
self.on_step(step)
Loading
Loading