Back to all posts

Driving Owl Browser with System 1 Models: Constrained Decisions, Local Bindings, and Zero Key Leakage

5 min read
Akram H. S.
Akram H. S.Founder & CTO

Owl Browser now speaks two distinct architectures of AI. Generative large language models read tool schemas and emit raw function calls like browser_click with a selector. That workflow works well with our REST API and MCP endpoints. System 1 models operate on a fundamentally different principle: they are constrained-decision evaluators. You provide them with structured state and typed questions, and they return one key from an explicit list of options alongside a probability distribution. They do not generate prose, code, CSS selectors, or JSON tool calls. Handing them an OpenAPI or MCP tool schema accomplishes nothing.

Owl Browser 1.4.0 introduces /sys1, a dedicated API surface that bridges the gap. It translates the active page into a closed set of complete, executable candidate actions described in plain English, and binds each option to a server-side tool call. When your evaluator picks a key, Owl executes the underlying action. The release also ships browser core support for Chrome 153 and 154, and upgrades our Python and TypeScript SDKs to version 2.2.0.

Why Evaluators Need a Different Surface

When an agent powered by a generative model drives a browser, it inspects an accessibility tree or OwlMark render and invents a tool invocation. An evaluator model cannot invent anything. It requires a bounded decision space: given the current goal and page state, which of the following concrete choices is correct?

Converting a modern DOM into an executable choice set requires strict constraints:

  • Every candidate must be an actionable, complete call. An option cannot be an abstract intent like 'type username'; it must pair a specific DOM element with a specific value ready for execution.
  • Interactive elements generate clicks: links, buttons, checkboxes, tabs, and menu items each receive an absolute candidate key.
  • Form fields require supplied values: if you do not pass a value definition to the observation call, the page produces clicks only. Owl will not guess what to type into a blank input.
  • Unenumerable controls yield nothing: continuous inputs like free-form sliders or drag handles cannot be turned into a safe closed set, so Owl omits them rather than guessing.
  • The none option is always present: every candidate set ends with a fallback choice. Without it, an evaluator model is forced to choose the least harmful bad option when a page is broken, stuck, or already finished.

The Four-Stage Loop

A core rule of Owl Browser is that the browser never contacts a model provider. There are no vendor SDKs in the browser binary, no outbound model requests on your network, and no API keys stored in configuration files. You own your provider credentials, and you decide which model evaluates your requests.

The interaction pattern splits into four explicit steps: observe, ask, answer, and execute. Only the final step applies side effects to the browser page.

# 1. Observe: generate the state and candidate questions
curl -X POST "$OWL/sys1/observe" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "context_id": "ctx_0001",
        "goal": "Fill in the email address field",
        "values": [
          {
            "id": "email",
            "text": "ada@example.com",
            "description": "the user's email address"
          }
        ]
      }'

The response provides a snapshot identifier, candidate definitions, and a ready-to-forward request payload tailored for evaluator models:

{
  "state": {
    "goal": "Fill in the email address field",
    "page": {
      "url": "https://example.com/signup",
      "title": "Create Account",
      "view": "<compact page render>"
    },
    "suppliedValues": [
      {
        "id": "email",
        "description": "the user's email address"
      }
    ]
  },
  "questions": {
    "action": {
      "type": "choice",
      "instructions": "Which single listed action best advances state.goal?",
      "criteria": {
        "c0": "Click link \"Terms of Service\"",
        "c2": "Fill textField \"Email Address *\" with supplied value \"email\": the user's email address",
        "c7": "Scroll to the bottom of the page",
        "none": "No listed action is suitable, more information is needed, or the goal is already achieved"
      }
    },
    "goal_visible": {
      "type": "noul",
      "instructions": "Does state.page visibly demonstrate that state.goal is ALREADY achieved?"
    }
  }
}

Zero-Leakage Credential Isolation

Notice what appears in state.suppliedValues above: only the identifier email and its description the user's email address. The actual text ada@example.com is retained entirely within the browser host snapshot memory.

When sending requests to an external model provider, sensitive data such as passwords, personal identifiers, and session tokens never leave your infrastructure. The model selects candidate key c2, and Owl injects the corresponding value locally when executing the action.

Separating Decisions From Execution

After querying your model provider, you post the response back to /sys1/answer. This endpoint is strictly inert: it validates the selection against the snapshot criteria and classifies the status, but does not interact with the page.

# 2. Answer: validate and score without executing
curl -X POST "$OWL/sys1/answer" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "snapshotId": "s1_98a7df8b",
        "answers": {
          "action": {
            "type": "choice",
            "choice": "c2",
            "confidence": 0.93,
            "probabilities": { "c2": 0.93, "none": 0.07 }
          },
          "goal_visible": {
            "type": "noul",
            "noul": 0.05
          }
        }
      }'

The reply indicates whether the candidate passed validation checks:

{
  "status": "proposed",
  "candidateId": "c2",
  "confidence": 0.93,
  "goalProbability": 0.05,
  "executed": false
}

Because /sys1/answer does not execute the action, you can inspect status before committing changes. If the model is uncertain, you can fall back to a human operator or abort the run:

StatusDescription
proposedActionable candidate above min_confidence (default 0.8)
needs_reviewCandidate selected, but confidence falls below threshold
no_actionModel selected the none escape option
goal_reportedGoal completion probe indicates the task is finished
expiredThe snapshot aged out past its 120-second validity window

When ready to apply the step, call /sys1/execute with the snapshot and candidate identifier:

# 3. Execute: trigger the side effect
curl -X POST "$OWL/sys1/execute" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"snapshotId":"s1_98a7df8b","candidateId":"c2"}'

Snapshots are single-use. The server-side binding is consumed upon execution, preventing duplicate submissions or replayed network calls from firing twice.

Handling Large Pages with Candidate Windows

When a complex web application contains hundreds of interactive elements, the candidate list can exceed single-call context budgets. Owl enforces candidate limits locally (up to 254 candidates plus the mandatory none option, roughly 60 KB total payload) so oversized pages fail fast at the browser rather than generating costly provider errors.

The observation call returns a coverage object indicating whether the page was windowed:

{
  "coverage": {
    "totalCandidates": 412,
    "candidateOffset": 0,
    "returned": 100,
    "nextCandidateOffset": 100
  }
}

Candidate keys are absolute across windows (c0, c1, up through c411), ensuring a key never refers to two different actions when walking pages. Calls can target specific page segments using candidate_offset, max_candidates, region (such as main, nav, form, or dialog), and detail='min'.

SDK 2.2.0: Python and TypeScript

Both official SDKs have been updated to version 2.2.0 with native support for the /sys1 workflow. The sys1.step() helper handles observation, provider invocation, confidence evaluation, and execution in a single call while keeping provider credentials in your code.

In Python:

from owl_browser import OwlBrowser, RemoteConfig, Sys1Value

async with OwlBrowser(RemoteConfig(url=OWL_URL, token=OWL_TOKEN)) as browser:
    ctx = await browser.create_context(render_mode="agent")
    await browser.navigate(context_id=ctx, url="https://example.com/signup")

    step = await browser.sys1.step(
        ctx,
        goal="Fill in the email address field",
        values=[
            Sys1Value(
                id="email",
                text="ada@example.com",
                description="the user's email address",
            )
        ],
        ask=call_model_provider,  # your private function and credentials
    )

    print(step.decision.status, step.execution.description)

In TypeScript:

import { OwlBrowser } from '@olib/owl-browser';

const browser = new OwlBrowser({ url: process.env.OWL_URL, token: process.env.OWL_TOKEN });
const ctx = await browser.createContext({ renderMode: 'agent' });
await browser.navigate({ contextId: ctx, url: 'https://example.com/signup' });

const step = await browser.sys1.step(ctx, {
  goal: 'Fill in the email address field',
  values: [
    {
      id: 'email',
      text: 'ada@example.com',
      description: "the user's email address",
    },
  ],
  ask: async (request) => callModelProvider(request),
});

console.log(step.decision.status, step.execution?.description);

Setting auto_execute=False in Python or autoExecute: false in TypeScript pauses after decision resolution, returning the proposed step for verification before any mutation occurs.

Chrome 153 and 154 Core Updates

Alongside the System 1 surface, version 1.4.0 updates our browser engine with support for Chrome 153 and 154. The profile roster, Client Hints generator, and anti-detect fingerprinting tables have been refreshed to match the latest stable browser builds. Context pinning now spans Chrome versions 143 through 154, ensuring your automation runs match modern browser signatures on the web.

Availability

Owl Browser 1.4.0 is available today. Instances expose the System 1 reference documentation at GET /sys1/docs, generated directly from active engine limits. Python and Node packages are published on PyPI and npm as version 2.2.0.

Test Owl on your workflow

Compare detection results, latency, and failure rates on the sites and network you use.

Get Started Now