DocsScreenshots & Visualbrowser_show_grid_overlay

browser_show_grid_overlay

browser_show_grid_overlay

Display an XY coordinate grid overlay on the page with position labels at intersections. Essential for finding exact pixel coordinates when using browser_click with coordinates or browser_drag_drop. The grid helps identify precise positions for mouse operations.

When to use browser_show_grid_overlay

Use browser_show_grid_overlay when you need to capture the visual state of a page for review or debugging. It is part of Owl Browser's Screenshots & Visual toolset and runs inside a self-hosted, source-level stealth engine, so every call inherits the same undetectable browser fingerprint as the rest of your automation — no separate anti-detect setup required.

Usage Example

1234567891011
import asyncio
from owl_browser import OwlBrowser, RemoteConfig
# Async usage
async with OwlBrowser(config) as browser:
context = await browser.create_context()
context_id = context["context_id"]
await browser.show_grid_overlay(
context_id=context_id
)

Parameters

Required

context_idstringrequired

The unique identifier of the browser context (e.g., 'ctx_000001')

Optional

horizontal_linesnumber

Number of horizontal lines to display from top to bottom. Default: 25. More lines = more precise coordinate reading but more visual clutter

vertical_linesnumber

Number of vertical lines to display from left to right. Default: 25

line_colorstring

CSS color for grid lines (use low alpha for visibility). Default: 'rgba(255, 0, 0, 0.15)'

text_colorstring

CSS color for coordinate labels at intersections. Default: 'rgba(255, 0, 0, 0.4)'

Response

Returns a JSON object with the operation result.

{
  "success": true,
  "result": <value>
}

Frequently Asked Questions

What does browser_show_grid_overlay do?

Display an XY coordinate grid overlay on the page with position labels at intersections. Essential for finding exact pixel coordinates when using browser_click with coordinates or browser_drag_drop. The grid helps identify precise positions for mouse operations. It belongs to Owl Browser's Screenshots & Visual category and is available through the REST API, the Python SDK (browser.show_grid_overlay()), the Node.js SDK, and the MCP server.

What parameters does browser_show_grid_overlay accept?

browser_show_grid_overlay accepts 1 required parameter (context_id) and 4 optional parameters. All parameters are sent as JSON in a POST request to /api/execute/browser_show_grid_overlay.

Is browser_show_grid_overlay detectable by anti-bot systems like Cloudflare or DataDome?

No. browser_show_grid_overlay executes inside Owl Browser's Chromium engine, which applies fingerprint spoofing at the C++ source level rather than through JavaScript patches. Every tool call shares the same consistent, human-like fingerprint, so anti-bot systems such as Cloudflare, DataDome, and Akamai see an ordinary browser.

Related Tools

browser_screenshot

Capture a PNG screenshot with configurable modes. 'viewport' (default) captures the current visible area, 'element' captures a specific element by CSS selector or natural language description, 'fullpage' captures the entire scrollable page. Returns base64-encoded image data. Screenshots capture exactly as rendered, including all dynamic content, images, and styling. Useful for visual verification, debugging, and AI vision analysis.

browser_highlight

Visually highlight an element on the page with a colored border and background overlay. Useful for debugging element selection - verify which element will be clicked before performing actions. The highlight persists until the page is navigated or refreshed.

browser_create_context

Create a new isolated browser context with its own cookies, storage, and optional proxy configuration. Each context acts as an independent browser session. Use this to create multiple isolated browsing sessions, configure proxy/Tor connections, load browser profiles with saved fingerprints, and enable/disable LLM features. Returns a context_id to use with other browser tools.

browser_navigate

Navigate the browser to a specified URL. This is a non-blocking operation that starts navigation and returns immediately. Use browser_wait_for_network_idle or browser_wait_for_selector to wait for the page to fully load. Supports HTTP, HTTPS, file, and data URLs. When wait_until is set (load, networkidle, fullscroll, domcontentloaded) and the page declares WebMCP tools, the response includes a webmcp_tools array containing the full tool definitions (name, description, inputSchema). Use browser_webmcp_call_tool to execute any of these tools directly.

browser_observe

Agent-native page observation. Returns the compacted OwlMark render (text-only structural view of the page), a handle table of interactive elements with stable tokens, page metadata, and a token estimate. Pass a handle token (e.g. 'b3') or 'pm:N' to browser_click/browser_type. Requires the context to be created with render_mode 'agent' or 'both'. ~20-100x fewer tokens than a screenshot for AI agent page understanding.

browser_click

Click on an element using CSS selector, XY coordinates, or natural language description. Supports semantic element finding using AI - describe what you want to click (e.g., 'login button', 'search icon') and the system will locate the right element. Simulates a real mouse click with proper event dispatch. Optionally hold the mouse button for press-and-hold interactions using hold_ms.

Browse the full Owl Browser API reference or get started with the Python SDK and Node.js SDK.