browserscale MCP server

Model Context Protocol endpoint (Streamable HTTP).

Endpoint: /mcp

Auth: Authorization: Bearer <your-api-key>

Docs: browserscale.cloud/docs

Lifecycle

Actions

Identity

Lifecycle

rent

Rent a fresh browser session and get back its sessionId + grpcUrl (pass both to the action tools). rentSeconds is the session lifetime. An upstream proxy is REQUIRED — the session's exit IP comes from it, and BOTH the geo/locale (countryCode) AND the timezone are auto-matched to that exit IP server-side. So leave countryCode and timezone empty in normal use — only set them to deliberately override the proxy-derived values. countryCode and timezone are both optional. The API key is taken from the request's Authorization header, not an argument. Remember to stop the session when done. Maps to SDK RentBrowser. Example: {'rentSeconds':600,'proxyHost':'gw.proxy.io','proxyPort':8080,'proxyUsername':'u','proxyPassword':'p'}

stop

Release a session (frees the browser and its proxy). Needs only sessionId (the API key comes from the Authorization header; no grpcUrl needed). Call this when the flow is finished; sessions also expire on their own after rentSeconds. Maps to SDK StopBrowser.

Actions

navigate

Navigate the session's page to a URL. Returns as soon as the main-frame navigation COMMITS (response received, new document selected) — NOT when the page is loaded or interactive. DOMContentLoaded / load / SPA hydration may still be in flight. Do NOT observe or click right after; follow with wait for the state you need, e.g. css for a known landmark or js 'location.pathname.startsWith(\'/login\')' / 'document.querySelector(\'#email\')!==null'. When writing code, the same pattern: Navigate then Wait for a known element. Maps to SDK Navigate. Example: {'sessionId':'...','url':'https://example.com/login'}

observe

First look at any page AND the cheapest way to re-read its current state. Each frame opens with header lines carrying the URL, the title and the scroll position, then one line per visible element: tag#id[backendNodeId] key="value" ... flags "label". Spans every frame, pierces open AND CLOSED shadow roots, enumerates <select> options, and reports live form state — value= is what is typed in right now (passwords as a length), checked= for boxes. The trailing quoted string is always the label or text, never the value, so an empty and a prefilled field are distinguishable. Re-run it after every navigation and after any action that changes the page: one call replaces a handful of evaluate probes. Acting on it: pick the target you can also use tomorrow — name=, then an id flagged with neither id-not-unique nor id-not-selectable, then a text/label relation as a js target, never a generated class name. Act through that css/js target right away and every call that worked is already a line of your script; click/fill return the backendNodeId they hit, so compare it against the observation to prove you targeted the right element. Use backendNodeId as a target where no durable anchor exists, for throwaway actions, or to reach what css cannot express — it dies with the document. Flags: click = interactive; offscreen = scroll_to it first; disabled/required/readonly/selected; hidden="reason" = interactive but not seeable. A '# budget spent' line means the rest of that frame is reduced to interactive elements only — scroll and observe again, or narrow with viewportOnly. Subtree: after the first full look, pass css / js / backendNodeId (+ frame when needed) to observe only that element's subtree — follow-up looks at a form then cost the form, not the ads around it. Child iframes inside the scope are still visited. Maps to SDK GetObservation.

evaluate

Run a JavaScript expression in the page's MAIN frame and get its JSON value back. Two jobs. (1) BUILD AND VERIFY A TARGET you will reuse: run the exact expression you intend to pass as click/fill js= — and to bake into generated code — and confirm it resolves to exactly one element. Doing that here once is far cheaper than debugging a wrong selector later. (2) ONE-OFF PROBES that observe does not cover — a computed style, a data-* attribute, an aria-state. NOT for page state: URL, title, scroll position, live input values, checkbox state, visible text and duplicate-id warnings all come back from a single observe call; re-deriving them here costs several times more per look. NOT for driving the UI: el.click(), el.value = '...', and .remove() on overlays/headers — click already evades sticky/fixed occluders; if still blocked, observe then dismiss via a real close control or scroll_to, never DOM surgery. Synthetic events skip the real focus/keyboard path and are detectable — use click/fill/type/press_key. If the expression returns a DOM element you get its backendNodeId instead of a value. Runs in the main frame only — for an element inside an iframe use click/fill/wait with the frame parameter. The expression runs in a __wrc-aware context: call __wrc.shadow(host) to reach into a CLOSED shadow root (el.shadowRoot is null there). Maps to SDK Evaluate. Deeper: docs/guides/evaluation.md, docs/guides/shadow-canvas.md. Example: {'sessionId':'...','expression':'document.querySelectorAll(\'input[name=loginId]\').length'}

wait

Wait until ANY of the given conditions matches — the first to match wins (this is the multi-condition race). Use it instead of sleeping or hand-rolling a JS polling loop. Each condition is a css selector or a js expression: a js value matches when truthy, a returned element when visible. Ideal for client-side navigation, e.g. js 'location.pathname.startsWith(\'/mypage\')'. On timeout you get a per-condition breakdown explaining why each never matched (including any occluding element). Actions do NOT wait, so wait before every click/fill. Maps to SDK Wait. Deeper: docs/guides/waiting.md. Example: {'sessionId':'...','conditions':[{'css':'.success'},{'js':'document.querySelector(\'.error\')!==null'}],'timeoutMs':8000}

click

Click an element. Target it with a css selector, OR with a js expression that returns the element — that is how you click by text/role without a brittle selector, e.g. js '[...document.querySelectorAll(\'button\')].find(b => b.textContent.trim() === \'Next\')'. For a duplicated id, pick the nth: js 'document.querySelectorAll(\'#loginID\')[2]'. If the click is blocked, the intercepting (occluding) element is returned (id/class/text/position). Click already tries to evade sticky/fixed occluders — if it still fails, observe, then dismiss via the blocker's real close control or scroll_to; do not evaluate .remove(). Don't blind-retry. Maps to SDK Click. Deeper: docs/guides/locators.md, docs/guides/interaction.md.

fill

Type text into ONE specific input, strictly bound to it: it clicks to focus then sends real, per-key keyboard events, re-verifying focus each key — the right choice for SPA forms that validate on input or enable a button as you type. Target with css or a js expression. clearFirst overwrites existing content. timeoutMs/steadyMs are ADVANCED knobs that tune how long it tries to make the field focusable (omit for defaults). For an OTP/code field that auto-advances across boxes use type instead (fill pins one field and would fail there); for a silent bulk paste use insert_text, which commits via the IME path — never assign .value through evaluate, that skips the events the page listens for and is detectable. Maps to SDK Fill.

type

Type text as a per-key keyboard stream into whatever currently has focus, WITHOUT targeting or pinning a field — so the page is free to move focus mid-stream. This is the tool for one-time-code / OTP inputs that auto-advance to the next box on each digit (fill would fail there because it stays bound to one field). Nothing is focused for you: click the first field first. clearFirst clears the focused field before typing. For a single field that must stay focused use fill; for a silent bulk commit use insert_text. Maps to SDK Type.

select

Pick an <option> in a <select>. Target the <select> with css/js, then give exactly ONE of index (0-based), value (the option's value attr) or text (visible text, exact trimmed match). Fires input+change unless noEvents. Returns which option was chosen. Maps to SDK SelectByIndex/Value/Text.

move_to

Move the mouse over an element without clicking — hover to reveal dropdowns/tooltips or warm up an element that only shows an action on mouseover. Target with css/js (frame supported). Maps to SDK MoveTo. Usually you observe → moveTo → wait → click.

scroll_to

Scroll an element into view without clicking it. Actions already scroll into view themselves, so this is mainly for what observe reports as offscreen, and before an observe/evaluate that won't trigger a scroll on its own. Target with css/js/backendNodeId. Maps to SDK ScrollTo.

drag

Drag an element. absolute=false drags BY an (x,y) offset from the pickup point (sliders, range inputs); absolute=true drops AT root-viewport (x,y) coordinates (reordering cards). Target with css/js. Maps to SDK DragBy/DragTo. Example: {'sessionId':'...','css':'.slider .handle','x':120,'y':0}

press_key

Send a full key press (down+up) to whatever currently has focus — Enter to submit, Tab to advance, or a shortcut with modifiers. modifiers is a bitmask: Alt=1, Ctrl=2, Meta=4, Shift=8 (combine by adding, e.g. 2 for Ctrl). Click/fill the field first to focus it. Maps to SDK PressKey+ReleaseKey. Example: {'sessionId':'...','key':'Enter'}

insert_text

Commit a whole string at once to the focused element via the IME path — NO per-key events fire (fast, silent). Use for big text blocks where you don't need input/keydown handlers to run; use fill instead when the page validates on real typing. Focus the field first. Maps to SDK InsertText.

screenshot

Capture the current viewport as an image and return it inline so you can SEE the page (great for verifying a step or diagnosing a stuck flow). format is png (default), jpeg or webp; quality 0-100 for jpeg/webp. Maps to SDK Screenshot.

solve_captcha

Attempt to solve a captcha challenge on the current page and return the resulting token. timeoutMs bounds the solve; retryAmount is how many attempts. Maps to SDK SolveCaptcha.

Identity

get_cookies

Dump every cookie in the session as a JSON array — an identity snapshot. Save the output and replay it later with set_cookies to restore a logged-in session on a fresh rental. Maps to SDK GetCookies.

set_cookies

Inject/restore cookies. Accepts the SAME JSON array shape get_cookies returns, so a saved dump can be fed back verbatim; an existing cookie with the same (name, domain, path) is overwritten. Set cookies BEFORE navigating to the origin so the first request already carries them. Maps to SDK SetCookies.

clear_cookies

Delete every cookie in the session — start from a clean identity. Maps to SDK ClearCookies.

get_storage

Dump localStorage grouped by origin as JSON — the other half of an identity snapshot (pair with get_cookies). Pass origin to limit to one site, empty for all. Reads directly, so no page needs to be open; sessionStorage (per-tab) is NOT included. Maps to SDK GetStorage.

set_storage

Restore localStorage. Accepts the same shape get_storage returns, grouped by origin; existing keys are overwritten. Pages already open won't see the writes until they reload, so set before navigating. Maps to SDK SetStorage.

clear_storage

Delete localStorage. Pass origin to wipe one site, empty to wipe all origins. Maps to SDK ClearStorage.