# fastbrowse public documentation > Every public usage document for fastbrowse at commit 4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5. Each document appears once, with its upstream source link. Source: https://fastbrowse.ai/ --- # First run Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#try-it Install fastbrowse with uv, choose a cloud or local browser, and run a first task from the command line. ## Try it Needs [uv](https://docs.astral.sh/uv/); uv fetches Python itself (3.13 or newer). Runs use a [Browser Use Cloud](https://cloud.browser-use.com) browser (`BROWSER_USE_API_KEY`) by default: it passes bot checks a fresh local Chrome fails. Local Chrome is fully supported with `--local`. A browser already running anywhere, from a container to a hosted browser with a CDP endpoint, is driven in place with `--cdp-url ws://…` or `--cdp-port 9222`: the run opens one tab and leaves the browser as it was found. To drive a window that is already open, such as an Electron app, see [Open windows and Electron apps](/docs/attach#open-windows-and-electron-apps). A Node program needs neither uv nor Python: see [Use it from JavaScript](/docs/javascript#use-it-from-javascript). ```sh export OPENROUTER_API_KEY=... # for Jev and the LLM that plans and reads # export AI_GATEWAY_API_KEY=... # Vercel AI Gateway: Jev backup, and the LLM when OPENROUTER_API_KEY is unset export BROWSER_USE_API_KEY=... # the cloud browser; or pass --local to use Chrome uvx fastbrowse "What is the title of the top story right now?" --start https://news.ycombinator.com/ ``` Jev uses direct TypeSafe when `TYPESAFE_API_KEY` is supplied; otherwise OpenRouter is primary. Vercel AI Gateway is supported as a backup or an explicit primary. `FASTBROWSE_JEV_SOURCE` overrides automatic selection; see [provider routing](https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/docs/jev.md#provider-failover). The LLM uses OpenRouter when `OPENROUTER_API_KEY` is set, otherwise the Vercel AI Gateway's OpenAI-compatible chat completions. `uvx` runs the published package in an isolated cached environment. `uv tool install fastbrowse` keeps it on your PATH, and `uv add fastbrowse` puts it in a project. Service keys can live in a `.env` file in the working directory; [`.env.example`](https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/.env.example) shows the settings. Values named by `--secret` must be in the process environment. Steps go to stderr; the status, cost, step count and answer go to stdout. `--json` prints every step, the quotes behind the answer, and cost by component. | Flag | Effect | |:--|:--| | `--start URL` | the page to open first; worked out from the task when omitted | | `--cloud` | force cloud Chrome even when `FASTBROWSE_PROFILE` or `FASTBROWSE_HEADED` is set | | `--local` | use local Chrome instead of a Browser Use Cloud browser. Cloud is the default: it passes bot checks a fresh Chrome fails, and prints a URL to watch the run live | | `--headed` | show the local Chrome window (implies `--local`) | | `--profile DIR` | keep the local Chrome profile in `DIR`, so a site signed into there stays signed in (implies `--local`) | | `--cloud-profile ID` | run on a Browser Use Cloud profile, signed in as whoever set it up | | `--cdp-url URL` | drive a browser already running at this `ws://` or `wss://` DevTools URL instead of starting one | | `--cdp-port PORT` | the same, for a browser or Electron app listening on `127.0.0.1:PORT`; the URL is read from its `/json/version` | | `--attach` | with `--cdp-url` or `--cdp-port`, drive a window already open instead of opening a tab, and leave it open after the run | | `--target-match TEXT` | attach to the first window whose title or URL contains `TEXT` (implies `--attach`) | | `--proxy-country CC` | browse from that country (Browser Use's codes: `uk`, `de`, ...; default `us`), so a shop shows its local delivery and prices | | `--authorize` | allow submit, pay, delete and send; without it the run stops at `needs_confirmation` first | | `--secret NAME=ENV_VAR[@ORIGIN]` | let the agent type `$ENV_VAR` on the declared origin, or the `--start` origin if omitted; models only see `NAME`. An explicit origin needs no `--start` | | `--bitwarden ITEM` | match the vault login's saved URIs against `--start`, then allow its `username`, `password` and, when the item holds an authenticator key, `one_time_code` only on that start origin | | `--max-steps N`, `--max-dollars N` | optionally bound steps and model spend; defaults are unlimited. Cloud browser charges are added when it stops | | `--downloads DIR` | keep downloaded files | | `--json` | full result instead of the answer | | `--record FILE` | save an MP4 of the tab, each step captioned, ending on the answer, time and cost (needs `ffmpeg`; the captions need its libass), e.g. `recordings/demo.mp4`, which git ignores; `demo.plain.mp4` beside it has no captions. It shows what the pages showed, so watch it before sharing | | `--cursor` | draw the agent's cursor over a visible Chrome (`--headed`, or one you attach to) with [Cua Driver](https://github.com/trycua/cua), so you can watch where it acts. Off by default. It needs `cua-driver` on `PATH` and an X11 display on Linux, and does nothing without them | ```sh export SAUCE_PASSWORD=secret_sauce uv run fastbrowse "Log in as standard_user with the saved password and add the backpack to the cart." \ --start https://www.saucedemo.com/ --secret password=SAUCE_PASSWORD --authorize ``` --- # Signed-in sites and secrets Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#signed-in-sites Reuse a signed-in Chrome profile, run a cloud profile, or type origin-scoped secrets from a Bitwarden vault. ## Signed-in sites Sign in once by hand in a profile of its own, then point runs at it: ```sh google-chrome --user-data-dir="$HOME/.fastbrowse/amazon" https://www.amazon.com/ # sign in, then close Chrome uv run fastbrowse "Add a UGREEN USB-A to USB-C cable, 2m, to my cart." \ --start https://www.amazon.com/ --profile ~/.fastbrowse/amazon --headed ``` On a cloud browser the profile lives on the [Browser Use Cloud](https://cloud.browser-use.com) account rather than on disk, and `--cloud-profile ID` runs as it. Whoever signed that profile in did so once, in a browser of their own; the run inherits the cookies and no model is shown a credential: ```sh uv run fastbrowse "Add a UGREEN USB-A to USB-C cable, 2m, to my cart." \ --start https://www.amazon.com/ --cloud-profile prof_1234 ``` Or from your vault, with the [Bitwarden CLI](https://bitwarden.com/help/cli/) unlocked. The item's saved URIs must match `--start`; values are then typed only on that origin, and models see only the names `username` and `password`, plus `one_time_code` when the item holds an authenticator key (the current code, computed as it is typed): ```sh export BW_SESSION="$(bw unlock --raw)" uv run fastbrowse "Add a UGREEN USB-A to USB-C cable, 2m, to my cart." \ --start https://www.amazon.com/ --bitwarden Amazon --profile ~/.fastbrowse/amazon --headed ``` To change a password, supply its replacement as an origin-scoped secret named `new_password`, alongside the existing `password`. Refer to `new_password` in the task and use `--authorize` to allow submission. Without the replacement secret, the run asks for input; it never generates a password or reuses the old one. --- # Use fastbrowse with your agent Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/docs/skill.md#use-fastbrowse-with-your-agent Install the fastbrowse skill, configure a browser and delegate cited website tasks with explicit authorization. ## Use fastbrowse with your agent The [fastbrowse skill](https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/../skills/fastbrowse/SKILL.md) lets an agent that supports Agent Skills delegate a website task to fastbrowse and interpret its status and citations. It prefers an existing local Chrome connection, then local Chrome. It uses an MCP server configured for that browser choice, or the CLI. Cloud browsing is available when requested or local browsing is unavailable. Agents without skill support can use the CLI or MCP server directly. Installing the skill installs instructions; the browser and model credentials still need configuration. ### Install From a checkout of this repository, install the skill for your agent with the [Skills CLI](https://github.com/vercel-labs/skills): ```sh npx skills add ./ --skill fastbrowse ``` The same installer can fetch the skill from GitHub without a checkout: ```sh npx skills add agent-labs-dev/fastbrowse --skill fastbrowse ``` Choose your agent when the installer prompts. The installer uses project scope by default. Add `--global` for all projects, or `--agent codex claude-code` to select those two agents explicitly. Review the skill before installing it. For a manual install, copy `skills/fastbrowse` into `.agents/skills/fastbrowse` for Codex or `.claude/skills/fastbrowse` for Claude Code in the project where you want to browse. Keep its `references` directory beside `SKILL.md`. Personal installs use `~/.agents/skills/fastbrowse` or `~/.claude/skills/fastbrowse`. Start a new agent session after installation. ### Configure and use Install [uv](https://docs.astral.sh/uv/getting-started/installation/) for the CLI path. Set `OPENROUTER_API_KEY` in the calling process environment or a local `.env`, outside the conversation. Local Chrome is the skill default and needs only the model key. Add `BROWSER_USE_API_KEY` for cloud browsing. An existing DevTools connection reuses your Chrome; `--local` without a connection starts a separate browser. The [CLI guide](https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/../README.md#try-it) covers browser profiles, secrets, and budgets. For MCP, use the [server setup](https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/../README.md#use-it-from-an-mcp-client) instead; credentials belong to the server's environment. In Codex: ```text Use $fastbrowse to find the top story on https://news.ycombinator.com/ and return its title with a citation. ``` In Claude Code: ```text /fastbrowse Find the top story on https://news.ycombinator.com/ and return its title with a citation. ``` The skill also supports automatic selection from its description. It uses the CLI's uncapped resource defaults unless you supply limits or an existing task budget requires them. MCP server ceilings still apply. Cloud browser charges are added when the browser stops, separately from any model-spend cap. For a form, include its URL and the values to fill. State whether the agent should stop before submission or submit the specified content. The skill does not authorize writes on its own. A result of `complete` means the task was verified; other statuses describe the missing input, authorization, evidence, or access. --- # Attach to a running browser Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#open-windows-and-electron-apps Drive a Chrome or Electron window over its DevTools port, or attach to a window that is already open. ## Open windows and Electron apps An Electron app is Chromium underneath, so a run can drive it once the app exposes a DevTools port. Quit the app, start it again with `--remote-debugging-port` (an app that is already running ignores the flag), then attach to its window by title or URL. VS Code titles each window after the folder it has open: ```sh open -a "Visual Studio Code" --args --remote-debugging-port=9222 ~/code/my-project # elsewhere, pass the flag to the binary uv run fastbrowse "Open the Extensions view and report how many extensions are installed." \ --cdp-port 9222 --target-match my-project ``` The same works for a Chrome you started with `--remote-debugging-port`. With `--attach`, the run takes over an existing window instead of opening a tab of its own. Without `--start` it begins on whatever the window shows, and the window stays open when the run ends. `--target-match` picks the first window whose title or URL contains the text; `--attach` alone takes the first window, skipping DevTools. Popups the window opens join the run. Windows that nothing opened, as an Electron app's main process opens them, join only with `--target-match`, because in a browser such a window could equally be a tab you opened yourself. A DevTools port gives any program on the machine full control of the app, including its signed-in sessions. Close the app, or restart it without the flag, when you are done. --- # Models Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#models Choose planning and reading models, override them for specific tasks, and set reasoning effort. ## Models The LLM defaults to `google/gemini-3.8-flash` at low reasoning effort, with `google/gemini-3.5-flash-lite` for proposing a direct address and typing field text. Override with `FASTBROWSE_LLM_MODEL` (every purpose), `FASTBROWSE_LLM_MODEL_` (`PLAN`, `READ`, `FIELD_TEXT`, `SHORTCUT`, `RECOVER`, `COMPOSE`, `VERIFY`) and `FASTBROWSE_LLM_REASONING` (`low`, `medium`, `high`). --- # Run statuses and exits Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#results Every status and exit code, what complete means, and how a run reports an exhausted budget. ## Results The exit code identifies the run status. Configuration refusals exit 1; invalid command syntax exits 2. Programs that only check zero versus nonzero continue to work. | Status | Exit | Meaning | |:--|--:|:--| | `complete` | 0 | every information requirement is backed by a quote, and every action is confirmed on the page | | `unverified` | 10 | it believes it finished but could not back every claim | | `needs_confirmation` | 3 | stopped before an irreversible action; re-run with `--authorize` | | `needs_login` | 4 | a sign-in wall that no `--secret` covers | | `blocked` | 6 | a bot check (a CAPTCHA) that did not clear; not a sign-in, so no secret passes it | | `needs_input` | 5 | a required value or file is missing, or an upload exceeds the configured size limit | | `stuck` | 9 | recovery ran out without reaching a page state the run had not seen | | `budget_exceeded` | 7 | a resource limit was reached, or a missing cost prevents enforcing the dollar cap | | `observation_limit` | 11 | the page or required evidence cannot fit the configured prompt budget | | `unavailable` | 8 | a model or browser provider stayed unavailable through every retry; the same run later may pass | | `error` | 1 | a model or browser failure | A `budget_exceeded` JSON result includes `budget.resource` (`steps`, `seconds`, `dollars`, `jev_calls` or `llm_calls`) and `budget.limit`. Other results have `budget: null`. The embedding and MCP APIs carry the same optional object, so callers can identify the affected limit without parsing an error message. A bot check can block a cloud browser before any task steps execute. Retry with an attached local browser or choose another source for the research. Credentials do not resolve a CAPTCHA. If a model completion omits its cost, fastbrowse looks up the matching OpenRouter or Vercel generation receipt. If that receipt is still unavailable, a run with a dollar cap stops with `budget_exceeded`; this does not mean the known charges reached the cap. Supplied `inputs` can provide exact field text without a field-writing generation. A browser command that remains unanswered for 120 seconds returns `unavailable`, even if the browser still answers health checks. Step and dollar limits do not bound total elapsed time. Set `Limits(max_seconds=...)` for a run deadline. When investigating a stalled run, enable trace logging and retain the last event alongside the result. --- # Use it from Python Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#embed-it Embed run_task, request schema-validated data, read citations and events, stream frames, and attach over CDP. ## Embed it `uv add fastbrowse` first, then: ```python import asyncio from pydantic import BaseModel from fastbrowse import run_task from fastbrowse.models import Limits class Release(BaseModel): package: str version: str async def main() -> None: result = await run_task( "Find the httpx package and report its name and latest released version.", start="https://pypi.org/", output_schema=Release, limits=Limits(max_dollars=0.10), ) print(result.status, result.data, f"${result.cost.known_dollars:.4f}") for evidence in result.evidence: print(f' "{evidence.quote}" from {evidence.url}') asyncio.run(main()) ``` `run_task(cdp_url=...)` drives a browser that is already running, wherever it is, instead of starting one: the run opens its own tab and closes the tabs it owns. It leaves the browser and pre-existing tabs open; cookies and other changes made by the task can persist. `cdp_port=` finds the same browser from its DevTools port on `127.0.0.1`. `attach=True` and `target_match=` drive a window already open and leave it open, as `--attach` and `--target-match` do. Pass `browser_api_key=` to start a cloud browser; with neither argument, it runs local Chrome. Passing both is an error. `cloud_extensions=[...]` loads up to three of your account's ready Browser Use Cloud extensions, by ID, into that cloud browser; it is an error with local or attached browsers. `connect_cdp()` hands a script of your own the attached page, without the agent. Downloads go to `downloads=`, or to a scratch directory removed on exit: ```python from fastbrowse import connect_cdp async def main() -> None: async with connect_cdp(9222, target_match="my-project") as page: print(await page.observe()) ``` `resolve_cdp_port(port)` returns the `ws://` URL behind a DevTools port, for a caller that passes `cdp_url=`. `RunResult.citations` is a tuple of `Citation` objects, also importable from `fastbrowse`. Each has `id` (the number in the answer), `text` (the Notes fact), `requirement_id` (or `None`), `url`, `quote` and `deep_link`. Each claim in `result.answer` carries numbered Markdown links to its supporting facts. Counts, totals and superlatives also cite the records they were derived from, including records read on earlier pages. Only verified Notes facts supply citation URLs and quotes; an answer citing an unknown reference fails the claim check, and the run falls back to an answer drafted from verified facts. Facts omitted from the answer have no citation, and citation numbers can have gaps. Deep links follow the [WICG Text Fragments syntax](https://wicg.github.io/scroll-to-text-fragment/#syntax): `url#existing-anchor:~:text=start`. Text is percent-encoded, including hyphens, ampersands and commas. Whitespace is collapsed for the link; `quote` keeps the verbatim capture. Quotes over 120 characters with more than ten words use the first and last five words as `text=start,end`. An existing anchor is preserved; an old text directive is replaced. Pages that change or require a session may no longer show the quote. To show a run as it happens, pass `on_event=`: a `BrowserEvent` arrives first with the live-view URL of a cloud browser, then a `StepEvent` per step. `Config(step_frames=True)` adds a PNG of the page each step acted on, for an interface that renders the run; a step whose page is showing a resolved secret sends no frame. `StepEvent.step.facts` (also `StepResult.facts`) holds only the facts added by that step: text, requirement id, quote, URL, deep link and reader (`jev_choice` or `llm`), with resolved secrets redacted before delivery. `StepResult.note` carries read outcomes, dispatch details, gate refusals or recovery guidance when available; it can be `None` for an ordinary successful action. For continuous live images, pass an async `on_frame` handler accepting JPEG bytes. Frames follow the active tab and are acknowledged after delivery, with no fixed frame rate. Only the latest pending frame is kept. Handler failures are logged without stopping the run. Live frames and recordings are held back while a resolved secret may show on the page, as PNG step frames are. No handler means no live capture. Jev uses direct TypeSafe when keyed, otherwise OpenRouter, with Vercel AI Gateway also supported. See [Jev routing](https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/docs/jev.md#provider-failover) for key precedence, overrides and failover. Any other source can be passed as `run_task(jev=...)`, implementing async `evaluate(state, questions)`; `run_task(llm=...)` accepts an implementation of the `LLMClient.generate(...)` protocol in `fastbrowse.llm`. --- # Use it from JavaScript Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#use-it-from-javascript Install the TypeScript SDK, run a task, and keep secrets, authorization and output validation in your process. ## Use it from JavaScript `npm install fastbrowse` gives a Node program the same agent, with no Python and no uv on the machine. The package is a TypeScript SDK with no dependencies. With it npm installs one of five packages that hold the agent as a native binary, the one for your platform: `@fastbrowse/darwin-arm64`, `darwin-x64`, `linux-arm64`, `linux-x64` or `win32-x64`. No install script runs and nothing is downloaded on first use, so it installs under pnpm and bun with scripts blocked, and it starts offline. It needs Node.js 20 or newer. The npm packages carry the PyPI package's version and are published by the same tag, starting with the release after 0.5.18. The binary brings no browser. It finds local Chrome, starts a Browser Use Cloud browser, or drives one already running, as the command line does. It reads the same keys from the environment of your process (`OPENROUTER_API_KEY`, and `BROWSER_USE_API_KEY` for a cloud browser); `env` in `Fastbrowse.start` adds to them. ```sh npm install fastbrowse zod export OPENROUTER_API_KEY=... ``` ```ts import { Fastbrowse } from 'fastbrowse'; import { z } from 'zod'; const Release = z.object({ package: z.string(), version: z.string() }); const fb = await Fastbrowse.start({ local: true }); try { const result = await fb.run('Find the httpx package and report its name and latest released version.', { start: 'https://pypi.org/', output: Release, limits: { maxDollars: 0.1 }, onEvent: event => { if (event.type === 'step') console.error(event.step.operation, event.step.target); }, signal: AbortSignal.timeout(120_000), }); console.log(result.status, result.answer); if (result.status === 'complete') console.log(result.output.package, result.output.version); for (const evidence of result.evidence) console.log(` "${evidence.quote}" from ${evidence.url}`); } finally { await fb.close(); } ``` `Fastbrowse.start` starts one fastbrowse process and `run` sends it a task. The process serves one run at a time: a second `run` while one is active rejects with the busy error, and a program that wants runs side by side starts more instances. `close()` shuts the process down and waits for it. The process also exits when yours does, however yours ended. `run` resolves with the result whatever status the run ended in, so `needs_login` or `stuck` is read from `status` and is not an exception. It rejects when no run took place or none finished: with `RpcError` when the server refuses the request before a browser opens (a bad option, a missing key, a busy server), with `AbortError` when `signal` stopped the run, and with `ProcessExitedError` when the process is gone. A run that fails after it has started resolves with status `error`. The result is the `RunResult` the Python library returns, with the same statuses, citations and cost lines. Its fields keep their Python names, such as `final_url`, since the types are generated from the Python models; the options are camelCase. The structured data is in `data`, as in Python. The SDK adds `output`: with a Zod (4.2 or newer) or ArkType schema, or a Valibot schema wrapped by `@valibot/to-json-schema`, `output` is `data` after the schema's own validation and transforms, and has the schema's output type. Data the schema refuses rejects with `OutputValidationError`, which carries the result. A JSON Schema object is accepted too; `output` is then `data`, typed `unknown`. What a run can fill today is narrower than what a schema can say. The server accepts nested objects, arrays, `enum`, `const` and optional values, and refuses any other keyword by name before a browser opens. A run fills a flat object of required string, number, integer and boolean fields, as the example has. With any other field the run ends `unverified` with no data. The callbacks that keep credentials and the last word in your code cross the process boundary: ```ts await fb.run('Log in as standard_user with the saved password and add the backpack to the cart.', { start: 'https://www.saucedemo.com/', authorization: { irreversibleActions: true }, secrets: { refs: [{ name: 'password', origins: ['https://www.saucedemo.com'] }], resolve: name => vault.get(name), }, until: url => url.endsWith('/cart.html'), }); ``` `secrets.resolve` is called each time a value is typed and never for an origin its ref does not cover, so the value leaves your vault at that moment and a one-time code is fresh. `until` gets the address the run ended on, and anything but `true` keeps the run from `complete`. `onFrame` receives JPEG frames of the active tab; without it no frame is sent. A callback that throws ends the run with status `error`. The other options are the command line's: `inputs`, `attachments` as bytes, `downloads`, `record`, and the browser choices `chrome`, `cloudProfile`, `cdpUrl`, `cdpPort`, `attach`, `targetMatch` and `proxyCountry`, which `Fastbrowse.start` also takes as defaults for every run. A run passes `null` for one of them to go without that default. There is no binary for Alpine or another musl system, and none for a platform outside the five. There `Fastbrowse.start` rejects with an error that names the platform, and a binary of your own is named with `binaryPath` or `FASTBROWSE_BINARY`. The macOS binaries carry an ad-hoc signature and are not notarized, which a Mac that allows programs only by the team that signed them refuses. The Windows binary is not signed. Bun and Deno are untested. ## Use fastbrowse with your agent Install the [fastbrowse skill](https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/docs/skill.md) to delegate website tasks from an agent that supports Agent Skills. It uses the MCP tool when connected, or the CLI, and preserves the task's citations, status, authorization, and budget. The guide covers agent selection, configuration, and example prompts for Codex and Claude Code. Agents without skill support can use the CLI or MCP server directly. ## Run fastbrowse through Nebula [Nebula](https://www.nebula.gg/) embeds fastbrowse in its agent harness, with browser steps, results and permitted vault logins. Use Bitwarden or 1Password through Nebula with credentials scoped to authorized websites. --- # Use it from an MCP client Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#use-it-from-an-mcp-client Serve the browse tool over MCP, with server-side limits, fixed browser choices and HTTP authorization. ## Use it from an MCP client `fastbrowse-mcp` serves one `browse` tool over [MCP](https://modelcontextprotocol.io), so Claude Code, Claude Desktop, Cursor or any other MCP client can hand it a task. It returns the answer, typed `fields` if asked for, the quotes behind them, the status and what to do about it, and reports progress on every step. The answer contains numbered Markdown links. MCP's `citations` list contains `quote` and `url` from the run's evidence; it does not expose the Python `Citation` ids, requirement ids or deep-link fields. ```sh claude mcp add fastbrowse -e OPENROUTER_API_KEY=... -e BROWSER_USE_API_KEY=... \ -- uvx --from 'fastbrowse[mcp]' fastbrowse-mcp --max-dollars 0.25 ``` For a client configured by JSON, such as Claude Desktop: ```json { "mcpServers": { "fastbrowse": { "command": "uvx", "args": ["--from", "fastbrowse[mcp]", "fastbrowse-mcp"], "env": { "OPENROUTER_API_KEY": "...", "BROWSER_USE_API_KEY": "..." } } } } ``` The server's flags decide what a calling model may do; a call can ask for less, never more. | Flag | Effect | |:--|:--| | `--local`, `--headed`, `--profile DIR`, `--downloads DIR`, `--cursor` | as for the CLI, fixed for every call | | `--cloud-profile ID` | every call runs signed in as that cloud profile; a calling model cannot choose it | | `--allow-authorize` | let a call pass `authorize` to go through irreversible actions; without it they always stop at `needs_confirmation` | | `--secret NAME=ENV_VAR@ORIGIN` | typed when a call's start page is on `ORIGIN` (`https://*.site.com` covers its hosts); the model sees `NAME` only | | `--bitwarden ITEM` | a vault login a call may name in `bitwarden` | | `--max-steps N`, `--max-dollars N`, `--max-seconds N` | optional ceilings per call; defaults are unlimited | | `--max-concurrent N` | runs at once, default 1; more calls wait their turn | | `--transport http`, `--host`, `--port` | streamable HTTP at `/mcp` instead of stdio | Over HTTP, set `FASTBROWSE_MCP_TOKEN` in the environment or `.env` and every request but `/healthz` needs `Authorization: Bearer `. The server refuses to bind anything but loopback without one. A run takes seconds to minutes, so raise the client's tool timeout if it has one (`MCP_TOOL_TIMEOUT` in Claude Code). --- # Safety model Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/README.md#safety-model How fastbrowse checks authorization, handles secrets and treats page content as data. ## Safety model - **Irreversible actions.** Jev judges clicks, including links, Enter presses and acceptance of confirm, prompt or before-unload dialogs. Code-selected pagination is exempt, as are authorized actions with sufficient confidence. A refusal appears as a failed step with a reason: an unauthorized, confident action stops at `needs_confirmation`; an uncertain action goes to recovery. This is a classifier, not a guarantee: a page can word a harmful control to look harmless. - **Secrets.** Models see secret names only. A value is resolved at the moment of typing, only for its declared origin, and redacted from everything the run returns. A password field is typed only from a stored secret, never generated. - **Page content is data.** Reader and verifier prompts say so, and completion is judged against quotes and page state rather than the model's say-so. --- # How it works Source: https://github.com/agent-labs-dev/fastbrowse/blob/4372bfd6bb8ff39f7b14b1cbb7dba2285b9761d5/docs/design.md The step loop, reading and citations, prompt limits, run events and images, and browser capabilities over CDP. fastbrowse splits a browser agent into three owners: - **Jev picks.** A batched request chooses an operation and its possible targets from indexed controls. It also judges whether unread evidence should be preserved before interaction, and checks sign-in and bot walls on fresh pages. On a page denser than the step's control budget, a batched relevance check keeps the controls the task can use rather than the first ones in document order. A page too dense even on screen offers the controls that fit and scrolls for the rest rather than ending the run. On a page with more quotable spans than one choice can offer, the same kind of check narrows the passages before Jev picks a short fact. Large target sets use a second choice within the selected group. Code can follow a list's next-page link without an action-choice call. - **An LLM reads and writes.** Proposing a direct address for the task while the start page loads, planning checkable requirements from the task alone, reading page content for answers when Jev cannot pick a short fact from quoted spans, writing non-secret field text, recovering when Jev is unsure, verifying completion in the uncertain band, composing the final answer. - **Code owns the gates.** Freshness and hit-tests before every input, no automatic retry of a mutation, authorization for irreversible actions, secret resolution and redaction, budgets, and the definition of success: only `COMPLETE`, which requires every requirement evidenced. ## Authorization Jev judges every model-selected click, including links, navigation to a caller-supplied address, every Enter press and acceptance of confirm, prompt and before-unload dialogs. Code-selected pagination is exempt. An authorized action whose confidence reaches `Thresholds.sensitive_act_from` proceeds without that classification. A refusal is recorded as a failed step with a reason. An unauthorized action with sufficient confidence stops at `needs_confirmation`; an uncertain action goes to recovery. The classifier can be wrong, so this gate is not a guarantee that every externally visible change is detected. ## Navigation The operation choice can open another literal HTTP(S) address from the caller's task. The address is a closed choice, revalidated before dispatch, and uses the browser's origin grant and live access checks. Page instructions cannot add addresses to that choice. Recovery can select the same supplied addresses; addresses inferred by the startup shortcut retain their separate verification rules. An editable search field can be filled directly, including when focusing it opens an editor. Opening or focusing the field does not enter a query. ## Reading and citations The policy's read assessment distinguishes useful evidence from editable query previews and irrelevant content. Navigation notes write each source URL once and refer to it from later claims. Evidence is read before another interaction can remove it. Reads are deduplicated by document, capture hash and unresolved information requirements, so changed content can be read again. A paraphrase of a collected quote does not restore the read budget; new source quotes and newly evidenced requirements do. Identical relevance questions share one score within a pass, while controls remain separate action targets. Top-page navigation links that fail freshness twice are excluded by document and destination, even when card labels change or intervening reads add facts. Same-page actions and framed controls keep their label and context identity. A new document permits another attempt; input freshness and authorization checks still apply. For a bounded set of short quoted spans, Jev chooses a scalar fact, requests synthesis, or judges the requirement absent from the page. Identical values from the same page and frame share one choice with all their source contexts. A selected repeated value retains every span through `Fact.basis`. Uncertain choices, comparisons, partial evidence and paginated lists reach the LLM reader. `FactReader` records `jev_choice` or `llm`. Neither reader writes page text: Jev picks a value from spans code cut from the capture, and the LLM reader cites the capture's source blocks by id (one block, or consecutive blocks of one frame from the chunk it was shown). Code copies the quote from those blocks, so a table's escaped pipe or a record read as two lines cannot drop a fact, and a claim citing a block it was not shown is rejected. A prose claim can name a literal excerpt within its cited blocks. Code requires a unique match in the offered chunk and copies the original span, retaining its source hash and offsets. Missing or ambiguous excerpts are rejected. Complete record citations remain unchanged. When quotes from the same address are already in notes, the short-fact batch can also ask whether the full page adds relevant evidence. Literal diffs against matching source blocks help distinguish changed values from cosmetic changes. A negative answer with at least 80% probability skips the read without marking any requirement evidenced. Missing or uncertain answers use the normal readers. This check is omitted when the full comparison does not fit the input budget or the read needs pagination, continuation or incomplete-comparison context. For a count, total or superlative, the reader cites every compared record on every page. Its conclusion lists those facts in `draws_on`, using evidence ids from collected notes or `claim:N` for earlier claims in the same response, indexed from zero. Code resolves these references and drops unknown ones with a debug log. `Fact.basis` keeps the resolved ids, including when a span is reused. A conclusion the page does not state (a count it never prints) cites no blocks: it is a derived fact with no evidence of its own, keyed `derived:`, and it is judged and linked through its basis records. Answer claims use numbered Markdown links built from those notes. `RunResult.citations` exposes the cited facts and text-fragment deep links; unused facts have no citation. An answer citing an unknown reference fails the claim check, and the run falls back to an answer drafted from verified facts. Jev checks the answer's claims against their quotes before completion; a derived fact contributes its basis records, never its own conclusion. Both drafted and composed claims include the transitive basis of each cited fact, deduplicated in read order. The claim check, numbered links and citation records all use those expanded ids. ## Prompt limits and progress `Config.observation` names the limits for viewport text, working notes and recent and earlier history. `Config.tokens` names the input budgets and reader/composer output limits. Working notes can omit facts with a count; verdict prompts retain all requirement evidence or stop at `observation_limit`. The notes budget keeps a retained fact's basis with it; requirement evidence includes its transitive basis. The Jev completion check reduces page text first to make room for that evidence. Cut page text carries a marker when there is room for one; an excerpt too small to carry the marker is empty. A money or time limit keeps collected facts and their citations in a partial answer. The status stays `budget_exceeded`, and producing the partial answer makes no additional model calls. A time budget stops active work; the return also waits for owned requests and browser cleanup. Injected clients must honor cancellation and finish their cleanup. Returning while a paid request still runs would lose its cost receipt. Only visible effects or added evidence count as progress. Rewriting the value already in the observed field cannot count, even when it opens an autocomplete popup. `StallRules` checks lack of progress, repeated interactions and consecutive unproductive steps with the same unresolved requirements. The latter two are recorded in `RunResult.would_fire` in `shadow` mode, the default; `armed` mode sends them to recovery. A productive step clears the plan-stagnation streak, and recovery resets the evidence used by all three checks. ## Run events and images `on_event` receives a `BrowserEvent` followed by `StepEvent` objects. Each step includes added facts, their reader and source links in `StepResult.facts`, and any available explanation in `StepResult.note`. `Config(step_frames=True)` adds a PNG after each step unless a fresh observation shows a resolved secret. `on_frame` receives JPEG bytes from the active tab. Frames are acknowledged after the async handler returns; there is no fixed frame rate, and only the latest pending frame is retained. Delivery runs separately from the agent, and handler failures are logged. Live images and MP4 recordings are held back from the moment a secret is typed, and whenever the page is read showing one, until a reading shows none; a recording holds its last clean frame meanwhile. ## Cursor feedback `--cursor` (`run_task(cursor=True)`) shows where the agent is about to act, through the Cua Driver's synthetic cursor. The click itself is still sent over CDP, so the overlay cannot change what an action did and is never retried. The page hook is in `CdpPage.act`, where the hit-tested point is known; a secret's field is never marked. `browser/cursor.py` owns one `cua-driver mcp` child per run, speaking JSON-RPC over stdio, and stops it when the run ends, including on cancellation. The driver's `move_cursor` is sent in `window` scope only, so the real pointer and focus never move. The page point becomes a screen point from what the page reports: `screenX/Y`, `outerWidth/Height`, `innerWidth/Height` and the zoom from `Page.getLayoutMetrics`. The top inset is `outerHeight - innerHeight * zoom`. The cursor is drawn only when exactly one driver window matches the page's reported frame (which also shows the two share a scale, so a scaled display is skipped), that window is in front of every other, the tab is visible and the point is inside the viewport. Otherwise it is hidden. Any driver failure turns the feature off for the run. Limits: Linux X11 only; native Wayland, macOS and Windows are not verified. It needs `cua-driver` 0.28.3 or newer with the session cursor tools. Cloud and headless browsers are skipped. The driver offers no click pulse through `move_cursor`, so only the glide is shown. A docked DevTools panel or a pinch zoom hides the cursor, and a window partly covered by another counts as covered. The overlay is a separate window, so CDP screenshots and recordings do not show it. ## Browser capabilities over plain CDP Verified with `cdp-use==1.4.5` against local headless Chrome and a Browser Use cloud browser: | Capability | Mechanism | Local | Cloud | |---|---|---|---| | Own tab rendered | `Target.createTarget` + `Target.activateTarget` | pass | pass | | Upload caller bytes | in-page `DataTransfer` + `File` on the input, `input`/`change` events (no host path needed) | pass | pass | | Download bytes | `Fetch.enable` at Response stage for Document responses (download-attribute anchors included), `Fetch.getResponseBody` on `Content-Disposition: attachment` | pass | pass | | Cross-origin iframe | `Target.setAutoAttach(flatten)` on the page session, evaluate in the iframe session | pass | pass | | Popup ownership | `Target.targetCreated.openerId` equals our target | pass | pass | | Dialogs | `Page.javascriptDialogOpening` + `Page.handleJavaScriptDialog` | pass | pass | The cloud browser ignores `Browser.setDownloadBehavior(deny)`, so bytes come from response interception, never from the remote filesystem. Host-path `DOM.setFileInputFiles` is only valid for a browser on the same machine. ### Attached windows A session normally opens its own tab and closes only the tabs it owns. With `BrowserConnection.attach` it opens none: it takes the first page target whose title or URL contains `target_match` (or the first page, skipping `devtools://` windows), and owns nothing, so `closeTarget` is never sent. That page loaded before the session's new-document scripts were registered, so they are also evaluated once in it; otherwise freshness tracking would start only at its next navigation. Turning target discovery on replays `targetCreated` for every window already open, so an attached session turns it on only after recording those, and none of them joins the run. Popups of a tracked window join without being owned. Windows with no opener join only when `target_match` is set: an Electron main process opens its windows that way, and so does a browser for each tab a person opens by hand. `allowed_origins` scopes documents with the same Fetch hook downloads use. A scoped session also pauses every `Document` request before it is sent and fails one outside the grant with `ERR_ABORTED`, which commits no page that could name the address. Redirect hops, iframes and popups all pause as documents, so one check covers each. Pages and child targets attach paused (`waitForDebuggerOnStart`) and resume only once their Fetch is on, so a popup or out-of-process frame cannot send its first request unchecked. Reads check the address each text was read from, the committed address of the active tab, and nested document addresses, so an opaque document (`about:blank`, `srcdoc`, `blob:` or `data:`) or a foreign page restored from the back-forward cache is refused rather than read. Nested blank documents remain refused even when `document.write` gives them their parent's address. A window already open outside the grant is never matched, and no window is navigated to enforce anything. Service workers are bypassed on every scoped session (`Network.setBypassServiceWorker`, undone on an attached window at close), so no navigation is answered without meeting the gate. A scoped run delivers no live frames, screenshots or recordings, because pixels cannot be matched to a document. The grant covers documents, not network egress: subresources and requests a granted page makes are not scoped, and neither are workers. `check_access`, when given, is awaited before browser startup and every browser read and action, scoped or not. Before a pointer press, the browser waits for a stable target and rechecks its guard and hit-test. A replacement control must match the original semantics and receiving document; ambiguous matches are refused. Focus and the receiving field are checked before typing. Settling waits for an interactive document and a quiet DOM, with a bounded extra wait for visible loading indicators. ---