benchmarks
Benchmarks
Each benchmark reports its own score, task completion, time and cost. Test settings are listed alongside the results.
back to fastbrowse.ai·how it works
| Benchmark | Measures | Tasks | Results |
|---|---|---|---|
| fastbrowse-internal | Task success | 54 x 3 | fastbrowse 86.4% |
fastbrowse-internal: method and detailed results
01 / fastbrowse-internal
Fastbrowse on the internal task set
All 54 supported tasks, three repeats each. Competitor results are separate references.
- fastbrowse86.4%
| Measurement | fastbrowse |
|---|---|
| Task success | 86.4% |
| Counted | 140 of 162 runs |
| Results completed | 137 of 162 |
| Browser runs | 162 for 162 scored results |
| Median time per result | 21.2s |
| Median cost per result | $0.0189 |
| Total cost | $7.22 |
Method and limits
- Task success uses the internal source graders. Agent completion is reported separately.
- Time and cost include every browser run, including failed retries.
- Internal live task suite, not an official Browser Use benchmark
- Fastbrowse only; competitor results are separate references
- Concurrency 4; local browser; no fixed task time, step or model-call caps
- Each task uses its worker remaining campaign share; idle shares are not transferred
- Time and cost total every physical run, including failed retries
- Source task success and agent completion are reported separately
- Recorded model for fastbrowse: jev openrouter
- Recorded text models for fastbrowse: google:gemini-3.5-flash-lite, google:gemini-3.8-flash
method and evidence
What these figures come from
Every score, time and cost above is computed from per-result rows. Each agent has exactly one scored result for every task and repeat, with no missing cost or grade. Time and cost include every browser run needed for a result, including failed retries. Full run records, answers and page text are kept private.
- Campaign
- release-0.5.20-3c40e6e
- Agent commit
- 3c40e6e18f9cfe87ed4c98c5c4b2fedc3e168353
- Runner commit
- 64a9e2b40f29269a11d4c6e47e7fcf10d3228dfc
- Task catalog sha256
- 55b5f7dae6a29bb301d519205f7f2ec472bd9f9c572d63a4855c6745f6734b2b
- Fixture check
- 108 of 108 passed at 3c40e6e18f9c, run 38023507680
- Run records
- Checked against source runs
- Data
- benchmarks.json at 4372bfd6bb8f
- Data sha256
- 27037c3b5df19f8c1116ef4593f9e7470670294c8621049f101f3257c5bf3490
- Source sha256, fastbrowse-internal
- 55b5f7dae6a29bb301d519205f7f2ec472bd9f9c572d63a4855c6745f6734b2b
published references
Published benchmark results
These links show results published by other projects. Their tasks, benchmark versions, run limits and scoring methods may differ from ours, so the scores cannot be compared directly.
Browser Use, BU Bench
Browser Use publishes its benchmark scores and chart with the benchmark.
Online-Mind2Web
The benchmark's authors maintain the task set and its leaderboard.