文件历史

提交图

13 次代码提交

作者 SHA1 备注 提交日期
Saurav Panda 67228da9a8 Merge pull request #50 from browser-use/feat/domain-skills-batch5
Add domain skills: SEC EDGAR, Coursera, Goodreads, DuckDuckGo, TrustPilot
2026-04-18 13:35:13 -07:00
sauravpanda 993323457d Add browser-harness validated domain skills for batch 10
Wellfound (DataDome + Cloudflare dual stack; browser CDP resolves silently; Rails not Next.js),
G2 (DataDome blocks all http_get; browser CDP; schema.org microdata stable; data.g2.com API needs vendor key),
Capterra (ClaudeBot UA returns Markdown not HTML; Chrome UA gets Cloudflare 403),
TradingView (scanner.tradingview.com POST open; symbol-search needs Origin header; data.tradingview.com dead),
Macrotrends (iframe PHP endpoints bypass main page; stock OHLCV direct; Referer required for economic API).
2026-04-18 05:18:38 -07:00
sauravpanda 33dcd86207 Add browser-harness validated domain skills for batch 9
CoinMarketCap (internal data-api/v3 fully open, no auth; 25 calls no rate limit),
Quora (full Chrome UA required; push() payloads double-encoded JSON; 3 SSR answers only),
Itch.io (http_get works; game cards via CSS selectors; RSS feeds exist),
Steam (appdetails single appid only; price in cents; ISteamApps/GetAppList dead in 2026),
HowLongToBeat (two-step token flow /api/find/init then POST; comp_* in seconds not hours).
2026-04-18 04:59:15 -07:00
sauravpanda cfa55ed15e Add browser-harness validated domain skills for batch 8
Letterboxd (http_get works on film pages; JSON-LD CDATA gotcha; API needs OAuth),
Gutenberg (Gutendex REST API; text via /cache/epub/; .opf is 404, use .rdf),
Metacritic (internal backend API key in HTML; Nuxt __NUXT_DATA__ not __NEXT_DATA__),
RAWG (API needs key; window.CLIENT_PARAMS in HTML has full game data without key),
OpenLibrary (full free API; missing cover = 43-byte GIF not 404; description dual type).
2026-04-18 04:46:39 -07:00
sauravpanda c2167e544e Add browser-harness validated domain skills for batch 7
Glassdoor (Cloudflare managed challenge; browser only; __NEXT_DATA__ + DOM fallbacks),
Medium (?format=json strips XSSI prefix; GraphQL /_/graphql no auth; RSS 10-item cap),
SoundCloud (oEmbed no-auth; __sc_hydration apiClient.id as client_id; API v2 with pagination),
Genius (OS token in UA bypasses 403; internal /api/songs no auth; strip first lyrics div header),
Dev.to (public REST API; burst limit 6 req then 429/1s; listings empty without auth).
2026-04-18 04:31:04 -07:00
sauravpanda f897a29d17 Add browser-harness validated domain skills for batch 6
Walmart (http_get works; bare Mozilla/5.0 UA bypasses PerimeterX; __NEXT_DATA__ via id= regex not type=),
Spotify (oEmbed no-auth; embed __NEXT_DATA__ has full trackList; anonymous token burns 22h Retry-After),
Weather (wttr.in format=j1; Open-Meteo current+hourly+daily; NWS two-call flow),
OpenStreetMap (Nominatim lat/lon are strings; Overpass bbox order differs from Nominatim; POST required),
Archive.org (Wayback availability API degraded; CDX sort=closest reliable; /search?output=json returns HTML).
2026-04-18 04:16:54 -07:00
sauravpanda 40291337ce Add Goodreads domain skill (http_get works; __NEXT_DATA__ Apollo state + JSON-LD) 2026-04-18 03:52:34 -07:00
sauravpanda a034cd0021 Add browser-harness validated domain skills for batch 5 (partial)
Coursera (public API no auth, q=search is POST-only/405 on GET),
DuckDuckGo (Instant Answer API, skip_disambig=1 essential, widget answers unusable),
SEC EDGAR (company UA required for www.sec.gov; 10 req/s; XBRL frames for cross-company),
TrustPilot (http_get works; __NEXT_DATA__ has reviews; 10-page cap per filter).
2026-04-18 03:50:39 -07:00
sauravpanda 4f05481d2a Replace drafted skills with browser-harness validated versions
Ran actual browser-harness sessions against each site and rewrote
the skill files from live test findings. Key corrections:

GitHub: wait(2) after wait_for_load() for React hydration; search API
  separate 10 req/min limit; search/code needs auth (401 unauthed)

HackerNews: athing also matches comment rows (use 'athing submission');
  job posts break naive score-zip; html.unescape() required for titles

Amazon: .zg-item-immersion gone from Best Sellers; #priceblock_ourprice
  returns null (legacy); review count selector collides with cross-sell
  widget — use [aria-label*='ratings'] instead

News: The Verge is Atom not RSS (namespace dict required); Reuters
  hard-blocks http_get with 403 even with User-Agent; BBC shows no
  consent banner from US IP; parallel fetch is 4.3x faster (0.16s vs 0.70s)

ProductHunt: goto() ERR_ABORTED — always use new_tab(); /posts/ URLs
  don't exist (it's /products/); homepage has 30 fixed items no lazy load;
  [data-test^='post-item-'] is the correct card selector
2026-04-18 01:22:30 -07:00
sauravpanda f0f70e7362 Add domain skills for 6 public, no-login-required sites
- github/scraping.md: GitHub API + trending page patterns, rate limiting
- hackernews/scraping.md: http_get + Algolia API for search/filtering
- producthunt/scraping.md: React SPA browser scraping + GraphQL token approach
- amazon/product-search.md: ASIN extraction, price parsing, CAPTCHA handling
- news-aggregation/multi-source.md: RSS-first approach, parallel fetch, consent banners
- job-boards/indeed-glassdoor.md: URL construction, job key extraction, salary normalization

All skills derived from real user task patterns. No credentials or sensitive data included.
2026-04-18 00:16:18 -07:00
Magnus Müller a1c35ba384 Add thetechgeeks domain-skill for Ubiquiti AU pricing (#40)
Captures three learnings from a real run that mis-reported $3,080 AUD for
a UACC-Rack-12U-Wall (real AU street ~$420-630): Tech Geeks is Shopify so
use /products/<handle>.js for canonical price and SKU; its .js `available`
flag is unreliable so cross-check the DOM for sold-out markers; and
sold-out pages there carry stale/junk prices that must never enter a
final table without a second-source sanity check.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-17 23:17:52 -07:00
Gregor Žunič 29c4dfa8e7 changed name 2026-04-17 20:50:33 -07:00
MagMueller d6942b61fe add interaction-skills and domain-skills structure
Two skill categories, pure markdown, no Python files:

- interaction-skills/ — generic browser patterns (dialogs, inputs, etc.)
  Flat .md files. Agent reads the relevant one before a task.

- domain-skills/ — per-site playbooks (tiktok/, linkedin/, etc.)
  Flat .md files per action (upload, schedule, post).
  Subfolders for large domains when needed.

Starting with:
- interaction-skills/dialogs.md — CDP vs JS dialog handling
- domain-skills/tiktok/upload.md — full upload flow with gotchas
- Placeholder folders for linkedin, spreadshirt, salesforce

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-17 13:46:37 -07:00