Scraping APIs for AI agents
Read, crawl and extract from pages and sites. 37 calls from 10 providers, each priced before it runs; provider errors are not charged. Not sure which to pick? vaaya/onescrape routes across them in one call.
npx @vaaya/mcp installBright Data
Diffbot
diffbot/analyze1¢Diffbot — Extract STRUCTURED, typed data from a URL: it classifies the page (article / product / discussion / image) and returns parsed fields — title…
diffbot/analyze_html1¢Diffbot — Same structured extraction as diffbot/analyze, but over HTML YOU already fetched rather than a URL Diffbot fetches itself. This is the pairing for…
FastCRW
crw/crawl10¢CRW — Start an ASYNC multi-page crawl from a seed URL, following links. Pass `url`; optional `maxPages` (1-100, default 100 — vendor bills 1 credit/page so…
crw/crawl_status1¢CRW — Poll an async crawl started by crw/crawl. Pass `id` (from the crawl response). Returns `{ status: scraping|completed|failed|cancelled, completed, total…
crw/extract5¢CRW — Structured extraction over up to 10 URLs using an LLM. Pass `urls` plus `prompt` (natural language) and/or `schema` (JSON schema); optional `basis…
crw/extract_status1¢CRW — Poll an async extraction started by crw/extract, on the rare occasions it returns an `id` instead of inline results. Pass `id`. Returns `{ status…
crw/map1¢CRW — Discover the URLs of a website without scraping content (sitemap + crawl fallback). Pass `url`; optional `maxDepth`, `useSitemap` (default true)…
crw/scrape1¢CRW — Scrape a single URL to clean markdown/HTML/JSON (Firecrawl-compatible). Pass `url`; optional `formats`…
crw/search1¢CRW — Search the web and optionally scrape the hits in one call. Pass `query`; optional `limit` (1-20, default 5), `tbs` (freshness: qdr:h|d|w|m|y), `sources`…
Firecrawl
firecrawl/crawl1¢Firecrawl — Crawl a website starting from a URL, following links.
firecrawl/extract1¢Firecrawl — Extract structured data from URLs using a schema.
firecrawl/map1¢Firecrawl — Map all URLs on a website without scraping content.
firecrawl/scrape1¢Firecrawl — Scrape a single URL and return clean markdown/HTML.
firecrawl/search1¢Firecrawl — Search the web and return scraped results.
Jina
jina/read1¢Jina Reader — fetch a URL and return LLM-ready markdown (r.jina.ai). Pass `url`. Handles JS rendering and boilerplate stripping automatically; returns `{…
jina/search1¢Jina Search — web search that returns the top hits WITH their full reader-processed page content in one call (s.jina.ai). Pass `q`; optional `num` (result…
Olostep
olostep/ai-visibilityvaries by paramsOlostep — ask AI search engines the prompts your customers ask, and get back what each one actually answered: the answer text, the sources it cited, and any…
olostep/answer5¢Olostep Answers — ask a question or hand it a data point to enrich; it searches and reads live pages, validates, and answers with sources in 3-30s. Pass…
olostep/batchvaries by paramsOlostep — scrape up to 100 URLs as one ASYNC batch (1-3 min for small batches, 5-8 min for large). Pass `items` ([{ url, custom_id? }]; custom_id defaults to…
olostep/batch-itemsfreeOlostep — the items of a completed batch, WITH content inline. Pass `id` ("batch_…"); optional `formats` (default ["markdown"]; use ["json"] for parser output…
olostep/batch-statusfreeOlostep — poll a batch started by olostep/batch or olostep/ai-visibility. Pass `id` ("batch_…"). Returns `{ status: in_progress|completed, total_urls…
olostep/crawlvaries by paramsOlostep — start an ASYNC crawl from `start_url`, following links. Required `max_pages` (1-100; it sets the price). Optional `include_urls` / `exclude_urls`…
olostep/crawl-pagesfreeOlostep — the pages of a crawl, WITH their content inline. Pass `id` ("crawl_…"); optional `formats` (default ["markdown"]; add "html" for source), `limit`…
olostep/crawl-statusfreeOlostep — poll a crawl started by olostep/crawl. Pass `id` ("crawl_…"). Returns `{ status: in_progress|completed, pages_count, credits_consumed }`. Free. Poll…
olostep/map1¢Olostep — list the URLs of a website without scraping them. Pass `url`; optional `search_query` + `top_n` (rank and keep the most relevant), `include_urls` /…
olostep/retrievefreeOlostep — fetch one stored page by `retrieve_id` (from olostep/scrape, crawl-pages or batch-items). Optional `formats` (markdown | html | json). Use it for a…
olostep/scrapevaries by paramsOlostep — scrape one URL (JS rendered, residential IPs) in about a second. Pass `url_to_scrape`; optional `formats` (markdown default | html | text | json |…
olostep/searchvaries by paramsOlostep — natural-language web search returning deduplicated links with title and description. Pass `query`; optional `limit` (1-25, default 12)…
oxylabs
oxylabs/scrapeup to 25¢Oxylabs — scrape a public URL with optional geo-targeting and JS rendering. Params: `url`, optionally `geo_location`, `render` ("html").
Scrape.do
scrapedo/scrape1¢Scrape.do — Fetch a page through a rotating datacenter-proxy pool with anti-bot handling. Surprisingly strong for the price: it returned the REAL page on a…
scrapedo/scrape_super2¢Scrape.do — The heavy rung: RESIDENTIAL/mobile proxy pool plus full JS rendering (`super` + `render`). For pages that bounce the plain scrapedo/scrape call…
scraping
scraping/scrapevaries by paramsScraping category endpoint
ScrapingAnt
scrapingant/extract20¢ScrapingAnt — AI data extraction WITHOUT a schema: describe the fields in plain English and get structured JSON back. Pass `url` and `extract_properties` — a…
scrapingant/markdown1¢ScrapingAnt — Scrape a URL and return LLM-ready markdown (rendered in headless Chrome, then converted). Pass `url`; optional…
scrapingant/scrape1¢ScrapingAnt — Scrape a URL through a managed headless-Chrome cluster (datacenter proxies). Pass `url`; optional `browser` (default true — set false for plain…
scrapingant/scrape_residential4¢ScrapingAnt — Scrape a HARD page through the 3M+ residential-proxy pool + headless Chrome: Cloudflare and anti-bot walls, geo-fenced content, sites that block…
Give your agents scraping.
Connect Vaaya to your agent or application, then call any action on this page on your Vaaya credit.
npx @vaaya/mcp install