Bright Data
Search, Crawl and Scrape any site, at scale, without getting blocked
0.5.1Bright Data is a web data platform for scraping, searching, and extracting structured data from any site at scale without getting blocked. This toolkit exposes Bright Data's proxy and scraping infrastructure through Arcade, enabling search, raw page scraping, and structured data extraction.
Capabilities
- Web scraping: Fetch any public webpage and receive clean Markdown output suitable for LLM consumption or downstream processing.
- Multi-engine search: Query Google, Bing, or Yandex with configurable result counts, country targeting, and search type (web or images).
- Structured data feeds: Extract normalized, schema'd data from major platforms including Amazon (products, reviews), LinkedIn (people, companies), Instagram, Facebook, X, Zillow, Booking.com, YouTube, and ZoomInfo — no scraping logic required.
Secrets
This toolkit requires two secrets to authenticate with Bright Data.
-
BRIGHTDATA_API_KEY— Your Bright Data API key, used to authenticate all requests. Obtain it from the Bright Data dashboard under Account Settings → API Keys. Generate a new key with sufficient permissions for the products you plan to use (scraping, SERP, datasets). A paid or trial Bright Data account is required. -
BRIGHTDATA_ZONE— The name of the Bright Data zone (proxy zone or scraping browser zone) that requests are routed through. Zones are created and named in the Bright Data control panel under Proxies & Scraping Infrastructure → Zones. The zone name must match a zone that is active on your account and has the appropriate product type enabled (e.g., a SERP API zone for search, a Web Unlocker or Scraping Browser zone for scraping).
For guidance on configuring secrets in Arcade, see the Arcade secrets docs. Secrets can also be managed at https://api.arcade.dev/dashboard/auth/secrets.
Available tools(3)
| Tool name | Description | Secrets | |
|---|---|---|---|
Scrape a webpage and return content in Markdown format using Bright Data.
Examples:
scrape_as_markdown("https://example.com") -> "# Example Page
Content..."
scrape_as_markdown("https://news.ycombinator.com") -> "# Hacker News
..."
| 2 | ||
Search using Google, Bing, or Yandex with advanced parameters using Bright Data.
Examples:
search_engine("climate change") -> "# Search Results
## Climate Change - Wikipedia
..."
search_engine("Python tutorials", engine="bing", num_results=5) -> "# Bing Results
..."
search_engine("cats", search_type="images", country_code="us") -> "# Image Results
..."
| 2 | ||
Extract structured data from various websites like LinkedIn, Amazon, Instagram, etc.
NEVER MADE UP LINKS - IF LINKS ARE NEEDED, EXECUTE search_engine FIRST.
Supported source types:
- amazon_product, amazon_product_reviews
- linkedin_person_profile, linkedin_company_profile
- zoominfo_company_profile
- instagram_profiles, instagram_posts, instagram_reels, instagram_comments
- facebook_posts, facebook_marketplace_listings, facebook_company_reviews
- x_posts
- zillow_properties_listing
- booking_hotel_listings
- youtube_videos
Examples:
web_data_feed("amazon_product", "https://amazon.com/dp/B08N5WRWNW")
-> "{"title": "Product Name", ...}"
web_data_feed("linkedin_person_profile", "https://linkedin.com/in/johndoe")
-> "{"name": "John Doe", ...}"
web_data_feed(
"facebook_company_reviews", "https://facebook.com/company", num_of_reviews=50
) -> "[{"review": "...", ...}]" | 2 |