We extract verified business listings, NAP records, categorisation tags, and reviews from Brownbook. Delivered as clean JSON, CSV, or Parquet to S3 or BigQuery on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Business Profiles objects from brownbook.net. All fields typed and schema-versioned.
"business_id": "bb_8472910", "name": "Apex Industrial Supplies", "category": "Manufacturing & Industry", "claimed_status": true, "date_added": "2021-04-12", "profile_url": "https://www.brownbook.net/business/8472910/apex-industrial-supplies"
| # | business_id | name | category | description | claimed_status | date_added |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for NAP & Location objects from brownbook.net. All fields typed and schema-versioned.
"business_id": "bb_8472910", "street": "142 Industrial Parkway", "city": "Birmingham", "state": "West Midlands", "postal_code": "B1 1AA", "country": "United Kingdom", "formatted_address": "142 Industrial Parkway, Birmingham, West Midlands, B1 1AA, United Kingdom"
| # | business_id | street | city | state | postal_code | country |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Operating Hours objects from brownbook.net. All fields typed and schema-versioned.
"business_id": "bb_8472910", "monday": "08:00 - 17:00", "tuesday": "08:00 - 17:00", "wednesday": "08:00 - 17:00", "thursday": "08:00 - 17:00", "friday": "08:00 - 16:00", "saturday": "Closed", "sunday": "Closed", "timezone": "Europe/London"
| # | business_id | monday | tuesday | wednesday | thursday | friday |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Reviews & Ratings objects from brownbook.net. All fields typed and schema-versioned.
"review_id": "rev_99281", "business_id": "bb_8472910", "reviewer_name": "James T.", "rating": 4.5, "review_text": "Reliable supplier for heavy machinery parts. Fast shipping.", "review_date": "2023-11-04"
| # | review_id | business_id | reviewer_name | rating | review_text | review_date |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Social & Media objects from brownbook.net. All fields typed and schema-versioned.
"business_id": "bb_8472910", "facebook_url": "https://facebook.com/apexindustrial", "linkedin_url": "https://linkedin.com/company/apex-industrial", "logo_url": "https://www.brownbook.net/images/logos/8472910.jpg", "photo_count": 4, "photo_urls": "['https://www.brownbook.net/images/photos/8472910_1.jpg', 'https://www.brownbook.net/images/photos/8472910_2.jpg']"
| # | business_id | facebook_url | twitter_url | linkedin_url | logo_url | photo_count |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our Brownbook scraper targets the core directory structure: business listings, contact information, category taxonomy, and review data. We handle pagination, geographic filtering, and stale record normalisation.
Extract company name, description, claimed status, and date added across millions of global directory listings.
Capture and structure Name, Address, and Phone number data into clean, queryable fields. We parse raw text into street, city, state, and postal code.
Extract primary and secondary business categories to map listings against your internal industry classification systems.
Pull reviewer names, star ratings, text content, and publication dates for sentiment analysis and reputation monitoring.
Structure weekly opening and closing times, including weekend variations and timezone alignment.
Extract linked Facebook, Twitter, and LinkedIn profiles associated with the business listing.
Target specific countries, regions, or postal codes to build localised business datasets.
Run continuous pipelines that only emit records when a business updates its address, phone number, or operating hours.
Capture logo URLs, photo counts, and image links to enrich POI databases.
Brief in. Clean data out.
Provide target countries, categories, or specific search queries. We design the extraction schema together.
We configure Scrapy crawlers, proxy rotation, and pagination logic to traverse Brownbook's directory structure.
Schema validation, null-rate checks, and NAP normalisation testing before full launch.
JSON / CSV / Parquet pushed to your S3 bucket or BigQuery dataset on agreed cadence.
Directory scraping involves deep pagination and unstructured data. Here is how we ensure data quality.
Brownbook organises data through deep category and location pagination. Our crawlers map the entire site taxonomy, ensuring no sub-category or regional listing is orphaned during the extraction process.
User-submitted directories often contain messy address formats. We apply regex patterns and location dictionaries to parse raw strings into structured street, city, state, and postal code fields.
Directories accumulate dead listings. We extract 'date added', claim status, and review recency to help you filter out stale or abandoned business profiles from your final dataset.
While Brownbook's bot protection is standard, aggressive scraping triggers IP bans. We distribute requests across rotating proxy pools and respect optimal concurrency limits to maintain 99.98% pipeline uptime.
We utilise multiple fallback chains for field extraction. If a listing lacks a standard address block, our fallback selectors parse the description or metadata tags to recover the missing information.
Agencies audit Brownbook listings to ensure NAP consistency for their clients across global directories.
Sales teams extract contact details and category tags to build targeted outreach lists by industry and region.
Analysts track business density by category and postal code to identify underserved markets.
Mapping and navigation companies cross-reference Brownbook data to validate existing Point of Interest databases.
Firms monitor new business registrations and category growth to identify macroeconomic trends at the local level.
Brands monitor user reviews and ratings across directory sites to manage public perception and respond to feedback.
"Brownbook offers a massive, crowdsourced global directory, but extracting clean, structured NAP data requires navigating deep pagination and inconsistent formatting."
Most teams struggle with directory scraping due to unstructured address fields, deep pagination walls, and stale records. DataFlirt normalises every NAP record, validates geographic coordinates, and handles all orchestration so your engineers receive clean, queryable data.
Everything supported by our brownbook.net scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles high-concurrency crawl orchestration, URL deduplication, and retry logic for traversing deep directory pagination.
We maintain pools of rotating proxies to distribute request load and prevent rate limiting during large-scale directory extractions.
Pipelines run on Kubernetes clusters. Airflow handles scheduling and dependency management. All state is stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About brownbook.net scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available business listings from directories like Brownbook is generally permissible. DataFlirt extracts only public, non-authenticated NAP data, categories, and reviews. We do not extract private user account data or bypass authentication walls.
Directory data is often user-submitted and unstructured. We apply custom parsing logic, regex patterns, and address normalisation libraries to split raw text into clean street, city, state, and postal code fields.
Yes. We can configure the pipeline to target specific geographic regions, postal codes, or industry categories based on your exact requirements.
We can run full catalogue refreshes on a weekly or monthly cadence. For specific target lists, we can configure daily runs to detect changes in operating hours or newly added reviews.
Yes. We capture Facebook, Twitter, and LinkedIn URLs associated with the business profile, along with website links and logo image URLs.
Yes. We provide a sample run of up to 1,000 listings for your target category or region during the scoping phase, allowing you to validate data structure and completeness.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a targeted list of local businesses or a global directory dump across millions of records - we scope, build, and operate the pipeline. Tell us what you need.