SYSTEM all green source brownbook.net queue 14,291 pages p99 latency 218ms dataflirt.com · scraper/brownbook-net
RUN - 42 active pipelines - brownbook.net live

Brownbook data,
at directory scale.

We extract verified business listings, NAP records, categorisation tags, and reviews from Brownbook. Delivered as clean JSON, CSV, or Parquet to S3 or BigQuery on your cadence.

Listings extracted
1.2M /day
Profile updates
340K /24h
Reviews parsed
89K /run
Active pipelines
42
Uptime
99.98%
Data Dictionary

Every field we extract from brownbook.net

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Business Profiles objects from brownbook.net. All fields typed and schema-versioned.

business_idnamecategorydescriptionclaimed_statusdate_addedprofile_urlphone
business_profiles
● 200 OK
"business_id": "bb_8472910",
"name": "Apex Industrial Supplies",
"category": "Manufacturing & Industry",
"claimed_status": true,
"date_added": "2021-04-12",
"profile_url": "https://www.brownbook.net/business/8472910/apex-industrial-supplies"
# business_idnamecategorydescriptionclaimed_statusdate_added
1
2
3

Complete list of extractable fields for NAP & Location objects from brownbook.net. All fields typed and schema-versioned.

business_idstreetcitystatepostal_codecountrylatitudelongitudeformatted_address
nap_& location
● 200 OK
"business_id": "bb_8472910",
"street": "142 Industrial Parkway",
"city": "Birmingham",
"state": "West Midlands",
"postal_code": "B1 1AA",
"country": "United Kingdom",
"formatted_address": "142 Industrial Parkway, Birmingham, West Midlands, B1 1AA, United Kingdom"
# business_idstreetcitystatepostal_codecountry
1
2
3

Complete list of extractable fields for Operating Hours objects from brownbook.net. All fields typed and schema-versioned.

business_idmondaytuesdaywednesdaythursdayfridaysaturdaysundaytimezone
operating_hours
● 200 OK
"business_id": "bb_8472910",
"monday": "08:00 - 17:00",
"tuesday": "08:00 - 17:00",
"wednesday": "08:00 - 17:00",
"thursday": "08:00 - 17:00",
"friday": "08:00 - 16:00",
"saturday": "Closed",
"sunday": "Closed",
"timezone": "Europe/London"
# business_idmondaytuesdaywednesdaythursdayfriday
1
2
3

Complete list of extractable fields for Reviews & Ratings objects from brownbook.net. All fields typed and schema-versioned.

review_idbusiness_idreviewer_nameratingreview_textreview_datehelpful_votesreply_text
reviews_& ratings
● 200 OK
"review_id": "rev_99281",
"business_id": "bb_8472910",
"reviewer_name": "James T.",
"rating": 4.5,
"review_text": "Reliable supplier for heavy machinery parts. Fast shipping.",
"review_date": "2023-11-04"
# review_idbusiness_idreviewer_nameratingreview_textreview_date
1
2
3

Complete list of extractable fields for Social & Media objects from brownbook.net. All fields typed and schema-versioned.

business_idfacebook_urltwitter_urllinkedin_urllogo_urlphoto_countphoto_urlsvideo_url
social_& media
● 200 OK
"business_id": "bb_8472910",
"facebook_url": "https://facebook.com/apexindustrial",
"linkedin_url": "https://linkedin.com/company/apex-industrial",
"logo_url": "https://www.brownbook.net/images/logos/8472910.jpg",
"photo_count": 4,
"photo_urls": "['https://www.brownbook.net/images/photos/8472910_1.jpg', 'https://www.brownbook.net/images/photos/8472910_2.jpg']"
# business_idfacebook_urltwitter_urllinkedin_urllogo_urlphoto_count
1
2
3

Capabilities

Everything you need from Brownbook - nothing you don't

Our Brownbook scraper targets the core directory structure: business listings, contact information, category taxonomy, and review data. We handle pagination, geographic filtering, and stale record normalisation.

Full Business Profiles

Extract company name, description, claimed status, and date added across millions of global directory listings.

NAP Normalisation

Capture and structure Name, Address, and Phone number data into clean, queryable fields. We parse raw text into street, city, state, and postal code.

Category Taxonomy Mapping

Extract primary and secondary business categories to map listings against your internal industry classification systems.

Review & Rating Extraction

Pull reviewer names, star ratings, text content, and publication dates for sentiment analysis and reputation monitoring.

Operating Hours Parsing

Structure weekly opening and closing times, including weekend variations and timezone alignment.

Social Media Discovery

Extract linked Facebook, Twitter, and LinkedIn profiles associated with the business listing.

Geographic Filtering

Target specific countries, regions, or postal codes to build localised business datasets.

Change Detection

Run continuous pipelines that only emit records when a business updates its address, phone number, or operating hours.

Media Metadata

Capture logo URLs, photo counts, and image links to enrich POI databases.

// engagement pipeline

From target region to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target countries, categories, or specific search queries. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, proxy rotation, and pagination logic to traverse Brownbook's directory structure.

Validation & QA
d 4–6

Schema validation, null-rate checks, and NAP normalisation testing before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket or BigQuery dataset on agreed cadence.

Under the hood

How our Brownbook pipeline handles the hard parts

Directory scraping involves deep pagination and unstructured data. Here is how we ensure data quality.

pipeline-monitor · brownbook.net · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Pagination traversal
Deep crawl architecture

Brownbook organises data through deep category and location pagination. Our crawlers map the entire site taxonomy, ensuring no sub-category or regional listing is orphaned during the extraction process.

Data normalisation
Structuring raw NAP text

User-submitted directories often contain messy address formats. We apply regex patterns and location dictionaries to parse raw strings into structured street, city, state, and postal code fields.

Stale record handling
Filtering inactive businesses

Directories accumulate dead listings. We extract 'date added', claim status, and review recency to help you filter out stale or abandoned business profiles from your final dataset.

Anti-bot layer
Rate limiting and proxy rotation

While Brownbook's bot protection is standard, aggressive scraping triggers IP bans. We distribute requests across rotating proxy pools and respect optimal concurrency limits to maintain 99.98% pipeline uptime.

Schema stability
Resilient DOM selectors

We utilise multiple fallback chains for field extraction. If a listing lacks a standard address block, our fallback selectors parse the description or metadata tags to recover the missing information.

Applications

Who uses Brownbook data - and how

Teams across industries use brownbook.net data to build competitive products and smarter operations.

01
Local SEO & Citation Building

Agencies audit Brownbook listings to ensure NAP consistency for their clients across global directories.

02
Lead Generation & B2B Sales

Sales teams extract contact details and category tags to build targeted outreach lists by industry and region.

03
Market Mapping

Analysts track business density by category and postal code to identify underserved markets.

04
POI Database Enrichment

Mapping and navigation companies cross-reference Brownbook data to validate existing Point of Interest databases.

05
Alternative Data for Investors

Firms monitor new business registrations and category growth to identify macroeconomic trends at the local level.

06
Reputation Management

Brands monitor user reviews and ratings across directory sites to manage public perception and respond to feedback.

Why DataFlirt

"Brownbook offers a massive, crowdsourced global directory, but extracting clean, structured NAP data requires navigating deep pagination and inconsistent formatting."

Most teams struggle with directory scraping due to unstructured address fields, deep pagination walls, and stale records. DataFlirt normalises every NAP record, validates geographic coordinates, and handles all orchestration so your engineers receive clean, queryable data.

Technical Spec

Brownbook scraper - technical capabilities

Everything supported by our brownbook.net scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Pagination traversal
Deep crawling across all category and location sub-pages
Supported
NAP normalisation
Parsing raw address strings into structured geographic fields
Supported
Category extraction
Mapping primary and secondary business classifications
Supported
Review parsing
Extracting text, rating, and author data from user reviews
Supported
Change detection
Hash-based diffing to emit only updated listings
Supported
Proxy rotation
Datacenter and residential IP pools to prevent rate limiting
Supported
Webhook delivery
HTTP POST per record for real-time ingestion
Supported
User account credentials
Extracting data from private user dashboards or billing pages
Partial
Direct messaging
Automated sending of messages to business owners via the platform
Partial
Infrastructure

Infrastructure powering the Brownbook pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy Orchestration

Scrapy handles high-concurrency crawl orchestration, URL deduplication, and retry logic for traversing deep directory pagination.

Proxy Infrastructure

We maintain pools of rotating proxies to distribute request load and prevent rate limiting during large-scale directory extractions.

Cloud-Native Delivery

Pipelines run on Kubernetes clusters. Airflow handles scheduling and dependency management. All state is stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested - schema versioned per run
CSV
Flat file with typed columns - Excel/Sheets compatible
XLS
Legacy spreadsheet format for business analysts
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery - compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoint to query your extracted directory dataset
PostgreSQL
Upsert into your existing schema with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About brownbook.net scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Brownbook legal?

Scraping publicly available business listings from directories like Brownbook is generally permissible. DataFlirt extracts only public, non-authenticated NAP data, categories, and reviews. We do not extract private user account data or bypass authentication walls.

How do you handle messy address formats?

Directory data is often user-submitted and unstructured. We apply custom parsing logic, regex patterns, and address normalisation libraries to split raw text into clean street, city, state, and postal code fields.

Can I filter extraction by specific countries or categories?

Yes. We can configure the pipeline to target specific geographic regions, postal codes, or industry categories based on your exact requirements.

How fresh is the directory data?

We can run full catalogue refreshes on a weekly or monthly cadence. For specific target lists, we can configure daily runs to detect changes in operating hours or newly added reviews.

Do you extract social media links?

Yes. We capture Facebook, Twitter, and LinkedIn URLs associated with the business profile, along with website links and logo image URLs.

Can I request a sample dataset?

Yes. We provide a sample run of up to 1,000 listings for your target category or region during the scoping phase, allowing you to validate data structure and completeness.

$ dataflirt scope --new-project --source=brownbook.net ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a targeted list of local businesses or a global directory dump across millions of records - we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in business directories

Services

Data Extraction for Every Industry

View All Services →