SYSTEM all green source superpages.com queue 12,948 pages p99 latency 185ms dataflirt.com · scraper/superpages-com
RUN * 114 active pipelines * superpages.com live

Superpages data,
at warehouse scale.

We extract business profiles, contact details, operating hours, ratings, and reviews from Superpages. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Businesses extracted
485K /day
Reviews processed
1.2M /week
Phone numbers
312K /run
Active pipelines
114
Uptime
99.94%
Data Dictionary

Every field we extract from superpages.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Business Profiles objects from superpages.com. All fields typed and schema-versioned.

business_idnameprimary_categoryaddresscitystatezip_codephonewebsiteclaim_status
business_profiles
● 200 OK
"business_id": "sp-98234",
"name": "Joes Plumbing",
"primary_category": "Plumbers",
"address": "123 Main St",
"city": "Austin",
"state": "TX",
"zip_code": "78701",
"phone": "512-555-0199"
# business_idnameprimary_categoryaddresscitystate
1
2
3

Complete list of extractable fields for Reviews & Ratings objects from superpages.com. All fields typed and schema-versioned.

review_idbusiness_idratingreview_textauthor_namereview_datesourcehelpful_votes
reviews_& ratings
● 200 OK
"review_id": "rev-456",
"business_id": "sp-98234",
"rating": 5,
"review_text": "Great service and fast response.",
"author_name": "Sarah Connor",
"review_date": "2023-10-12",
"source": "Superpages",
"helpful_votes": 12
# review_idbusiness_idratingreview_textauthor_namereview_date
1
2
3

Complete list of extractable fields for Operating Hours objects from superpages.com. All fields typed and schema-versioned.

business_idday_of_weekopen_timeclose_timeis_closedis_holidaytimezonelast_verified
operating_hours
● 200 OK
"business_id": "sp-98234",
"day_of_week": "Monday",
"open_time": "08:00",
"close_time": "18:00",
"is_closed": false,
"timezone": "America/Chicago"
# business_idday_of_weekopen_timeclose_timeis_closedis_holiday
1
2
3

Complete list of extractable fields for Geolocation objects from superpages.com. All fields typed and schema-versioned.

business_idlatitudelongitudeneighborhoodmap_urldirections_urlservice_areaaccuracy
geolocation
● 200 OK
"business_id": "sp-98234",
"latitude": 30.2672,
"longitude": -97.7431,
"neighborhood": "Downtown",
"map_url": "https://maps.example.com",
"accuracy": "rooftop"
# business_idlatitudelongitudeneighborhoodmap_urldirections_url
1
2
3

Complete list of extractable fields for Search Results objects from superpages.com. All fields typed and schema-versioned.

keywordlocationpositionbusiness_idnamesponsoredratingreview_countlisting_url
search_results
● 200 OK
"keyword": "plumber",
"location": "Austin, TX",
"position": 1,
"business_id": "sp-98234",
"name": "Joes Plumbing",
"sponsored": true,
"rating": 4.8,
"review_count": 45
# keywordlocationpositionbusiness_idnamesponsored
1
2
3

Capabilities

Everything you need from Superpages: nothing you do not

Our Superpages scraper handles every layer of the directory: business listings, dynamic search results, category tracking, and the review corpus with JavaScript rendering and anti-bot circumvention built in.

Full Business Profile Extraction

Business name, address, phone number, website, claim status, and every metadata field Superpages surfaces scraped at the listing level.

Contact Information Mining

Capture primary and secondary phone numbers, email addresses where available, and external website links.

Category & Taxonomy Mapping

Extract primary and secondary business categories. Track listings across multiple service taxonomies.

Review & Rating Aggregation

Full review text, star ratings, helpful vote counts, and review dates paginated across all review pages.

Geolocation & Service Areas

Latitude, longitude, neighborhood identifiers, and defined service areas for local businesses.

Search Result Scraping

Track organic vs sponsored position for any keyword and location combination.

Operating Hours Normalisation

Extract and normalise daily operating hours, holiday exceptions, and timezone data.

Sponsored Listing Detection

Distinguish between paid advertisements and organic local search results.

Scheduled & Streaming Modes

Run one-off bulk exports or configure continuous pipelines at hourly, daily, or weekly cadences.

// engagement pipeline

From location lists to warehouse records

Brief in. Clean data out.

Define Scope
d 0

Provide locations, business categories, or keyword sets. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy and Playwright crawlers, proxy rotation, session management, and CAPTCHA handling for superpages.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, and sample data reviews before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Superpages pipeline handles the hard parts

Superpages employs scraping detection for high-volume requests. Here is how we stay resilient and why teams choose managed infrastructure over DIY.

pipeline-monitor · superpages.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
Residential proxy rotation and fingerprint spoofing

Superpages bot detection operates on IP reputation and request volume. Our crawlers use residential ISP proxies with realistic browser fingerprints and randomised request timing.

JavaScript rendering
Full Playwright execution for dynamic content

Superpages search results and contact reveals often require JavaScript execution. We run full Playwright browser sessions to trigger lazy-loads and reveal hidden contact data.

Schema stability
Resilient selectors with fallback chains

Directory structures change. Our selector strategy uses multiple fallback chains per field so a layout change does not break your data pipeline overnight.

Pagination handling
Deep crawl execution

We navigate complex category and location pagination trees to ensure comprehensive data capture across thousands of local search result pages.

Monitoring & alerting
24/7 pipeline health monitoring

Every run emits structured logs to our observability stack. We alert on null-rate spikes and coverage drops. SLA uptime is contractual.

Applications

Who uses Superpages data and how

Teams across industries use superpages.com data to build competitive products and smarter operations.

01
Local SEO Monitoring

Agencies track NAP consistency, citation building, and competitor rankings across local directories.

02
Lead Generation

B2B sales teams extract verified local business contact details to build targeted outreach lists.

03
Market Research

Analysts track business density, category growth, and regional service availability to identify market opportunities.

04
Data Enrichment

Data providers append Superpages operating hours, ratings, and category data to their existing business records.

05
Competitor Analysis

Franchises monitor local competitor presence, review sentiment, and sponsored ad placements.

06
AI Training Data

Machine learning teams use structured business profile datasets to train location-based recommendation engines.

Why DataFlirt

"Superpages contains millions of verified local business records, but extracting accurate NAP data across thousands of categories requires a resilient infrastructure."

Most teams underestimate the investment required: reliable Superpages scraping requires residential proxies, full JavaScript rendering, CAPTCHA handling, daily selector maintenance, and anomaly monitoring. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.

Technical Spec

Superpages scraper: technical capabilities

Everything supported by our superpages.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for dynamic phone number reveals and map data
Supported
CAPTCHA bypass
Automated 2Captcha and CapSolver integration with fallback to manual queue
Supported
Residential proxy rotation
ISP-grade residential IPs from US pools rotated per request
Supported
Category pagination
Deep extraction across all sub-categories and location modifiers
Supported
Review extraction
Full review corpus including pagination across all user reviews
Supported
Sponsored ad detection
Distinguishes organic vs sponsored placements in search results
Supported
Change detection
Hash-based diff: only emit records with changed fields since last run
Supported
Webhook delivery
HTTP POST per record or batch for downstream processing
Supported
User account data
Extraction of private user profiles or saved business lists
Partial
Private messaging extraction
Access to direct messages between users and businesses
Partial
Infrastructure

Infrastructure powering the Superpages pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy & Playwright Stack

Scrapy handles crawl orchestration and retry logic. Playwright handles JavaScript rendering and interaction flows for dynamic contact reveals.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across US regions. Rotation happens per-request with sticky sessions where required.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling and dependency management. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested schema versioned per run
CSV
Flat file with typed columns Excel compatible
XLS
Standard spreadsheet format for business teams
Parquet
Columnar format for BigQuery and Snowflake
AWS S3
Direct bucket delivery compatible with any data lake
Webhook
HTTP POST per record for immediate downstream processing
API
RESTful endpoints for on-demand data retrieval
PostgreSQL
Upsert into your existing schema with conflict resolution
Snowflake
Stage and COPY INTO workflow
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About superpages.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Superpages legal?

Scraping publicly available information from Superpages is generally permissible under applicable law. DataFlirt targets only public, non-authenticated business profile and review data. We do not extract personal user data or circumvent authentication walls.

How do you handle Superpages bot protection?

We use residential ISP proxies, full Playwright browser sessions with realistic fingerprints, and request timing modelled on human behaviour. We monitor for rate spikes in real time and trigger pool rotation automatically.

Can you extract hidden phone numbers?

Yes. Superpages often masks phone numbers behind JavaScript click events. Our Playwright integration executes these events to capture the underlying contact information accurately.

How fresh is the data?

Full category refreshes at weekly or monthly cadences complete within a defined execution window. Real-time extraction for specific keyword searches can be configured for sub-60-minute latency.

Do you support review scraping?

Yes. Each review record includes rating, text, author name, review date, and helpful votes paginated across all available review pages for a listing.

What is the minimum viable engagement?

Our smallest packages start at a defined category or location list with weekly delivery. For larger national directories, we price based on volume and delivery frequency.

Can I request a sample dataset?

Absolutely. We provide a sample run of up to 500 business listings as part of the pre-engagement scoping process so you can validate schema fit and data quality.

$ dataflirt scope --new-project --source=superpages.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off directory dump or a continuous local SEO monitoring feed, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in business directories

Services

Data Extraction for Every Industry

View All Services →