SYSTEM all green source paginasamarillas.es queue 12,844 pages p99 latency 318ms dataflirt.com · scraper/paginasamarillas-es
RUN · 37 active pipelines · paginasamarillas.es live

Spanish business data,
at warehouse scale.

We extract company profiles, verified phone numbers, addresses, reviews, and operating hours from paginasamarillas.es. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Businesses extracted
1.2M /month
Contact updates
48K /day
Review records
312K /run
Active pipelines
37
Uptime
99.94%
Data Dictionary

Every field we extract from paginasamarillas.es

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Company Profiles objects from paginasamarillas.es. All fields typed and schema-versioned.

business_idnamelegal_namedescriptionprimary_categorysubcategoriesvat_idyear_establishedwebsite_urllogo_url
company_profiles
● 200 OK
"business_id": "ES-MAD-847291",
"name": "Fontanería Martínez",
"primary_category": "Fontaneros",
"subcategories": "['Instalaciones', 'Urgencias 24h']",
"year_established": 1998,
"website_url": "https://fontaneriamartinez.es",
"logo_url": "https://paginasamarillas.es/logos/847291.jpg"
# business_idnamelegal_namedescriptionprimary_categorysubcategories
1
2
3

Complete list of extractable fields for Contact Information objects from paginasamarillas.es. All fields typed and schema-versioned.

business_idphone_primaryphone_secondarymobilewhatsappemail_addresscontact_personsocial_facebooksocial_linkedinsocial_twitter
contact_information
● 200 OK
"business_id": "ES-MAD-847291",
"phone_primary": "+34 91 555 12 34",
"mobile": "+34 600 123 456",
"whatsapp": "+34 600 123 456",
"email_address": "contacto@fontaneriamartinez.es",
"social_facebook": "facebook.com/fontaneriamartinez",
"social_linkedin": "linkedin.com/company/fontaneriamartinez"
# business_idphone_primaryphone_secondarymobilewhatsappemail_address
1
2
3

Complete list of extractable fields for Location Data objects from paginasamarillas.es. All fields typed and schema-versioned.

business_idstreet_addresscityprovincepostal_coderegionlatitudelongitudeneighbourhoodmap_url
location_data
● 200 OK
"business_id": "ES-MAD-847291",
"street_address": "Calle de Alcalá 142",
"city": "Madrid",
"province": "Madrid",
"postal_code": "28009",
"latitude": 40.4245,
"longitude": -3.6741,
"region": "Comunidad de Madrid"
# business_idstreet_addresscityprovincepostal_coderegion
1
2
3

Complete list of extractable fields for Operating Hours objects from paginasamarillas.es. All fields typed and schema-versioned.

business_idmondaytuesdaywednesdaythursdayfridaysaturdaysundayholiday_exceptionstimezone
operating_hours
● 200 OK
"business_id": "ES-MAD-847291",
"monday": "09:00-18:00",
"tuesday": "09:00-18:00",
"wednesday": "09:00-18:00",
"saturday": "10:00-14:00",
"sunday": "Closed",
"timezone": "Europe/Madrid"
# business_idmondaytuesdaywednesdaythursdayfriday
1
2
3

Complete list of extractable fields for Reviews & Ratings objects from paginasamarillas.es. All fields typed and schema-versioned.

review_idbusiness_idauthor_namerating_scorereview_textreview_dateresponse_textresponse_dateplatform_sourcehelpful_votes
reviews_& ratings
● 200 OK
"review_id": "REV-847291-001",
"business_id": "ES-MAD-847291",
"author_name": "Carlos Ruiz",
"rating_score": 4.5,
"review_text": "Servicio rápido y profesional.",
"review_date": "2025-11-04",
"helpful_votes": 3
# review_idbusiness_idauthor_namerating_scorereview_textreview_date
1
2
3

Capabilities

Extract Spanish directory data with precision

Our paginasamarillas.es scraper handles complex regional blocking, JavaScript-rendered contact details, and diverse listing templates to deliver structured B2B datasets.

Complete Business Profiles

Extract name, description, categories, and core metadata for millions of Spanish SMEs across all provinces.

Contact Detail Extraction

Capture phone numbers, emails, and website URLs directly from directory listings, bypassing JavaScript obfuscation.

Address & Geolocation

Parse structured addresses into street, city, province, postal code, and coordinate data for spatial analysis.

Operating Hours Processing

Normalise unstructured opening hours into standard ISO formats for 7-day schedules.

Review & Rating Aggregation

Collect user ratings, textual reviews, and business responses across active listings.

Category & Taxonomy Mapping

Maintain the exact category hierarchy used by paginasamarillas.es for accurate market segmentation.

Regional Filtering

Target specific provinces like Madrid, Barcelona, or Valencia for localised data acquisition.

Anti-Bot Circumvention

Bypass rate limits and CAPTCHAs using Spanish residential proxies and fingerprint spoofing.

Scheduled Updates

Run continuous pipelines to detect new business registrations and closed entities on a weekly or monthly cadence.

// engagement pipeline

From target provinces to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target categories, provinces, or search queries. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, and CAPTCHA handling for paginasamarillas.es.

Validation & QA
d 4–6

Schema validation, null-rate checks, and data quality tests before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our directory pipeline handles the hard parts

Extracting data from paginasamarillas.es requires defeating regional blocks and navigating inconsistent listing formats. Here is how we build resilient pipelines.

pipeline-monitor · paginasamarillas.es · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
Spanish residential proxies

Paginasamarillas.es aggressively blocks non-EU datacenter traffic. Our crawlers use localised Spanish residential proxies with realistic browser fingerprints to maintain high success rates.

DOM layout variations
Resilient selectors for premium vs free listings

The directory uses multiple listing templates depending on the business subscription tier. We use fallback chains to extract core fields regardless of the visual layout.

Contact obfuscation
Playwright execution for hidden fields

Phone numbers and emails are often masked behind JavaScript events. We run full Playwright browser sessions to trigger the necessary interactions and capture the revealed data.

Pagination limits
Granular grid search strategies

Search results are capped at a specific page limit. We bypass this by dividing broad queries into hyper-localised geographic grids to ensure full category coverage.

Monitoring & alerting
24/7 pipeline health

We track null rates on critical fields like phone numbers and physical addresses. Any layout change triggers an immediate alert to our engineering team.

Applications

Who uses Spanish directory data

Teams across industries use paginasamarillas.es data to build competitive products and smarter operations.

01
B2B Lead Generation

Sales teams use extracted contact details to build targeted outreach lists across specific Spanish provinces and industries.

02
Local SEO & Citation Building

Agencies audit directory consistency across the Spanish local search ecosystem to improve client visibility.

03
Market Research

Analysts track business density and category growth across different regions to identify market trends.

04
Mapping & Navigation

GIS teams integrate accurate coordinate and address data into spatial applications and routing software.

05
Competitor Analysis

Businesses monitor local competitors, their ratings, and service offerings to adjust their own positioning.

06
Master Data Management

Enterprises enrich their existing CRM records with verified directory data to maintain database accuracy.

Why DataFlirt

"Paginasamarillas.es holds the most comprehensive registry of Spanish SMEs, but extracting clean, structured contact data requires bypassing aggressive regional rate limits."

Extracting business directory data requires handling inconsistent listing templates, JavaScript-obfuscated contact details, and strict regional blocking. DataFlirt manages the residential proxy rotation and selector maintenance so you receive clean data directly in your warehouse.

Technical Spec

Paginasamarillas.es scraper — technical capabilities

Everything supported by our paginasamarillas.es scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Playwright sessions required for revealing masked contact details
Supported
Spanish residential proxies
Localised IPs to bypass non-EU geographic blocks
Supported
CAPTCHA bypass
Automated 2Captcha + CapSolver integration
Supported
Deep pagination
Granular grid search to bypass 50-page limits
Supported
Change detection (diffs)
Hash-based diff to track updated phone numbers and addresses
Supported
Webhook delivery
HTTP POST per record for immediate CRM ingestion
Supported
Premium account analytics
Traffic stats provided only to business owners
Partial
User account passwords
Gated user profiles and billing information
Partial
Infrastructure

Infrastructure powering the directory pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering and interaction flows for hidden contact data.

Residential Proxy Infrastructure

We maintain pools of Spanish residential ISP proxies to bypass geographic restrictions and rate limits.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling and dependency management for large-scale directory crawls.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested objects
CSV
Flat file with typed columns for CRM imports
XLS
Excel format for manual sales team review
Parquet
Columnar format for BigQuery and Snowflake
AWS S3
Direct bucket delivery
Webhook
HTTP POST per record
API
REST endpoint to query extracted records
BigQuery
Streamed directly into your dataset
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About paginasamarillas.es scraping, legality, and pipeline operations.

Ask us directly →
Is scraping paginasamarillas.es legal?

Scraping publicly available directory information is generally permissible under applicable law. DataFlirt targets only public business data and contact details. We do not extract personal user data or circumvent authentication walls.

How do you extract hidden phone numbers?

We use Playwright to simulate user interaction, clicking the necessary buttons to reveal JavaScript-obfuscated phone numbers and email addresses.

Can you scrape specific Spanish provinces?

Yes. We can target specific regions, cities, or postal codes to build localised datasets based on your exact requirements.

How fresh is the data?

Pipelines can be scheduled to run weekly or monthly to capture new business registrations, closed entities, and updated contact information.

Do you extract business reviews?

Yes. We extract user ratings, review text, and business responses from the directory listings.

Can I request a sample dataset?

Yes. We provide a sample run of up to 500 business listings during the scoping process to validate schema fit and data quality.

$ dataflirt scope --new-project --source=paginasamarillas.es ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a specific category extract or a full national database sync, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in business directories

Services

Data Extraction for Every Industry

View All Services →