We extract company profiles, verified phone numbers, addresses, reviews, and operating hours from paginasamarillas.es. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Company Profiles objects from paginasamarillas.es. All fields typed and schema-versioned.
"business_id": "ES-MAD-847291", "name": "Fontanería Martínez", "primary_category": "Fontaneros", "subcategories": "['Instalaciones', 'Urgencias 24h']", "year_established": 1998, "website_url": "https://fontaneriamartinez.es", "logo_url": "https://paginasamarillas.es/logos/847291.jpg"
| # | business_id | name | legal_name | description | primary_category | subcategories |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Contact Information objects from paginasamarillas.es. All fields typed and schema-versioned.
"business_id": "ES-MAD-847291", "phone_primary": "+34 91 555 12 34", "mobile": "+34 600 123 456", "whatsapp": "+34 600 123 456", "email_address": "contacto@fontaneriamartinez.es", "social_facebook": "facebook.com/fontaneriamartinez", "social_linkedin": "linkedin.com/company/fontaneriamartinez"
| # | business_id | phone_primary | phone_secondary | mobile | email_address | |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Location Data objects from paginasamarillas.es. All fields typed and schema-versioned.
"business_id": "ES-MAD-847291", "street_address": "Calle de Alcalá 142", "city": "Madrid", "province": "Madrid", "postal_code": "28009", "latitude": 40.4245, "longitude": -3.6741, "region": "Comunidad de Madrid"
| # | business_id | street_address | city | province | postal_code | region |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Operating Hours objects from paginasamarillas.es. All fields typed and schema-versioned.
"business_id": "ES-MAD-847291", "monday": "09:00-18:00", "tuesday": "09:00-18:00", "wednesday": "09:00-18:00", "saturday": "10:00-14:00", "sunday": "Closed", "timezone": "Europe/Madrid"
| # | business_id | monday | tuesday | wednesday | thursday | friday |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Reviews & Ratings objects from paginasamarillas.es. All fields typed and schema-versioned.
"review_id": "REV-847291-001", "business_id": "ES-MAD-847291", "author_name": "Carlos Ruiz", "rating_score": 4.5, "review_text": "Servicio rápido y profesional.", "review_date": "2025-11-04", "helpful_votes": 3
| # | review_id | business_id | author_name | rating_score | review_text | review_date |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our paginasamarillas.es scraper handles complex regional blocking, JavaScript-rendered contact details, and diverse listing templates to deliver structured B2B datasets.
Extract name, description, categories, and core metadata for millions of Spanish SMEs across all provinces.
Capture phone numbers, emails, and website URLs directly from directory listings, bypassing JavaScript obfuscation.
Parse structured addresses into street, city, province, postal code, and coordinate data for spatial analysis.
Normalise unstructured opening hours into standard ISO formats for 7-day schedules.
Collect user ratings, textual reviews, and business responses across active listings.
Maintain the exact category hierarchy used by paginasamarillas.es for accurate market segmentation.
Target specific provinces like Madrid, Barcelona, or Valencia for localised data acquisition.
Bypass rate limits and CAPTCHAs using Spanish residential proxies and fingerprint spoofing.
Run continuous pipelines to detect new business registrations and closed entities on a weekly or monthly cadence.
Brief in. Clean data out.
Provide target categories, provinces, or search queries. We design the extraction schema together.
We configure Scrapy / Playwright crawlers, proxy rotation, and CAPTCHA handling for paginasamarillas.es.
Schema validation, null-rate checks, and data quality tests before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Extracting data from paginasamarillas.es requires defeating regional blocks and navigating inconsistent listing formats. Here is how we build resilient pipelines.
Paginasamarillas.es aggressively blocks non-EU datacenter traffic. Our crawlers use localised Spanish residential proxies with realistic browser fingerprints to maintain high success rates.
The directory uses multiple listing templates depending on the business subscription tier. We use fallback chains to extract core fields regardless of the visual layout.
Phone numbers and emails are often masked behind JavaScript events. We run full Playwright browser sessions to trigger the necessary interactions and capture the revealed data.
Search results are capped at a specific page limit. We bypass this by dividing broad queries into hyper-localised geographic grids to ensure full category coverage.
We track null rates on critical fields like phone numbers and physical addresses. Any layout change triggers an immediate alert to our engineering team.
Sales teams use extracted contact details to build targeted outreach lists across specific Spanish provinces and industries.
Agencies audit directory consistency across the Spanish local search ecosystem to improve client visibility.
Analysts track business density and category growth across different regions to identify market trends.
GIS teams integrate accurate coordinate and address data into spatial applications and routing software.
Businesses monitor local competitors, their ratings, and service offerings to adjust their own positioning.
Enterprises enrich their existing CRM records with verified directory data to maintain database accuracy.
"Paginasamarillas.es holds the most comprehensive registry of Spanish SMEs, but extracting clean, structured contact data requires bypassing aggressive regional rate limits."
Extracting business directory data requires handling inconsistent listing templates, JavaScript-obfuscated contact details, and strict regional blocking. DataFlirt manages the residential proxy rotation and selector maintenance so you receive clean data directly in your warehouse.
Everything supported by our paginasamarillas.es scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering and interaction flows for hidden contact data.
We maintain pools of Spanish residential ISP proxies to bypass geographic restrictions and rate limits.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling and dependency management for large-scale directory crawls.
Data delivered to where your team already works — no new tooling required.
About paginasamarillas.es scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available directory information is generally permissible under applicable law. DataFlirt targets only public business data and contact details. We do not extract personal user data or circumvent authentication walls.
We use Playwright to simulate user interaction, clicking the necessary buttons to reveal JavaScript-obfuscated phone numbers and email addresses.
Yes. We can target specific regions, cities, or postal codes to build localised datasets based on your exact requirements.
Pipelines can be scheduled to run weekly or monthly to capture new business registrations, closed entities, and updated contact information.
Yes. We extract user ratings, review text, and business responses from the directory listings.
Yes. We provide a sample run of up to 500 business listings during the scoping process to validate schema fit and data quality.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a specific category extract or a full national database sync, we scope, build, and operate the pipeline. Tell us what you need.