SYSTEM all green source elmundo.es queue 12,491 URLs p99 latency 185ms dataflirt.com · scraper/elmundo-es
RUN · 42 active pipelines · elmundo.es live

El Mundo data,
at warehouse scale.

We extract publication archives, breaking news feeds, author profiles, and comment structures from elmundo.es. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Articles extracted
14.2K /day
Comment records
85.4K /24h
Author profiles
1,240 /run
Active pipelines
42
Uptime
99.98%
Data Dictionary

Every field we extract from elmundo.es

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Articles objects from elmundo.es. All fields typed and schema-versioned.

urlheadlinesubheadlineauthorpublish_dateupdate_datesectionbody_texttagsis_premiumimage_urlcomment_count
articles
● 200 OK
"url": "https://www.elmundo.es/espana/2026/05/12/example.html",
"headline": "El Gobierno aprueba la nueva reforma fiscal",
"author": "Juan M. Lamet",
"publish_date": "2026-05-12T08:30:00Z",
"section": "España",
"is_premium": true,
"comment_count": 342,
"tags": "['Política', 'Impuestos', 'Congreso']"
# urlheadlinesubheadlineauthorpublish_dateupdate_date
1
2
3

Complete list of extractable fields for Authors objects from elmundo.es. All fields typed and schema-versioned.

author_idnameroletwitter_handlearticle_countlatest_article_urlbio_textprofile_image
authors
● 200 OK
"author_id": "juan-m-lamet",
"name": "Juan M. Lamet",
"role": "Redactor",
"twitter_handle": "@juanmlamet",
"article_count": 845,
"latest_article_url": "https://www.elmundo.es/espana/2026/05/12/example.html"
# author_idnameroletwitter_handlearticle_countlatest_article_url
1
2
3

Complete list of extractable fields for Comments objects from elmundo.es. All fields typed and schema-versioned.

comment_idarticle_urlusernametimestampcomment_textupvotesdownvotesreplies_countis_nested
comments
● 200 OK
"comment_id": "c_9823749",
"article_url": "https://www.elmundo.es/espana/2026/05/12/example.html",
"username": "LectorHabitual",
"timestamp": "2026-05-12T09:15:22Z",
"comment_text": "Esta medida afectará principalmente a las pymes.",
"upvotes": 45,
"replies_count": 3
# comment_idarticle_urlusernametimestampcomment_textupvotes
1
2
3

Complete list of extractable fields for Sections & Feeds objects from elmundo.es. All fields typed and schema-versioned.

section_namepositionheadlineurlis_breakingpinnedscrape_timestamprelated_links
sections_& feeds
● 200 OK
"section_name": "Economía",
"position": 1,
"headline": "El Ibex 35 cierra en verde tras la decisión del BCE",
"url": "https://www.elmundo.es/economia/2026/05/12/ibex.html",
"is_breaking": false,
"scrape_timestamp": "2026-05-12T18:05:00Z"
# section_namepositionheadlineurlis_breakingpinned
1
2
3

Complete list of extractable fields for Multimedia objects from elmundo.es. All fields typed and schema-versioned.

media_idarticle_urlmedia_typetitledurationthumbnail_urlembed_codeprovider
multimedia
● 200 OK
"media_id": "vid_48291",
"article_url": "https://www.elmundo.es/deportes/2026/05/12/video.html",
"media_type": "video",
"title": "Resumen de la jornada de Champions",
"duration": "00:03:45",
"provider": "Unidad Editorial"
# media_idarticle_urlmedia_typetitledurationthumbnail_url
1
2
3

Capabilities

Everything you need from El Mundo — nothing you don't

Our El Mundo scraper handles every layer of the platform: frontpage feeds, historical archives, author tracking, and comment threads — with JavaScript rendering and anti-bot circumvention built in.

Full Article Extraction

Extract headlines, subheadings, full body text, publication timestamps, and metadata tags from any section.

Frontpage & Section Tracking

Monitor article placement, breaking news banners, and pinned content across the main homepage and sub-sections.

Author & Journalist Tracking

Capture bylines, author bios, social media handles, and historical publication records for specific journalists.

Comment & Sentiment Mining

Extract user comments, timestamps, upvote/downvote ratios, and nested reply structures for sentiment analysis.

Paywall Detection

Accurately flag articles marked as 'Premium' to separate free public discourse from gated subscriber content.

Regional Edition Scraping

Crawl specific local editions including Madrid, Andalucía, Cataluña, and Comunidad Valenciana.

Archive Crawling

Paginate through El Mundo's historical hemeroteca to build deep retrospective datasets.

Multimedia Metadata

Extract image URLs, video embed codes, captions, and media provider details attached to articles.

Scheduled + Streaming Modes

Run continuous pipelines for breaking news or one-off bulk exports for historical analysis.

// engagement pipeline

From URL list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target sections, author lists, or keyword parameters. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, European proxy rotation, and CMP consent bypass logic.

Validation & QA
d 4–6

Schema validation, null-rate checks, and text parsing verification before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our El Mundo pipeline handles the hard parts

News sites deploy strict caching and bot protections. Here's how we stay resilient — and why teams choose managed infrastructure over DIY.

pipeline-monitor · elmundo.es · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Bot protection
Akamai bypass and residential IPs

El Mundo uses enterprise CDN and bot mitigation layers. Our crawlers utilize Spanish residential proxies with realistic browser fingerprints to bypass request blocking and IP bans.

CMP Handling
Automated EU cookie consent

We programmatically accept or dismiss Didomi cookie consent modals required under GDPR, ensuring the underlying DOM is fully accessible for scraping.

Dynamic Comments
Playwright for lazy-loaded threads

Comment sections on elmundo.es load asynchronously via JavaScript. We use full browser rendering to trigger these network requests and capture the complete discussion tree.

Schema stability
Normalising varied article templates

Opinion pieces, multimedia galleries, and standard news articles use different DOM structures. We maintain fallback selector chains to extract clean text regardless of the presentation format.

Change detection
Tracking article updates

News stories evolve. We track 'update_date' timestamps and content hashes to emit new records only when an article is modified, reducing redundant data.

Applications

Who uses El Mundo data — and how

Teams across industries use elmundo.es data to build competitive products and smarter operations.

01
Media Monitoring

PR agencies and brands track mentions, quotes, and narrative framing across Spain's leading daily newspaper.

02
Sentiment Analysis

Researchers analyse user comments on political and economic articles to gauge public reaction and polarity.

03
NLP & LLM Training

AI teams build high-quality Spanish language corpuses using decades of professionally edited journalistic text.

04
Competitor Intelligence

Other media organisations monitor El Mundo's publishing velocity, topic coverage, and breaking news latency.

05
Political & Social Research

Academics track editorial trends, keyword frequency, and author bias across election cycles.

06
Financial Signals

Quant funds extract market news and corporate announcements from the Economy section to inform trading models.

Why DataFlirt

"El Mundo represents one of the most critical records of Spanish public discourse — but extracting clean text from its varied digital templates requires specialized infrastructure."

Most teams underestimate the investment required: reliable news scraping requires bypassing aggressive CDN caching, handling EU cookie consent popups, rendering dynamic comment sections, and normalising dozens of distinct article layouts. DataFlirt absorbs that complexity so your engineers can focus on the analysis — not the infrastructure.

Technical Spec

El Mundo scraper — technical capabilities

Everything supported by our elmundo.es scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for dynamic comment threads and interactive graphics
Supported
EU Cookie Consent bypass
Automated interaction with Didomi CMP to access article text
Supported
Residential proxy rotation
Spanish and European residential IPs rotated to avoid CDN blocks
Supported
Archive pagination
Deep crawling of historical hemeroteca indices by date
Supported
Paywall detection
Accurate flagging of articles locked behind the Premium tier
Supported
Comment thread extraction
Capture of nested replies, usernames, and upvote metrics
Supported
Change detection
Hash-based diffing to capture article updates and headline changes
Supported
Premium Content Text
Full text extraction of El Mundo Premium paywalled articles
Partial
Orbyt Digital Kiosk
Extraction of PDF or e-reader replica editions from Orbyt
Partial
Infrastructure

Infrastructure powering the El Mundo pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles broad crawling of section feeds and archives. Playwright executes JavaScript for comment loading and consent management.

Residential Proxy Infrastructure

We maintain pools of European residential proxies. Rotation happens per-request to bypass Akamai and Cloudflare protections.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling for continuous news feeds. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — ideal for complex article structures
CSV
Flat file with typed columns for metadata analysis
XLS
Excel compatible format for editorial teams
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time breaking news alerts
API
REST endpoint to query your extracted datasets
BigQuery
Streamed directly into your dataset
Snowflake
Stage + COPY INTO workflow
Postgres
Upsert into your existing schema
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About elmundo.es scraping, legality, and pipeline operations.

Ask us directly →
Is scraping El Mundo legal?

Scraping publicly available news articles and headlines is generally permissible. DataFlirt extracts only public, non-authenticated data. We do not bypass paywalls to steal Premium content, nor do we extract personal user data beyond public comment usernames. Clients should review terms of service and consult legal counsel for specific use cases.

How do you handle El Mundo's bot protection?

We use Spanish residential proxies and full browser rendering to mimic legitimate user behaviour, effectively bypassing CDN-level bot mitigation and request rate limits.

Can you extract historical articles?

Yes. We can crawl El Mundo's hemeroteca (archive) to extract historical articles based on specified date ranges, sections, or keyword parameters.

How fast can you deliver breaking news?

For continuous monitoring pipelines targeting the frontpage or specific sections, we can achieve sub-5-minute latency using Webhook delivery.

Do you extract comments from all articles?

We extract comments from articles where the discussion thread is enabled and publicly visible. This includes paginating through the thread to capture replies.

Can you bypass the Premium paywall?

No. We do not circumvent authentication or payment gateways. We extract the headline, subheadline, and any publicly visible introductory text, and flag the record as 'is_premium: true'.

Can I request a sample dataset?

Yes. We provide a sample run of up to 500 articles across various sections during the scoping process to validate schema fit and data quality.

$ dataflirt scope --new-project --source=elmundo.es ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off archive dump or a continuous news-monitoring feed across specific sections — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in news and media

Services

Data Extraction for Every Industry

View All Services →