SYSTEM all green source spiegel.de queue 12,408 URLs p99 latency 184ms dataflirt.com · scraper/spiegel-de
RUN: 42 active pipelines: spiegel.de live

Spiegel data,
at warehouse scale.

We extract news articles, author metadata, publication timestamps, and comment threads from spiegel.de. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Articles extracted
14.2K /day
Comment threads
84.1K /24h
Author profiles
1.2K /run
Active pipelines
42
Uptime
99.94%
Data Dictionary

Every field we extract from spiegel.de

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Articles objects from spiegel.de. All fields typed and schema-versioned.

urlheadlinesubheadlineauthorpublished_atupdated_atcategorytagsbody_textword_countis_premium
articles
● 200 OK
"url": "https://www.spiegel.de/politik/example-article",
"headline": "Bundesregierung plant neue Richtlinien",
"author": "Markus Becker",
"published_at": "2026-05-12T08:30:00Z",
"category": "Politik",
"is_premium": false
# urlheadlinesubheadlineauthorpublished_atupdated_at
1
2
3

Complete list of extractable fields for Frontpage objects from spiegel.de. All fields typed and schema-versioned.

positionsectionheadlineurlimage_urlbadgeis_breakingscraped_at
frontpage
● 200 OK
"position": 1,
"section": "Top News",
"headline": "Wirtschaftswachstum übertrifft Erwartungen",
"url": "https://www.spiegel.de/wirtschaft/example",
"is_breaking": true,
"scraped_at": "2026-05-12T09:14:33Z"
# positionsectionheadlineurlimage_urlbadge
1
2
3

Complete list of extractable fields for Comments objects from spiegel.de. All fields typed and schema-versioned.

comment_idarticle_urluser_nameuser_idtimestamptextupvotesdownvotesreply_to_id
comments
● 200 OK
"comment_id": "c_987654321",
"user_name": "BerlinReader88",
"timestamp": "2026-05-12T10:05:12Z",
"text": "Das ist eine sehr interessante Entwicklung.",
"upvotes": 42,
"downvotes": 3
# comment_idarticle_urluser_nameuser_idtimestamptext
1
2
3

Complete list of extractable fields for Authors objects from spiegel.de. All fields typed and schema-versioned.

author_idnameprofile_urlbiotwitter_handlearticle_countrecent_articlesrole
authors
● 200 OK
"author_id": "a_12345",
"name": "Markus Becker",
"profile_url": "https://www.spiegel.de/impressum/autor-12345",
"article_count": 342,
"role": "Redakteur",
"bio": "Berichtet über Technologie und Wissenschaft."
# author_idnameprofile_urlbiotwitter_handlearticle_count
1
2
3

Complete list of extractable fields for Multimedia objects from spiegel.de. All fields typed and schema-versioned.

article_urlimage_urlcaptioncopyrightwidthheightformatvideo_urlduration
multimedia
● 200 OK
"article_url": "https://www.spiegel.de/panorama/example",
"image_url": "https://cdn.prod.www.spiegel.de/images/example.jpg",
"caption": "Demonstration in der Innenstadt",
"copyright": "DPA / Reuters",
"format": "image/jpeg",
"width": 1920
# article_urlimage_urlcaptioncopyrightwidthheight
1
2
3

Capabilities

Everything you need from Spiegel, nothing you do not

Our Spiegel scraper handles the entire news platform: frontpage hierarchies, historical article archives, author profiles, and paginated comment threads, with strict anti-bot circumvention built in.

Full Article Text

Headline, intro, body paragraphs, and blockquotes parsed cleanly into structural JSON.

Author and Metadata Tracking

Capture bylines, publication dates, modification timestamps, and assigned categories.

Frontpage Hierarchy

Track article placement, section prominence, and breaking news badges over time.

Comment Section Mining

Extract paginated user comments, timestamps, and upvote metrics for sentiment analysis.

Tag and Keyword Extraction

Map articles to Spiegel internal taxonomy and keyword tags.

Historical Archive Access

Crawl historical sitemaps to build comprehensive German language datasets.

Paywall Detection

Identify Spiegel+ premium articles and flag them in the metadata schema.

Multimedia Extraction

Capture image URLs, captions, copyright credits, and embedded video metadata.

Scheduled and Streaming Modes

Run continuous pipelines at 5-minute cadences for breaking news or daily for archives.

// engagement pipeline

From URL list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target sections, author URLs, or search terms. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, proxy rotation, session management, and parsing logic for spiegel.de.

Validation & QA
d 4–6

Schema validation, null-rate checks, text-encoding verification, and sample datasets before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Spiegel pipeline handles the hard parts

News sites employ aggressive caching and bot mitigation. Here is how we stay resilient, and why teams choose managed infrastructure over DIY.

pipeline-monitor · spiegel.de · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
Residential proxy rotation and fingerprint spoofing

Spiegel employs rate limiting and bot detection. Our crawlers use German residential ISP proxies with realistic browser fingerprints and randomised request timing.

Dynamic comments
Full Playwright execution for React content

Comment sections load dynamically via JavaScript. We run full Playwright browser sessions with JavaScript execution to capture user discussions.

Layout volatility
Resilient selectors with fallback chains

Spiegel feature articles use custom layouts. Our selector strategy uses multiple fallback chains per field so a layout change does not break your data pipeline.

Change detection
Track stealth edits

We maintain a hash index of last-seen values for article text. Subsequent runs push diffs, allowing you to track editorial changes post-publication.

Paywall logic
Spiegel+ metadata extraction

We detect paywalled content automatically, extracting available metadata and previews without triggering authentication blocks or crawler traps.

Applications

Who uses Spiegel data, and how

Teams across industries use spiegel.de data to build competitive products and smarter operations.

01
Media Monitoring

PR teams track brand mentions, executive coverage, and crisis events across primary news sections.

02
Sentiment Analysis

Quantifying public reaction via comment threads on political and economic articles.

03
NLP Model Training

Building high-quality German language models using editorial text and verified grammar.

04
Competitor Intelligence

Rival publishers analyse Spiegel content strategy, publication cadence, and author output.

05
Political Discourse Tracking

Researchers monitor election coverage, narrative framing, and tag frequencies.

06
Financial Signal Extraction

Hedge funds parse business news for macroeconomic indicators and corporate announcements.

Why DataFlirt

"Der Spiegel represents the pinnacle of German editorial content, a critical corpus for any serious European media monitoring or NLP initiative."

Extracting news at scale requires more than basic HTTP requests. You must navigate aggressive edge caching, dynamic comment loading, varied article templates, and strict rate limits. DataFlirt manages this complexity so your engineering team receives structured text, not HTML errors.

Technical Spec

Spiegel scraper: technical capabilities

Everything supported by our spiegel.de scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for comment sections and dynamic embeds
Supported
Residential proxy rotation
ISP-grade residential IPs from DE pools rotated per request
Supported
Full-text extraction
Clean paragraph parsing without ad injection or tracking scripts
Supported
Frontpage rank tracking
Position capture on spiegel.de index at 5-minute intervals
Supported
Archive crawling
Sitemap traversal for historical articles dating back decades
Supported
Change detection
Hash-based diff to emit records with changed text since last run
Supported
Webhook delivery
HTTP POST per record for real-time breaking news workflows
Supported
Spiegel+ Premium Content
Full text of paywalled articles without user credentials
Partial
User Account Profiles
Scraping private user account data, reading history, or saved articles
Partial
Infrastructure

Infrastructure powering the Spiegel pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across DE regions. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.

Cloud-Native Orchestration

Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested, schema versioned per run
CSV
Flat file with typed columns, Excel/Sheets compatible
XLS
Legacy spreadsheet format for offline analysis
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery, compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoints for on-demand query access
BigQuery
Streamed directly into your dataset with schema auto-detect
PostgreSQL
Upsert into your existing schema with conflict resolution
Snowflake
Stage and COPY INTO workflow, incremental or full-replace
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About spiegel.de scraping, legality, and pipeline operations.

Ask us directly →
Is scraping spiegel.de legal?

Scraping publicly available information from spiegel.de is generally permissible under applicable law. DataFlirt targets only public, non-authenticated article and metadata. We do not extract personal data, circumvent authentication walls, or violate GDPR.

Can you bypass the Spiegel+ paywall?

No. We detect Spiegel+ articles and extract the publicly available headline, intro, and metadata, but we do not circumvent the paywall to access premium body text.

Do you extract user comments?

Yes. We extract paginated user comments including upvotes, downvotes, timestamps, and usernames using headless browser automation.

How fast can you deliver breaking news?

Real-time streaming pipelines achieve sub-5-minute latency for frontpage tracking and new article detection.

What languages are supported?

We primarily extract the German language corpus, but we also support the Spiegel International English section using the same schema.

Can you track article edits over time?

Yes. Every pipeline run produces timestamped snapshots. We maintain a hash of the article text and can deliver diffs when editorial changes occur post-publication.

Do you extract images and video?

We extract the image URLs, captions, copyright metadata, and video embed links. We do not download the binary media files to your warehouse.

$ dataflirt scope --new-project --source=spiegel.de ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive dump or a continuous frontpage monitoring feed, we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in news and media

Services

Data Extraction for Every Industry

View All Services →