SYSTEM all green source zeit.de queue 12,841 URLs p99 latency 204ms dataflirt.com · scraper/zeit-de
RUN : 37 active pipelines : zeit.de live

Zeit.de data,
at warehouse scale.

We extract full text articles, author profiles, comment threads, and metadata from zeit.de. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Articles extracted
45.2K /day
Comments parsed
1.2M /day
Author profiles
8.4K /run
Active pipelines
37
Uptime
99.94%
Data Dictionary

Every field we extract from zeit.de

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Article Metadata objects from zeit.de. All fields typed and schema-versioned.

article_urlheadlinesubheadlineauthor_nameauthor_urlpublish_datelast_updatedsectionis_zpluscomment_counttopic_tags
article_metadata
● 200 OK
"article_url": "https://www.zeit.de/politik/deutschland/2026-10/bundestagswahl-ergebnisse",
"headline": "Die neuen Machtverhaeltnisse im Bundestag",
"author_name": "Anna Sauerbrey",
"publish_date": "2026-10-24T18:30:00Z",
"section": "Politik",
"is_zplus": false,
"comment_count": 1452
# article_urlheadlinesubheadlineauthor_nameauthor_urlpublish_date
1
2
3

Complete list of extractable fields for Full Text Content objects from zeit.de. All fields typed and schema-versioned.

article_urlheadlinelead_paragraphbody_textword_countimage_urlsembedded_linkspull_quotesis_truncatedpaywall_hit
full_text content
● 200 OK
"article_url": "https://www.zeit.de/wirtschaft/2026-10/inflation-ezb",
"word_count": 1240,
"lead_paragraph": "Die Europaeische Zentralbank senkt den Leitzins erneut.",
"is_truncated": true,
"paywall_hit": true,
"body_text": "Die Europaeische Zentralbank (EZB) hat am Donnerstag beschlossen..."
# article_urlheadlinelead_paragraphbody_textword_countimage_urls
1
2
3

Complete list of extractable fields for Comment Threads objects from zeit.de. All fields typed and schema-versioned.

comment_idarticle_urluser_nameuser_profile_urltimestampcomment_textupvotesis_recommendedreply_countparent_comment_id
comment_threads
● 200 OK
"comment_id": "cid-984210",
"user_name": "PolitikBeobachter99",
"timestamp": "2026-10-24T19:15:22Z",
"comment_text": "Ein sehr treffender Kommentar zur aktuellen Lage.",
"upvotes": 42,
"is_recommended": true,
"reply_count": 3
# comment_idarticle_urluser_nameuser_profile_urltimestampcomment_text
1
2
3

Complete list of extractable fields for Author Profiles objects from zeit.de. All fields typed and schema-versioned.

author_urlfull_namerole_titlebiographytwitter_handlearticle_countrecent_articlesprimary_topicsprofile_image_url
author_profiles
● 200 OK
"author_url": "https://www.zeit.de/autoren/S/Anna_Sauerbrey/index",
"full_name": "Anna Sauerbrey",
"role_title": "Koordinatorin Meinung",
"twitter_handle": "@AnnaSauerbrey",
"article_count": 342,
"primary_topics": "['Politik', 'Deutschland', 'USA']"
# author_urlfull_namerole_titlebiographytwitter_handlearticle_count
1
2
3

Complete list of extractable fields for Frontpage Rankings objects from zeit.de. All fields typed and schema-versioned.

scrape_timestampposition_rankarticle_urlheadlineblock_nameis_breaking_newsis_zplustime_on_homepage_minutes
frontpage_rankings
● 200 OK
"scrape_timestamp": "2026-10-25T08:00:00Z",
"position_rank": 1,
"article_url": "https://www.zeit.de/politik/deutschland/2026-10/koalitionsverhandlungen",
"block_name": "Aufmacher",
"is_breaking_news": true,
"is_zplus": false
# scrape_timestampposition_rankarticle_urlheadlineblock_nameis_breaking_news
1
2
3

Capabilities

Extract German journalism data with precision

Our zeit.de scraper handles consent walls, dynamic paywall detection, infinite scroll comments, and complex German character encoding, delivering clean text corpora for NLP and media analysis.

Full Article Extraction

Extract headlines, subheadlines, lead paragraphs, and full body text from free articles, with precise paragraph separation and image link capture.

Z+ Paywall Detection

Accurately flag Z+ premium articles. Capture available preview text and metadata without triggering false positives or pipeline failures.

Comment Thread Mining

Paginate through thousands of user comments per article. Capture timestamps, upvotes, editor recommendations, and nested reply structures.

Author Network Mapping

Scrape author directories and profile pages. Link journalists to their full article history, primary topics, and biographical metadata.

Section & Topic Tracking

Categorise content by section (Politik, Wirtschaft, Gesellschaft) and extract granular topic tags attached to every article.

German Encoding Standardisation

Automatic normalisation of umlauts and special characters (UTF-8) to ensure clean ingestion into your NLP models and databases.

Frontpage Rank Tracking

Monitor the zeit.de homepage at high frequency to track article placement, breaking news banners, and editorial prioritisation over time.

Consent Wall Circumvention

Automated handling of the Pur-Abo consent banners via cookie injection and session management, ensuring uninterrupted access to public pages.

Scheduled Archiving

Run daily or hourly pipelines to build a comprehensive historical archive of German media discourse and publication trends.

// engagement pipeline

From URL list to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target sections, author lists, or historical date ranges. We map the required data fields.

Pipeline Build
d 2–4

We configure crawlers to handle the Pur-Abo consent wall, Z+ detection, and comment pagination.

Validation & QA
d 4–6

Schema validation, character encoding checks, and paywall flag verification before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Navigating zeit.de infrastructure

European media sites deploy aggressive consent walls and dynamic paywalls. Here is how we maintain reliable extraction.

pipeline-monitor · zeit.de · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Consent Wall Bypass
Automated Pur-Abo handling

Zeit.de uses a strict Pur-Abo model requiring users to accept tracking or pay for a subscription. Our infrastructure programmatically manages consent cookies and session states to access public content without manual intervention.

Paywall Dynamics
Accurate Z+ flagging

Articles frequently switch from free to Z+ premium based on traffic velocity. We detect paywall states dynamically, capturing full text when available and gracefully degrading to preview text and metadata when gated.

Comment Pagination
Deep thread extraction

Popular articles generate thousands of comments loaded via complex asynchronous requests. We trace the API calls to extract complete, deeply nested discussion threads rather than just the top ten visible comments.

DOM Variability
Resilient article selectors

Live blogs, interactive graphics, and standard articles use different DOM structures. We deploy content-type specific parsing logic to ensure clean text extraction regardless of the editorial format.

German IP Pools
Localised residential proxies

To avoid geo-blocking and receive the correct regional content variants, we route all requests through high-quality German residential proxies.

Applications

Who uses zeit.de data

Teams across industries use zeit.de data to build competitive products and smarter operations.

01
Media Sentiment Analysis

Track editorial tone and public reaction in comment sections regarding political events and corporate news.

02
German NLP Training

Ingest high-quality, editorially reviewed German text to train large language models and translation engines.

03
Political Discourse Tracking

Analyse topic frequency, author bias, and keyword prominence during election cycles or major policy shifts.

04
Competitor Paywall Analysis

Publishers monitor zeit.de to understand which topics drive Z+ conversions and how long articles remain free.

05
Author & Journalist Mapping

PR firms and researchers track specific journalists, their output frequency, and their primary coverage areas.

06
Topic Trend Forecasting

Identify emerging narratives by tracking the velocity of new tags and section categorisations over time.

Why DataFlirt

"Zeit.de represents a foundational corpus of German journalism and political discourse, but extracting it requires navigating consent walls and dynamic paywalls."

Most teams fail at scraping German news media because they cannot handle the Pur-Abo consent banners or reliably separate free content from Z+ premium articles. DataFlirt manages the session cookies, proxies, and selector maintenance so your data science teams receive clean text.

Technical Spec

Zeit.de scraper technical specifications

Everything supported by our zeit.de scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Consent Wall Bypass
Automated acceptance of tracking cookies to view free content
Supported
Z+ Premium Full Text
Extraction of text hidden behind the paid subscription wall
Partial
Comment Thread Extraction
Full pagination of user comments, including nested replies
Supported
Live Blog Parsing
Continuous extraction of timestamped updates on breaking news pages
Supported
Author Directory Scraping
Capture of journalist bios, roles, and article histories
Supported
Frontpage Rank Tracking
High-frequency monitoring of article positions on the homepage
Supported
Registered User Account Details
Private email addresses or billing info of commenters
Partial
UTF-8 Encoding Normalisation
Clean output of German umlauts and special characters
Supported
Infrastructure

Infrastructure powering the pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Session & Cookie Management

We maintain persistent cookie jars to navigate the Pur-Abo consent walls, ensuring every request returns valid HTML rather than a redirect to the consent banner.

Localised Proxy Routing

Traffic is routed exclusively through German residential IPs to mimic authentic local readership and avoid aggressive geo-blocking.

Asynchronous API Parsing

Instead of slow browser automation for comments, we reverse-engineer zeit.de internal APIs to extract discussion threads rapidly and reliably.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Nested structures ideal for articles and comment threads
CSV
Flat files for author lists and metadata analysis
XLS
Excel format for manual editorial review
Parquet
Columnar format for fast querying in data warehouses
AWS S3
Direct bucket delivery for your data lake
Webhook
HTTP POST for real-time breaking news alerts
API
On-demand REST endpoints for specific article queries
PostgreSQL
Direct database inserts for structured archives
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About zeit.de scraping, legality, and pipeline operations.

Ask us directly →
Is scraping zeit.de legal?

Scraping publicly available news articles and comments is generally permissible for research and analysis, provided it complies with copyright laws regarding republication. DataFlirt extracts factual metadata, public text, and topic tags. We do not bypass cryptographic paywalls or extract private user data. Clients must ensure their downstream use cases comply with German copyright and GDPR regulations.

Do you bypass the Z+ paywall?

No. We do not hack or bypass paid subscription walls. If an article is flagged as Z+, we extract the headline, metadata, and whatever preview text is publicly visible before the paywall cuts off the content.

How do you handle the Pur-Abo consent wall?

Our infrastructure automatically accepts the necessary tracking cookies required to view the free version of the site. We manage these session tokens programmatically across our proxy pool to maintain continuous access.

Can you extract historical articles?

Yes. We can crawl the zeit.de sitemaps and search archives to extract articles published years ago, provided the URLs are still active and the content is not retroactively paywalled.

What about German character encoding?

All pipelines enforce strict UTF-8 encoding. Umlauts (ä, ö, ü) and the eszett (ß) are correctly preserved in the final JSON, CSV, or Parquet output, preventing corruption in your NLP pipelines.

How fast can you scrape breaking news?

For frontpage monitoring, we can configure pipelines to run at sub-5-minute intervals, capturing breaking news banners and position changes in near real-time.

Can you extract user comments?

Yes. We extract complete comment threads, including nested replies, timestamps, upvote counts, and author names, which is highly valuable for sentiment analysis.

$ dataflirt scope --new-project --source=zeit.de ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. From daily frontpage monitoring to massive historical NLP corpora, we build and manage the pipeline. Tell us your data requirements.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in news and media

Services

Data Extraction for Every Industry

View All Services →