We extract full text articles, author profiles, comment threads, and metadata from zeit.de. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Article Metadata objects from zeit.de. All fields typed and schema-versioned.
"article_url": "https://www.zeit.de/politik/deutschland/2026-10/bundestagswahl-ergebnisse", "headline": "Die neuen Machtverhaeltnisse im Bundestag", "author_name": "Anna Sauerbrey", "publish_date": "2026-10-24T18:30:00Z", "section": "Politik", "is_zplus": false, "comment_count": 1452
| # | article_url | headline | subheadline | author_name | author_url | publish_date |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Full Text Content objects from zeit.de. All fields typed and schema-versioned.
"article_url": "https://www.zeit.de/wirtschaft/2026-10/inflation-ezb", "word_count": 1240, "lead_paragraph": "Die Europaeische Zentralbank senkt den Leitzins erneut.", "is_truncated": true, "paywall_hit": true, "body_text": "Die Europaeische Zentralbank (EZB) hat am Donnerstag beschlossen..."
| # | article_url | headline | lead_paragraph | body_text | word_count | image_urls |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Comment Threads objects from zeit.de. All fields typed and schema-versioned.
"comment_id": "cid-984210", "user_name": "PolitikBeobachter99", "timestamp": "2026-10-24T19:15:22Z", "comment_text": "Ein sehr treffender Kommentar zur aktuellen Lage.", "upvotes": 42, "is_recommended": true, "reply_count": 3
| # | comment_id | article_url | user_name | user_profile_url | timestamp | comment_text |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Author Profiles objects from zeit.de. All fields typed and schema-versioned.
"author_url": "https://www.zeit.de/autoren/S/Anna_Sauerbrey/index", "full_name": "Anna Sauerbrey", "role_title": "Koordinatorin Meinung", "twitter_handle": "@AnnaSauerbrey", "article_count": 342, "primary_topics": "['Politik', 'Deutschland', 'USA']"
| # | author_url | full_name | role_title | biography | twitter_handle | article_count |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Frontpage Rankings objects from zeit.de. All fields typed and schema-versioned.
"scrape_timestamp": "2026-10-25T08:00:00Z", "position_rank": 1, "article_url": "https://www.zeit.de/politik/deutschland/2026-10/koalitionsverhandlungen", "block_name": "Aufmacher", "is_breaking_news": true, "is_zplus": false
| # | scrape_timestamp | position_rank | article_url | headline | block_name | is_breaking_news |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our zeit.de scraper handles consent walls, dynamic paywall detection, infinite scroll comments, and complex German character encoding, delivering clean text corpora for NLP and media analysis.
Extract headlines, subheadlines, lead paragraphs, and full body text from free articles, with precise paragraph separation and image link capture.
Accurately flag Z+ premium articles. Capture available preview text and metadata without triggering false positives or pipeline failures.
Paginate through thousands of user comments per article. Capture timestamps, upvotes, editor recommendations, and nested reply structures.
Scrape author directories and profile pages. Link journalists to their full article history, primary topics, and biographical metadata.
Categorise content by section (Politik, Wirtschaft, Gesellschaft) and extract granular topic tags attached to every article.
Automatic normalisation of umlauts and special characters (UTF-8) to ensure clean ingestion into your NLP models and databases.
Monitor the zeit.de homepage at high frequency to track article placement, breaking news banners, and editorial prioritisation over time.
Automated handling of the Pur-Abo consent banners via cookie injection and session management, ensuring uninterrupted access to public pages.
Run daily or hourly pipelines to build a comprehensive historical archive of German media discourse and publication trends.
Brief in. Clean data out.
Provide target sections, author lists, or historical date ranges. We map the required data fields.
We configure crawlers to handle the Pur-Abo consent wall, Z+ detection, and comment pagination.
Schema validation, character encoding checks, and paywall flag verification before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
European media sites deploy aggressive consent walls and dynamic paywalls. Here is how we maintain reliable extraction.
Zeit.de uses a strict Pur-Abo model requiring users to accept tracking or pay for a subscription. Our infrastructure programmatically manages consent cookies and session states to access public content without manual intervention.
Articles frequently switch from free to Z+ premium based on traffic velocity. We detect paywall states dynamically, capturing full text when available and gracefully degrading to preview text and metadata when gated.
Popular articles generate thousands of comments loaded via complex asynchronous requests. We trace the API calls to extract complete, deeply nested discussion threads rather than just the top ten visible comments.
Live blogs, interactive graphics, and standard articles use different DOM structures. We deploy content-type specific parsing logic to ensure clean text extraction regardless of the editorial format.
To avoid geo-blocking and receive the correct regional content variants, we route all requests through high-quality German residential proxies.
Track editorial tone and public reaction in comment sections regarding political events and corporate news.
Ingest high-quality, editorially reviewed German text to train large language models and translation engines.
Analyse topic frequency, author bias, and keyword prominence during election cycles or major policy shifts.
Publishers monitor zeit.de to understand which topics drive Z+ conversions and how long articles remain free.
PR firms and researchers track specific journalists, their output frequency, and their primary coverage areas.
Identify emerging narratives by tracking the velocity of new tags and section categorisations over time.
"Zeit.de represents a foundational corpus of German journalism and political discourse, but extracting it requires navigating consent walls and dynamic paywalls."
Most teams fail at scraping German news media because they cannot handle the Pur-Abo consent banners or reliably separate free content from Z+ premium articles. DataFlirt manages the session cookies, proxies, and selector maintenance so your data science teams receive clean text.
Everything supported by our zeit.de scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
We maintain persistent cookie jars to navigate the Pur-Abo consent walls, ensuring every request returns valid HTML rather than a redirect to the consent banner.
Traffic is routed exclusively through German residential IPs to mimic authentic local readership and avoid aggressive geo-blocking.
Instead of slow browser automation for comments, we reverse-engineer zeit.de internal APIs to extract discussion threads rapidly and reliably.
Data delivered to where your team already works — no new tooling required.
About zeit.de scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available news articles and comments is generally permissible for research and analysis, provided it complies with copyright laws regarding republication. DataFlirt extracts factual metadata, public text, and topic tags. We do not bypass cryptographic paywalls or extract private user data. Clients must ensure their downstream use cases comply with German copyright and GDPR regulations.
No. We do not hack or bypass paid subscription walls. If an article is flagged as Z+, we extract the headline, metadata, and whatever preview text is publicly visible before the paywall cuts off the content.
Our infrastructure automatically accepts the necessary tracking cookies required to view the free version of the site. We manage these session tokens programmatically across our proxy pool to maintain continuous access.
Yes. We can crawl the zeit.de sitemaps and search archives to extract articles published years ago, provided the URLs are still active and the content is not retroactively paywalled.
All pipelines enforce strict UTF-8 encoding. Umlauts (ä, ö, ü) and the eszett (ß) are correctly preserved in the final JSON, CSV, or Parquet output, preventing corruption in your NLP pipelines.
For frontpage monitoring, we can configure pipelines to run at sub-5-minute intervals, capturing breaking news banners and position changes in near real-time.
Yes. We extract complete comment threads, including nested replies, timestamps, upvote counts, and author names, which is highly valuable for sentiment analysis.
20-minute scoping call. Pilot dataset within the week. Production within two. From daily frontpage monitoring to massive historical NLP corpora, we build and manage the pipeline. Tell us your data requirements.