SYSTEM all green source tagesschau.de queue 12,844 articles p99 latency 218ms dataflirt.com · scraper/tagesschau-de
RUN . 31 active pipelines . tagesschau.de live

German news data,
structured for analysis.

We extract article text, broadcast metadata, Faktenfinder reports, and regional updates from Tagesschau.de. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Articles extracted
3.2K /day
Eilmeldungen tracked
84 /24h
Video metadata
412 /run
Active pipelines
31
Uptime
99.98%
Data Dictionary

Every field we extract from tagesschau.de

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for News Articles objects from tagesschau.de. All fields typed and schema-versioned.

article_idurlheadlinesubheadlineauthorpublish_dateupdate_datebody_texttagscategoryregional_tagrelated_links
news_articles
● 200 OK
"article_id": "ts-112458",
"url": "https://www.tagesschau.de/inland/gesellschaft/beispiel-artikel.html",
"headline": "Bundestag verabschiedet neues Gesetz",
"author": "ARD Hauptstadtsstudio",
"publish_date": "2026-05-12T14:30:00Z",
"category": "Inland",
"tags": "['Bundestag', 'Gesetz', 'Politik']",
"body_text": "Der Bundestag hat heute mit breiter Mehrheit..."
# article_idurlheadlinesubheadlineauthorpublish_date
1
2
3

Complete list of extractable fields for Breaking News (Eilmeldungen) objects from tagesschau.de. All fields typed and schema-versioned.

alert_idtimestampheadlinesummaryarticle_urlpush_notification_flagseverityactive_duration_minutessuperseded_by
breaking_news (eilmeldungen)
● 200 OK
"alert_id": "eil-88392",
"timestamp": "2026-05-12T08:15:22Z",
"headline": "Leitzins bleibt unverändert",
"summary": "Die EZB belässt den Leitzins auf dem aktuellen Niveau.",
"article_url": "https://www.tagesschau.de/wirtschaft/ezb-zinsentscheid.html",
"push_notification_flag": true,
"severity": "high"
# alert_idtimestampheadlinesummaryarticle_urlpush_notification_flag
1
2
3

Complete list of extractable fields for Faktenfinder objects from tagesschau.de. All fields typed and schema-versioned.

report_idurlclaimverdictauthorpublish_datebody_textreferences_listsocial_media_links
faktenfinder
● 200 OK
"report_id": "ff-9921",
"url": "https://www.tagesschau.de/faktenfinder/beispiel-faktencheck.html",
"claim": "Angebliches Zitat auf Social Media geteilt",
"verdict": "Falsch",
"author": "Faktenfinder-Team",
"publish_date": "2026-05-11T10:00:00Z",
"references_list": "['dpa', 'Statistisches Bundesamt']"
# report_idurlclaimverdictauthorpublish_date
1
2
3

Complete list of extractable fields for Video Broadcasts objects from tagesschau.de. All fields typed and schema-versioned.

broadcast_idtitleair_dateduration_secondsvideo_urlsubtitle_urltopics_coveredanchor_nameview_count_estimate
video_broadcasts
● 200 OK
"broadcast_id": "tv-2000-20260512",
"title": "tagesschau 20:00 Uhr",
"air_date": "2026-05-12T20:00:00Z",
"duration_seconds": 915,
"anchor_name": "Jens Riewa",
"topics_covered": "['Wetter', 'Politik', 'Sport']",
"video_url": "https://media.tagesschau.de/video/2026/0512/TV-20260512-2000.mp4"
# broadcast_idtitleair_dateduration_secondsvideo_urlsubtitle_url
1
2
3

Complete list of extractable fields for Regional News objects from tagesschau.de. All fields typed and schema-versioned.

region_idregion_nameard_stationheadlinepublish_dateurlbody_textlocal_tagscoordinate_data
regional_news
● 200 OK
"region_id": "reg-ndr-441",
"region_name": "Niedersachsen",
"ard_station": "NDR",
"headline": "Neue Windparks in der Nordsee geplant",
"publish_date": "2026-05-12T09:45:00Z",
"local_tags": "['Windenergie', 'Nordsee', 'Wirtschaft']",
"url": "https://www.tagesschau.de/inland/regional/niedersachsen/ndr-windparks.html"
# region_idregion_nameard_stationheadlinepublish_dateurl
1
2
3

Capabilities

Extract the German news cycle at scale

Our pipeline handles the complexities of public broadcasting data: high-frequency breaking news alerts, regional ARD syndication mapping, and video metadata extraction. Built for media monitoring and NLP training.

Full Text Article Extraction

Capture headlines, subheadlines, author bylines, and complete body text across all categories including Inland, Ausland, and Wirtschaft.

Eilmeldungen Tracking

High-frequency polling for breaking news alerts. Track duration, severity, and the final linked article for every push notification event.

Faktenfinder Corpus

Extract structured fact-checking reports including the original claim, the verdict, and all cited reference links.

Broadcast & Video Metadata

Collect air dates, durations, anchor names, and covered topic lists for the 20:00 Uhr broadcast and Tagesthemen.

Regional ARD Station Mapping

Identify and categorise local news syndicated from regional broadcasters like NDR, WDR, BR, and SWR.

Update Timestamp Normalisation

Track article revisions. Capture both the original publication date and the last modified timestamp to monitor editorial changes.

Tag & Taxonomy Scraping

Extract the internal tagging system to cluster articles by topic, geographic region, or political entity.

Election Data (Wahlen)

Scrape structured polling data, regional results, and coalition projections during state and federal election cycles.

Continuous Archiving

Run hourly or daily pipelines to build a permanent, queryable archive of German public news coverage.

// engagement pipeline

From front page to structured dataset

Brief in. Clean data out.

Define Scope
d 0

Select target categories, regional filters, or specific broadcast types. We map the required schema.

Pipeline Build
d 2–4

We configure Scrapy crawlers, handle rate limits, and normalise ARD regional syndication formats.

Validation & QA
d 4–6

Schema validation, null-rate checks on article bodies, and timestamp normalisation before deployment.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on your cadence.

Under the hood

Handling public broadcaster infrastructure

Tagesschau.de uses aggressive caching and specific regional syndication patterns. Here is how we ensure comprehensive data capture.

pipeline-monitor · tagesschau.de · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
High-frequency polling
Capturing Eilmeldungen without bans

Breaking news requires minute-by-minute polling. We distribute requests across German residential proxy pools to monitor the front page JSON feeds without triggering rate limits or IP blocks.

Syndication mapping
Normalising regional ARD content

Regional news is syndicated from different ARD stations (WDR, NDR, etc.) with varying DOM structures. Our selectors normalise these disparate layouts into a single, consistent regional news schema.

Revision tracking
Monitoring editorial changes

Articles are frequently updated as stories develop. We maintain a hash index of article bodies. When an update timestamp changes, we emit a new versioned record to track the editorial narrative over time.

Media metadata
Extracting video and audio context

Broadcast pages rely heavily on embedded media players. We intercept the backend API calls that populate these players to extract raw metadata, subtitle URLs, and topic segment timestamps.

Archive pagination
Navigating historical depth

Deep archiving requires traversing complex date-based navigation. Our crawlers systematically map the historical index to ensure zero missed articles during full retrospective scrapes.

Applications

Who uses Tagesschau data and how

Teams across industries use tagesschau.de data to build competitive products and smarter operations.

01
Media Monitoring & PR

Agencies track brand mentions, political entities, and corporate coverage across national and regional public broadcasting.

02
NLP & LLM Training

AI teams use high-quality, editorially vetted German text from Tagesschau to train language models and sentiment classifiers.

03
Political Analysis

Researchers monitor election coverage, topic frequency, and Faktenfinder reports to study political discourse and media bias.

04
Event Detection

Financial and supply chain analysts ingest Eilmeldungen via webhook for real-time alerts on geopolitical events and economic policy.

05
Disinformation Research

Think tanks analyse the Faktenfinder corpus to track the spread of specific claims and the effectiveness of public fact-checking.

06
Regional Trend Analysis

Marketers aggregate local ARD news data to understand regional economic developments and sentiment variations across Germany.

Why DataFlirt

"Tagesschau.de is the definitive record of German public broadcasting. Extracting its historical archive and real-time alerts requires a resilient pipeline."

News cycles move fast. Polling the front page for Eilmeldungen requires high-frequency scraping without triggering rate limits. DataFlirt handles the proxy rotation, video metadata extraction, and timestamp normalisation so your data science team can focus on NLP and sentiment modelling instead of maintaining brittle selectors.

Technical Spec

Tagesschau scraper technical capabilities

Everything supported by our tagesschau.de scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Article full text
Complete body text extraction with HTML formatting stripped or retained
Supported
Eilmeldungen polling
High-frequency checks for breaking news alerts with webhook delivery
Supported
Faktenfinder extraction
Structured capture of claims, verdicts, and reference lists
Supported
Regional ARD mapping
Normalisation of syndicated content from NDR, WDR, BR, SWR, etc.
Supported
Video metadata
Extraction of broadcast topics, anchor names, and subtitle URLs
Supported
Change detection (diffs)
Version tracking for articles that are updated post-publication
Supported
Election data (Wahlen)
Capture of polling figures and election results during active cycles
Supported
German residential proxies
DE-specific IP pools to bypass regional geo-blocking for certain live streams
Supported
ARD Mediathek user history
Personalised watch history and user account data requires authentication
Partial
Internal editorial CMS data
Pre-publication drafts and internal editorial notes are not publicly accessible
Partial
Infrastructure

Infrastructure powering the news pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles high-throughput article extraction. Playwright executes JavaScript to capture dynamic media player metadata and interactive election graphics.

Localised Proxy Infrastructure

We route requests through German residential IPs to ensure access to regionally restricted broadcasts and avoid rate limits during high-frequency polling.

Cloud-Native Orchestration

Pipelines run on AWS infrastructure. Airflow manages polling schedules for breaking news versus deep historical archive crawls. All state is stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested arrays for hierarchical article data
CSV
Flat file with typed columns for metadata analysis
XLS
Excel compatible format for editorial and PR teams
Parquet
Columnar format optimised for BigQuery and Snowflake
AWS S3
Direct bucket delivery for data lake integration
Webhook
HTTP POST delivery for real-time Eilmeldungen alerts
API
REST endpoints to query historical archive extracts
BigQuery
Streamed directly into your dataset with schema auto-detect
Snowflake
Stage and COPY INTO workflow for continuous ingestion
Postgres
Upsert into your existing relational schema
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About tagesschau.de scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Tagesschau.de legal?

Scraping publicly available news articles and metadata from Tagesschau.de is generally permissible for analysis and research purposes. DataFlirt extracts only public, non-authenticated data. We do not bypass DRM on video files or extract personal user data. Clients must ensure their downstream use complies with copyright law and ARD terms of service.

How fast can you deliver breaking news (Eilmeldungen)?

Our high-frequency pipelines poll the front page feeds every minute. We push breaking news alerts via Webhook within seconds of detection, allowing your systems to react in real time.

Can you extract historical news archives?

Yes. We can configure deep crawls to traverse the historical index, extracting articles, metadata, and Faktenfinder reports dating back to the limits of the public archive.

Do you extract the actual video files?

We extract comprehensive video metadata, subtitle URLs, and direct media stream URLs. We do not download or host the heavy MP4/HLS video files themselves, but we provide the links required for your systems to process them.

How do you handle articles that are updated after publication?

We monitor the update timestamps on target articles. If the content changes, we capture the new version and emit a diff record, allowing you to track editorial revisions over time.

Do you cover regional ARD stations?

Yes. Tagesschau.de syndicates content from regional broadcasters (NDR, WDR, SWR, etc.). We normalise these various layouts into a consistent regional news schema, including geographic tags.

$ dataflirt scope --new-project --source=tagesschau.de ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive dump for NLP training or a real-time webhook for breaking news alerts. We scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in news and media

Services

Data Extraction for Every Industry

View All Services →