SYSTEM all green source nos.nl queue 3,492 pages p99 latency 118ms dataflirt.com · scraper/nos-nl
RUN · 42 active pipelines · nos.nl live

nos.nl data,
at warehouse scale.

We extract breaking news, live blog updates, match statistics, and Teletekst records from nos.nl. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Articles extracted
4,291 /day
Live blog updates
12,405 /24h
Teletekst records
812 /run
Active pipelines
42
Uptime
99.98%
Data Dictionary

Every field we extract from nos.nl

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for News Articles objects from nos.nl. All fields typed and schema-versioned.

article_idurltitlesummarybody_textpublished_atmodified_atauthorscategoriestagsimage_urlsrelated_articles
news_articles
● 200 OK
"article_id": "2501934",
"title": "Kabinet presenteert nieuwe klimaatplannen",
"published_at": "2026-05-12T09:14:00Z",
"categories": "['Politiek', 'Binnenland']",
"authors": "['NOS Nieuws']",
"tags": "['Klimaat', 'Den Haag']"
# article_idurltitlesummarybody_textpublished_at
1
2
3

Complete list of extractable fields for Live Blogs objects from nos.nl. All fields typed and schema-versioned.

blog_idtitlestatuslast_updatepost_idpost_timestamppost_titlepost_contentpost_mediais_pinnedsource_url
live_blogs
● 200 OK
"blog_id": "liveblog-849201",
"status": "active",
"post_id": "post-9921",
"post_timestamp": "2026-05-12T10:05:22Z",
"post_title": "Update vanuit de rechtbank",
"is_pinned": false
# blog_idtitlestatuslast_updatepost_idpost_timestamp
1
2
3

Complete list of extractable fields for NOS Sport objects from nos.nl. All fields typed and schema-versioned.

match_idsport_typetournamenthome_teamaway_teamscore_homescore_awaystatusmatch_datematch_summaryminute_by_minute
nos_sport
● 200 OK
"match_id": "voetbal-eredivisie-3921",
"sport_type": "Voetbal",
"tournament": "Eredivisie",
"home_team": "Ajax",
"away_team": "Feyenoord",
"status": "finished"
# match_idsport_typetournamenthome_teamaway_teamscore_home
1
2
3

Complete list of extractable fields for Teletekst objects from nos.nl. All fields typed and schema-versioned.

page_numbersub_pagecontent_textcategorylast_updatedpage_urllinked_pagescolor_codesraw_html
teletekst
● 200 OK
"page_number": 101,
"sub_page": 1,
"category": "Nieuws",
"last_updated": "2026-05-12T10:15:00Z",
"linked_pages": "[102, 103, 104]",
"page_url": "https://nos.nl/teletekst#101"
# page_numbersub_pagecontent_textcategorylast_updatedpage_url
1
2
3

Complete list of extractable fields for Video Metadata objects from nos.nl. All fields typed and schema-versioned.

video_idtitledescriptionduration_secondsbroadcast_dateprogram_namethumbnail_urlview_countformat_typeembed_url
video_metadata
● 200 OK
"video_id": "video-39201",
"title": "Samenvatting Formule 1 Grand Prix",
"duration_seconds": 345,
"program_name": "NOS Sport",
"broadcast_date": "2026-05-11T16:30:00Z",
"format_type": "highlight"
# video_idtitledescriptionduration_secondsbroadcast_dateprogram_name
1
2
3

Capabilities

Complete coverage of the NOS platform

Our nos.nl scraper handles every data layer: static articles, real-time live blogs, legacy Teletekst structures, and dynamic sports scoreboards. Built with automated polling and schema normalisation.

Article Extraction

Capture headlines, body text, publication timestamps, authors, and category tags from all NOS news sections.

Live Blog Polling

Monitor breaking news live blogs with sub-minute polling. Extract individual posts, timestamps, and embedded media.

Teletekst Parsing

Extract structured text and navigation links from the NOS Teletekst web interface, converting legacy formats into clean JSON.

Sports Results & Data

Track match scores, tournament standings, and minute-by-minute updates across football, cycling, Formula 1, and more.

Video Metadata

Collect broadcast metadata, descriptions, durations, and program associations from NOS video players.

Taxonomy Mapping

Extract and map internal NOS categories, regional tags, and topic clusters for accurate content classification.

Change Detection

Monitor articles for post-publication edits. Capture modification timestamps and text diffs across updates.

Regional News Integration

Aggregate regional broadcaster feeds surfaced on the NOS platform, tagged by province and municipality.

High-Frequency Cadence

Configure pipelines for daily archival sweeps or real-time streaming for breaking news alerts.

// engagement pipeline

From target URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target sections, live blog URLs, or Teletekst pages. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, polling frequencies, and API interceptors for nos.nl dynamic content.

Validation & QA
d 4–6

Schema validation, null-rate checks, and duplicate detection before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Handling dynamic news infrastructure

Modern news sites use a mix of static generation and real-time websockets. Here is how we extract structured data reliably.

pipeline-monitor · nos.nl · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Live blog extraction
Real-time XHR interception

NOS live blogs push updates via background API calls. We intercept these XHR requests directly, parsing the raw JSON payloads rather than scraping the DOM, ensuring zero latency on breaking news.

Teletekst parsing
Legacy format normalisation

Teletekst relies on fixed-width character grids. Our parsers reconstruct these grids into structured text, mapping colour codes to semantic meaning and extracting valid page links.

Article updates
Diff tracking for edited content

News articles are frequently updated after initial publication. We maintain a hash index of article bodies, emitting new records only when the modified_at timestamp or content hash changes.

Media extraction
Video and image metadata capture

We extract high-resolution image URLs and video player metadata from the NOS content management system, linking media assets directly to their parent articles.

Monitoring
Uptime and schema drift alerts

News layouts change during major events. We monitor selector success rates in real time, alerting our engineers to DOM shifts before they impact your data delivery.

Applications

Who uses NOS data and how

Teams across industries use nos.nl data to build competitive products and smarter operations.

01
Media Monitoring

PR agencies and corporate communications teams track brand mentions, political developments, and public sentiment across Dutch national news.

02
LLM Training & NLP

Machine learning teams ingest high-quality Dutch language corpora from NOS articles to train and fine-tune regional language models.

03
Sports Analytics

Analysts track Eredivisie scores, match statistics, and sports reporting for betting models and historical performance analysis.

04
Event Detection

Financial institutions monitor breaking news and live blogs for macro-economic events, natural disasters, or political shifts affecting European markets.

05
Archival Research

Academic researchers build historical datasets of Dutch public broadcasting output for sociological and political science studies.

06
Competitor News Tracking

Publishers monitor NOS editorial decisions, publication timing, and topic coverage to benchmark their own newsroom performance.

Why DataFlirt

"The NOS platform holds the most authoritative real-time record of Dutch news and sports events, requiring sub-minute polling for live blogs."

Tracking breaking news across nos.nl requires handling heavily dynamic live blogs, legacy Teletekst structures, and rapid article updates. DataFlirt manages the polling frequency, deduplication, and schema normalisation so your data science teams receive clean, queryable JSON feeds rather than raw HTML.

Technical Spec

NOS scraper technical capabilities

Everything supported by our nos.nl scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Live blog XHR parsing
Direct interception of background API calls for sub-minute updates
Supported
Teletekst grid parsing
Conversion of fixed-width character grids into structured text
Supported
Article body extraction
Clean text extraction stripping out ad units and navigation elements
Supported
Image metadata
Capture of high-resolution asset URLs and alt-text descriptions
Supported
Category taxonomy
Mapping of internal NOS topic tags and regional identifiers
Supported
Change detection (diffs)
Hash-based diff tracking for post-publication article edits
Supported
Archive traversal
Pagination through historical news categories and search results
Supported
Video transcripts
Extraction of closed captions when available on the video player
Supported
NPO ID authenticated streams
Gated premium content requiring NPO Plus user credentials
Partial
User comments
Extraction of user interaction data (NOS does not host native comments)
Partial
Infrastructure

Infrastructure powering the NOS pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles high-throughput article extraction. Playwright intercepts XHR requests for live blogs and handles dynamic video player rendering.

High-Frequency Polling

Redis-backed deduplication ensures that sub-minute polling on breaking news live blogs only emits new posts, preventing downstream data bloat.

Cloud-Native Orchestration

Pipelines run on AWS Lambda for burst scaling during major news events. Airflow manages dependencies and delivery schedules.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested arrays for live blog posts
CSV
Flat file format for static article metadata
XLS
Excel compatible exports for editorial teams
Parquet
Columnar format optimized for analytics workloads
AWS S3
Direct delivery to your cloud storage buckets
Webhook
HTTP POST for real-time breaking news alerts
API
REST endpoints to query historical news datasets
BigQuery
Direct streaming inserts into Google Cloud
Snowflake
Automated staging and ingestion workflows
Postgres
Relational upserts with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About nos.nl scraping, legality, and pipeline operations.

Ask us directly →
Is scraping nos.nl legal?

Scraping publicly available news articles, sports data, and Teletekst pages is generally permissible under web scraping precedents, provided it does not disrupt the host servers. DataFlirt targets only public, unauthenticated content. Clients must ensure their downstream use cases comply with copyright laws regarding journalistic content.

How fast can you deliver breaking news updates?

For targeted live blogs, we configure polling intervals down to 30 seconds. New posts are extracted, deduplicated, and pushed via Webhook within milliseconds of appearing on the NOS platform.

Can you extract historical news articles?

Yes. We can traverse NOS category archives and search results to build historical datasets spanning back to the limits of their public index.

How do you handle Teletekst formatting?

Our parsers map the fixed-width character grids of the Teletekst web interface into structured JSON. We capture page numbers, sub-pages, text content, and valid navigation links.

Do you extract video content?

We extract video metadata including titles, descriptions, broadcast dates, and thumbnails. We do not download or host the raw MP4 video files.

What is the minimum viable engagement?

Engagements typically start with a defined daily extraction volume or specific category monitoring. Contact us with your target sections and latency requirements for a scoped quote.

Can I request a sample dataset?

Yes. We provide sample JSON extracts of news articles, live blogs, and Teletekst pages during the scoping phase so your engineering team can validate the schema.

$ dataflirt scope --new-project --source=nos.nl ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive of Dutch news or real-time webhooks for live blog updates, we build and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in news and media

Services

Data Extraction for Every Industry

View All Services →