SYSTEM all green source dn.se queue 12,841 URLs p99 latency 184ms dataflirt.com · scraper/dn-se
RUN · 41 active pipelines · dn.se live

Dagens Nyheter data,
at warehouse scale.

We extract article metadata, headlines, author profiles, category taxonomy, and public content from dn.se. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Articles extracted
4.2K /day
Metadata updates
18.1K /24h
Author profiles
840 /run
Active pipelines
41
Uptime
99.94%
Data Dictionary

Every field we extract from dn.se

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Article Metadata objects from dn.se. All fields typed and schema-versioned.

article_idurlheadlinesubheadlineauthor_namepublished_atupdated_atcategorytagspaywall_statusword_count
article_metadata
● 200 OK
"article_id": "dn-1234567",
"headline": "Riksbanken sänker styrräntan",
"author_name": "Anna Andersson",
"published_at": "2026-05-12T08:30:00Z",
"category": "Ekonomi",
"paywall_status": true,
"word_count": 842
# article_idurlheadlinesubheadlineauthor_namepublished_at
1
2
3

Complete list of extractable fields for Author Profiles objects from dn.se. All fields typed and schema-versioned.

author_idnameprofile_urltwitter_handleemailrolearticle_countlatest_article_datebiography
author_profiles
● 200 OK
"author_id": "auth-8921",
"name": "Anna Andersson",
"profile_url": "https://www.dn.se/av/anna-andersson/",
"role": "Ekonomireporter",
"article_count": 412,
"latest_article_date": "2026-05-12"
# author_idnameprofile_urltwitter_handleemailrole
1
2
3

Complete list of extractable fields for Frontpage Placements objects from dn.se. All fields typed and schema-versioned.

scrape_timepositionsectionheadlineurlis_breakinghas_videoimage_urlrelated_links
frontpage_placements
● 200 OK
"scrape_time": "2026-05-12T09:00:00Z",
"position": 1,
"section": "Nyheter",
"headline": "Riksbanken sänker styrräntan",
"is_breaking": true,
"has_video": false
# scrape_timepositionsectionheadlineurlis_breaking
1
2
3

Complete list of extractable fields for Taxonomy & Categories objects from dn.se. All fields typed and schema-versioned.

category_idnameparent_categoryurl_slugarticle_count_24htrending_scorelast_updateddescriptionrss_feed_url
taxonomy_& categories
● 200 OK
"category_id": "cat-ekonomi",
"name": "Ekonomi",
"parent_category": "Nyheter",
"url_slug": "/ekonomi/",
"article_count_24h": 45,
"last_updated": "2026-05-12T08:45:00Z"
# category_idnameparent_categoryurl_slugarticle_count_24htrending_score
1
2
3

Complete list of extractable fields for Media & Images objects from dn.se. All fields typed and schema-versioned.

image_idarticle_urlimage_urlcaptionphotographerwidthheightalt_textformat
media_& images
● 200 OK
"image_id": "img-998213",
"article_url": "https://www.dn.se/ekonomi/riksbanken-sanker/",
"image_url": "https://images.dn.se/v1/image/123.jpg",
"caption": "Riksbankschefen under presskonferensen.",
"photographer": "Lars Larsson / TT",
"width": 1200
# image_idarticle_urlimage_urlcaptionphotographerwidth
1
2
3

Capabilities

Swedish media data extracted at scale

Our dn.se scraper handles dynamic frontpage layouts, strict bot protection, and complex article metadata structures. We extract clean text and metadata across the entire publication archive.

Article Metadata Extraction

Capture headlines, subheadlines, publication timestamps, and update histories across all news sections.

Frontpage Position Tracking

Monitor layout changes, breaking news banners, and article positioning on the main dn.se index over time.

Author & Byline Data

Extract journalist profiles, contact information, role descriptions, and historical publication records.

Paywall Status Detection

Identify premium versus open articles to optimise downstream processing and content aggregation logic.

Category & Tag Taxonomy

Map the entire site structure, extracting section hierarchies and article tags for precise topical filtering.

Historical Archive Scraping

Traverse date-based archives to build longitudinal datasets of Swedish media coverage over past decades.

Real-Time News Monitoring

Poll RSS feeds and section indexes at high frequency to capture breaking news within seconds of publication.

Multi-Format Delivery

Push structured data to your warehouse via S3, BigQuery, or PostgreSQL in JSON, CSV, or Parquet formats.

Scheduled + Streaming Modes

Configure pipelines for daily batch exports or continuous real-time extraction with change detection.

// engagement pipeline

From target URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide section URLs, author lists, or historical date ranges. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, proxy rotation, and session management for dn.se infrastructure.

Validation & QA
d 4–6

Schema validation, null-rate checks, and text encoding verification before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Handling modern news site architecture

Media organisations deploy aggressive caching and bot protection. Here is how we maintain reliable extraction pipelines for dn.se.

pipeline-monitor · dn.se · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
Regional proxy rotation

News sites often restrict access based on geography or flag data centre IPs. We route requests through Swedish residential proxies to ensure consistent access and avoid rate limits.

JavaScript rendering
Playwright for dynamic content

Modern news frontpages load content asynchronously. We deploy Playwright to execute JavaScript, ensuring we capture lazy-loaded articles and dynamic breaking news banners.

Paywall handling
Clear premium content delineation

Dagens Nyheter places significant content behind a strict paywall. Our pipeline detects paywall markers in the DOM, extracting available metadata and public lead paragraphs without triggering authentication errors.

Schema stability
Resilient selector chains

Editorial teams frequently alter layouts for major events. We use multiple XPath and CSS fallback chains to ensure data extraction continues even when the DOM structure shifts.

Change detection
Optimised article diffing

We hash article content and metadata. Subsequent runs only emit records when an article is updated or newly published, preventing duplicate data in your warehouse.

Applications

Who uses Dagens Nyheter data

Teams across industries use dn.se data to build competitive products and smarter operations.

01
Media Monitoring

PR agencies and corporate communications teams track brand mentions, sentiment, and crisis development in real time.

02
NLP & LLM Training

Machine learning teams ingest high-quality Swedish editorial text to train language models and sentiment classifiers.

03
Sentiment Analysis

Financial analysts process economic news and editorial opinions to gauge market sentiment and predict trends.

04
Author Network Analysis

Researchers map journalist beats, publication frequency, and topic specialisation across the media landscape.

05
Competitive Intelligence

Competing publishers analyse dn.se publication velocity, frontpage curation strategies, and paywall conversion tactics.

06
Academic Research

Political scientists compile longitudinal datasets of news coverage to study media bias, framing, and agenda setting.

Why DataFlirt

"Dagens Nyheter represents the historical and contemporary pulse of Swedish media, but extracting structured text requires navigating strict paywalls and dynamic DOM structures."

Most teams underestimate the investment required. Reliable news scraping requires residential proxies, full JavaScript rendering, and daily selector maintenance to adapt to editorial layout changes. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.

Technical Spec

Dagens Nyheter scraper technical specifications

Everything supported by our dn.se scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions required for dynamic frontpage elements and lazy-loaded images
Supported
CAPTCHA bypass
Automated solver integration for rate-limit challenges
Supported
Residential proxy rotation
Swedish ISP IPs rotated per request to prevent blocking
Supported
Article text extraction
Publicly available article text, lead paragraphs, and subheadlines
Supported
Author metadata
Byline extraction, role identification, and profile linking
Supported
Paywall detection
Boolean flag indicating if the article requires a premium subscription
Supported
Change detection
Hash-based diffing to emit only new or updated articles
Supported
Webhook delivery
HTTP POST per article for real-time media monitoring alerts
Supported
Premium subscriber content
Full text of paywalled articles requires client-provided authentication credentials
Partial
Authenticated user comments
Extraction of reader comments hidden behind the login wall
Partial
Infrastructure

Infrastructure powering the extraction pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and retry logic. Playwright manages JavaScript execution and dynamic DOM hydration for complex editorial layouts.

Residential Proxy Infrastructure

We route traffic through Swedish residential proxies to maintain geographic relevance and avoid data centre IP bans enforced by media firewalls.

Cloud-Native Orchestration

Pipelines execute on AWS Lambda and ECS. Airflow manages scheduling and dependency trees. Postgres stores state and deduplication hashes.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested structures per run
CSV
Flat file with typed columns for simple ingestion
XLS
Excel compatible format for analyst teams
Parquet
Columnar format optimised for BigQuery and Snowflake
AWS S3
Direct delivery to your cloud storage buckets
Webhook
HTTP POST payloads for real-time article alerts
API
REST endpoints to query extracted historical data
PostgreSQL
Direct database upserts with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About dn.se scraping, legality, and pipeline operations.

Ask us directly →
Can you scrape full articles behind the dn.se paywall?

We extract all publicly available metadata, headlines, and lead paragraphs. Accessing full premium text requires valid subscriber credentials, which we do not provide. If you supply authenticated session cookies, we can configure the pipeline to extract full text for your internal use.

How quickly can you detect breaking news?

Our real-time pipelines can poll dn.se RSS feeds, section indexes, and the frontpage at sub-minute intervals, delivering new article payloads via Webhook almost immediately after publication.

Do you extract historical article archives?

Yes. We can traverse the dn.se sitemap and historical date archives to construct comprehensive datasets of past media coverage, subject to the site's historical availability.

How do you handle changes to the website layout?

We utilise resilient selector chains with multiple fallbacks. Our telemetry alerts us to schema drift or null-rate spikes, allowing our engineers to update selectors before data quality degrades.

Is it legal to scrape news websites?

Scraping public factual data, headlines, and metadata is generally permissible. However, reproducing full copyrighted article text for commercial redistribution may violate copyright law. DataFlirt extracts data for internal analysis, NLP training, and monitoring. Clients must ensure their specific use case complies with local copyright regulations.

Can you track how long an article stays on the frontpage?

Yes. By scheduling high-frequency scrapes of the index page, we build a time-series dataset tracking article position, section placement, and total duration on the frontpage.

$ dataflirt scope --new-project --source=dn.se ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive dump or a real-time news monitoring feed. We scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in news and media

Services

Data Extraction for Every Industry

View All Services →