We extract article metadata, headlines, author profiles, category taxonomy, and public content from dn.se. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Article Metadata objects from dn.se. All fields typed and schema-versioned.
"article_id": "dn-1234567", "headline": "Riksbanken sänker styrräntan", "author_name": "Anna Andersson", "published_at": "2026-05-12T08:30:00Z", "category": "Ekonomi", "paywall_status": true, "word_count": 842
| # | article_id | url | headline | subheadline | author_name | published_at |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Author Profiles objects from dn.se. All fields typed and schema-versioned.
"author_id": "auth-8921", "name": "Anna Andersson", "profile_url": "https://www.dn.se/av/anna-andersson/", "role": "Ekonomireporter", "article_count": 412, "latest_article_date": "2026-05-12"
| # | author_id | name | profile_url | twitter_handle | role | |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Frontpage Placements objects from dn.se. All fields typed and schema-versioned.
"scrape_time": "2026-05-12T09:00:00Z", "position": 1, "section": "Nyheter", "headline": "Riksbanken sänker styrräntan", "is_breaking": true, "has_video": false
| # | scrape_time | position | section | headline | url | is_breaking |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Taxonomy & Categories objects from dn.se. All fields typed and schema-versioned.
"category_id": "cat-ekonomi", "name": "Ekonomi", "parent_category": "Nyheter", "url_slug": "/ekonomi/", "article_count_24h": 45, "last_updated": "2026-05-12T08:45:00Z"
| # | category_id | name | parent_category | url_slug | article_count_24h | trending_score |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Media & Images objects from dn.se. All fields typed and schema-versioned.
"image_id": "img-998213", "article_url": "https://www.dn.se/ekonomi/riksbanken-sanker/", "image_url": "https://images.dn.se/v1/image/123.jpg", "caption": "Riksbankschefen under presskonferensen.", "photographer": "Lars Larsson / TT", "width": 1200
| # | image_id | article_url | image_url | caption | photographer | width |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our dn.se scraper handles dynamic frontpage layouts, strict bot protection, and complex article metadata structures. We extract clean text and metadata across the entire publication archive.
Capture headlines, subheadlines, publication timestamps, and update histories across all news sections.
Monitor layout changes, breaking news banners, and article positioning on the main dn.se index over time.
Extract journalist profiles, contact information, role descriptions, and historical publication records.
Identify premium versus open articles to optimise downstream processing and content aggregation logic.
Map the entire site structure, extracting section hierarchies and article tags for precise topical filtering.
Traverse date-based archives to build longitudinal datasets of Swedish media coverage over past decades.
Poll RSS feeds and section indexes at high frequency to capture breaking news within seconds of publication.
Push structured data to your warehouse via S3, BigQuery, or PostgreSQL in JSON, CSV, or Parquet formats.
Configure pipelines for daily batch exports or continuous real-time extraction with change detection.
Brief in. Clean data out.
Provide section URLs, author lists, or historical date ranges. We design the extraction schema together.
We configure Scrapy crawlers, proxy rotation, and session management for dn.se infrastructure.
Schema validation, null-rate checks, and text encoding verification before full launch.
JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Media organisations deploy aggressive caching and bot protection. Here is how we maintain reliable extraction pipelines for dn.se.
News sites often restrict access based on geography or flag data centre IPs. We route requests through Swedish residential proxies to ensure consistent access and avoid rate limits.
Modern news frontpages load content asynchronously. We deploy Playwright to execute JavaScript, ensuring we capture lazy-loaded articles and dynamic breaking news banners.
Dagens Nyheter places significant content behind a strict paywall. Our pipeline detects paywall markers in the DOM, extracting available metadata and public lead paragraphs without triggering authentication errors.
Editorial teams frequently alter layouts for major events. We use multiple XPath and CSS fallback chains to ensure data extraction continues even when the DOM structure shifts.
We hash article content and metadata. Subsequent runs only emit records when an article is updated or newly published, preventing duplicate data in your warehouse.
PR agencies and corporate communications teams track brand mentions, sentiment, and crisis development in real time.
Machine learning teams ingest high-quality Swedish editorial text to train language models and sentiment classifiers.
Financial analysts process economic news and editorial opinions to gauge market sentiment and predict trends.
Researchers map journalist beats, publication frequency, and topic specialisation across the media landscape.
Competing publishers analyse dn.se publication velocity, frontpage curation strategies, and paywall conversion tactics.
Political scientists compile longitudinal datasets of news coverage to study media bias, framing, and agenda setting.
"Dagens Nyheter represents the historical and contemporary pulse of Swedish media, but extracting structured text requires navigating strict paywalls and dynamic DOM structures."
Most teams underestimate the investment required. Reliable news scraping requires residential proxies, full JavaScript rendering, and daily selector maintenance to adapt to editorial layout changes. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.
Everything supported by our dn.se scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and retry logic. Playwright manages JavaScript execution and dynamic DOM hydration for complex editorial layouts.
We route traffic through Swedish residential proxies to maintain geographic relevance and avoid data centre IP bans enforced by media firewalls.
Pipelines execute on AWS Lambda and ECS. Airflow manages scheduling and dependency trees. Postgres stores state and deduplication hashes.
Data delivered to where your team already works — no new tooling required.
We extract all publicly available metadata, headlines, and lead paragraphs. Accessing full premium text requires valid subscriber credentials, which we do not provide. If you supply authenticated session cookies, we can configure the pipeline to extract full text for your internal use.
Our real-time pipelines can poll dn.se RSS feeds, section indexes, and the frontpage at sub-minute intervals, delivering new article payloads via Webhook almost immediately after publication.
Yes. We can traverse the dn.se sitemap and historical date archives to construct comprehensive datasets of past media coverage, subject to the site's historical availability.
We utilise resilient selector chains with multiple fallbacks. Our telemetry alerts us to schema drift or null-rate spikes, allowing our engineers to update selectors before data quality degrades.
Scraping public factual data, headlines, and metadata is generally permissible. However, reproducing full copyrighted article text for commercial redistribution may violate copyright law. DataFlirt extracts data for internal analysis, NLP training, and monitoring. Clients must ensure their specific use case complies with local copyright regulations.
Yes. By scheduling high-frequency scrapes of the index page, we build a time-series dataset tracking article position, section placement, and total duration on the frontpage.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive dump or a real-time news monitoring feed. We scope, build, and operate the pipeline. Tell us what you need.