We extract full article text, author profiles, comment sections, and publication metadata from Politiken. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Articles objects from politiken.dk. All fields typed and schema-versioned.
"article_id": "9384712", "headline": "Nye klimamål kræver massive investeringer", "author_name": "Lars Jensen", "published_at": "2026-03-14T08:30:00Z", "section": "Indland", "paywalled": true, "tags": "['Klima', 'Politik', 'Økonomi']"
| # | article_id | url | headline | subheadline | author_name | published_at |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Authors objects from politiken.dk. All fields typed and schema-versioned.
"author_id": "auth_8472", "name": "Lars Jensen", "role": "Klimakorrespondent", "twitter_handle": "@larsjensen_pol", "article_count": 342, "latest_article_date": "2026-03-14T08:30:00Z"
| # | author_id | name | profile_url | role | twitter_handle | |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Comments objects from politiken.dk. All fields typed and schema-versioned.
"comment_id": "c_993821", "article_id": "9384712", "user_name": "Mette Nielsen", "timestamp": "2026-03-14T09:15:22Z", "text": "Dette er et vigtigt skridt for fremtiden.", "likes": 42, "replies_count": 3
| # | comment_id | article_id | user_name | timestamp | text | likes |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Sections objects from politiken.dk. All fields typed and schema-versioned.
"section_id": "sec_indland", "section_name": "Indland", "parent_section": "Nyheder", "url": "https://politiken.dk/indland/", "top_headline": "Nye klimamål kræver massive investeringer", "last_updated": "2026-03-14T10:05:00Z"
| # | section_id | section_name | parent_section | url | top_headline | article_count |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Search Results objects from politiken.dk. All fields typed and schema-versioned.
"query": "klima investering", "position": 1, "article_url": "https://politiken.dk/indland/art9384712/", "headline": "Nye klimamål kræver massive investeringer", "date": "2026-03-14", "author": "Lars Jensen"
| # | query | position | article_url | headline | snippet | date |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our Politiken scraper handles dynamic pagination, section hierarchies, and paywall detection, delivering clean, structured journalism data for your models and monitoring tools.
Extract headlines, subheadlines, bylines, publication dates, and body text from both free and publicly available summary sections.
Track journalist output, role designations, social handles, and publication frequency across all Politiken sections.
Automatically detect and flag premium (Plus) content versus free-to-read articles for accurate dataset filtering.
Extract user comments, timestamps, like counts, and reply threads from article discussion sections.
Capture and structure internal Politiken tags, keywords, and category assignments for precise topical clustering.
Monitor articles for post-publication edits by tracking the updated_at timestamp and comparing body text diffs.
Map the entire site structure from top-level categories down to specific regional or topical sub-sections.
Configure high-frequency polls on the front page and RSS feeds to capture breaking news within minutes of publication.
Run daily or weekly batch jobs to build historical archives of Danish news coverage and opinion pieces.
Brief in. Clean data out.
Provide target sections, author lists, or historical date ranges. We design the extraction schema together.
We configure Scrapy / Playwright crawlers, proxy rotation, and dynamic content handling for politiken.dk.
Schema validation, null-rate checks, and paywall detection logic verification before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
News sites employ dynamic rendering and strict rate limits. Here is how we maintain data flow.
Politiken loads comments and certain dynamic widgets via JavaScript after the initial page load. We use Playwright to execute these scripts and capture the full DOM state.
We identify Politiken Plus articles via meta tags and DOM structure, ensuring your dataset accurately reflects what is publicly readable versus paywalled.
To prevent IP bans from heavy scraping, we distribute requests across a pool of Danish residential proxies, maintaining realistic request rates.
Media sites frequently run A/B tests on article layouts. Our extraction logic uses multiple fallback selectors, relying on JSON-LD structured data where available.
News articles evolve. We track publication and modification timestamps, capturing diffs when an article is significantly updated post-publish.
PR firms and corporate communications teams track brand mentions, executive coverage, and crisis developments in real time.
Machine learning teams use structured Danish text corpora to train and fine-tune language models.
Financial analysts and political researchers gauge public sentiment by analysing opinion pieces and comment sections.
Media analysts track journalist output, topic focus, and bias across different publications.
Rival media organisations monitor Politiken's publishing cadence, section popularity, and paywall strategies.
Institutions build searchable historical archives of Danish news for academic research and regulatory compliance.
"Politiken represents the pulse of Danish journalism, but extracting structured historical archives requires navigating dynamic frontend frameworks and paywalls."
Most teams underestimate the investment required: reliable media scraping requires residential proxies, full JavaScript rendering for comment sections, and daily selector maintenance. DataFlirt absorbs that complexity so your engineers can focus on the analysis, not the infrastructure.
Everything supported by our politiken.dk scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering for dynamic comments and lazy-loaded images.
We maintain pools of residential ISP proxies across DK regions to prevent rate limiting and IP blocks during high-volume archival scrapes.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. State is stored in Postgres.
Data delivered to where your team already works — no new tooling required.
About politiken.dk scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available information is generally permissible. DataFlirt targets only public, non-authenticated article and metadata. We do not extract personal data beyond public bylines, circumvent authentication walls, or violate GDPR. Clients should review Politiken's ToS and consult legal counsel for specific use cases.
Our crawlers detect paywall metadata and DOM flags. For paywalled articles, we extract the publicly available headline, subheadline, author, and metadata, flagging the record as 'paywalled: true'. We do not bypass the paywall to access premium body text.
Yes. We can configure pipelines to crawl historical sitemaps and search archives, extracting decades of available news articles subject to the same paywall restrictions.
For breaking news monitoring, we can configure high-frequency polling on specific sections or RSS feeds, achieving sub-5-minute latency for new publications.
Yes. We use Playwright to render the dynamic comment sections on articles, extracting user names, timestamps, comment text, and engagement metrics.
Absolutely. We provide a sample run of up to 500 articles as part of the pre-engagement scoping process so you can validate schema fit and data quality.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive dump or a continuous news monitoring feed — we scope, build, and operate the pipeline. Tell us what you need.