We extract national news, regional reporting, SVT Play metadata, and broadcast schedules from svt.se. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for News Articles objects from svt.se. All fields typed and schema-versioned.
"article_id": "39482104", "title": "Riksbanken sänker styrräntan", "sub_headline": "Historisk sänkning väntas påverka bolånen", "author": "Anna Andersson", "published_at": "2026-05-12T09:30:00Z", "category": "Ekonomi", "regional_tag": "Riks"
| # | article_id | url | title | sub_headline | author | published_at |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for SVT Play Metadata objects from svt.se. All fields typed and schema-versioned.
"video_id": "8x9a2b1", "program_name": "Uppdrag granskning", "title": "Spelet om skogen", "duration_seconds": 3480, "broadcast_date": "2026-05-11T20:00:00Z", "expires_at": "2026-11-11T23:59:59Z", "manifest_url": "https://svt-vod.akamaized.net/.../master.m3u8", "subtitle_url": "https://svt-vod.akamaized.net/.../sv.vtt"
| # | video_id | title | description | program_name | season | episode |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Broadcast Schedules objects from svt.se. All fields typed and schema-versioned.
"channel": "SVT1", "program_title": "Rapport", "start_time": "2026-05-12T19:30:00Z", "end_time": "2026-05-12T20:00:00Z", "is_live": true, "genre": "Nyheter", "rerun_status": false
| # | channel | program_title | start_time | end_time | is_live | description |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Authors & Journalists objects from svt.se. All fields typed and schema-versioned.
"name": "Lars Larsson", "role": "Inrikespolitisk kommentator", "email": "lars.larsson@svt.se", "twitter_handle": "@larslarssonsvt", "article_count": 412, "profile_url": "https://www.svt.se/nyheter/profiler/lars-larsson"
| # | author_id | name | role | twitter_handle | article_count | |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Regional News objects from svt.se. All fields typed and schema-versioned.
"region_name": "Västerbotten", "article_id": "39482188", "title": "Nytt sjukhusbygge försenas", "municipality": "Umeå", "published_at": "2026-05-12T10:15:00Z", "local_video_attached": true
| # | region_name | article_id | title | published_at | municipality | tags |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our SVT scraper handles every layer of the platform: text journalism, regional filtering, SVT Play video metadata, and dynamic schedule feeds. We manage the Next.js hydration and Swedish geo-routing internally.
Title, sub-headline, body text, publication timestamps, and author bylines scraped across national and regional news sections.
Extract program names, episode descriptions, broadcast dates, expiration windows, and HLS manifest URLs for video streams.
Capture daily television schedules for SVT1, SVT2, SVT Barn, and Kunskapskanalen with precise start and end timestamps.
Locate and download VTT or SRT subtitle files associated with SVT Play content for NLP and accessibility analysis.
Isolate news by specific Swedish regions and municipalities, tracking local reporting trends and priority levels.
Track output by specific journalists, capturing their contact information, roles, and historical article catalogues.
Map how SVT links internal articles together to understand editorial focus and narrative clustering.
Monitor the SVT homepage and RSS feeds at high frequency to capture breaking news alerts and push notifications.
Run one-off historical archive exports or configure continuous pipelines at hourly cadences with change-detection diffing.
Brief in. Clean data out.
Provide category URLs, specific programs on SVT Play, or regional news sections. We design the extraction schema together.
We configure Scrapy and Playwright crawlers, Swedish proxy routing, and Next.js state extraction logic.
Schema validation, null-rate checks, timestamp normalisation, and sample data reviews before full launch.
JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Extracting data from modern public broadcasters requires handling geo-blocks and complex frontend frameworks. Here is how we manage the SVT infrastructure.
SVT Play restricts access to video manifests and certain metadata based on geographic location. Our crawlers route requests through Swedish residential and mobile IP pools to ensure consistent access to restricted content catalogues.
The svt.se frontend relies heavily on React and Next.js. Instead of scraping the raw DOM, our pipeline intercepts the Next.js hydration state (JSON data embedded in the page source), resulting in faster execution and perfectly structured data.
Extracting playable video URLs requires parsing complex HLS (.m3u8) or DASH (.mpd) manifests. We extract these endpoint URLs and associated subtitle tracks, delivering them as clean string fields in your database.
Category pages and regional news feeds use infinite scroll mechanics. We map the underlying API calls and paginate through the raw JSON endpoints, bypassing the browser entirely for high-volume historical archive extraction.
Broadcasters frequently update their internal APIs. Every run emits structured logs to our observability stack. We alert on schema drift, null-rate spikes, and missing fields to repair selectors before your downstream processes fail.
Agencies track mentions of brands, politicians, and organisations across national and regional Swedish news broadcasts.
AI researchers extract high-quality Swedish text from articles and subtitles to train regional language models.
Commercial broadcasters analyse SVT schedule patterns, content genres, and publication frequencies to optimise their own programming.
Political scientists analyse topic frequency, regional bias, and editorial focus in public service broadcasting over time.
Organisations build subtitle corpora from SVT Play to train speech-to-text models for the Swedish language.
Production companies track the longevity, rerun frequency, and placement of specific documentary and entertainment properties.
"Sveriges Television produces the most comprehensive public record of Swedish news and culture, but extracting it requires navigating geo-blocks and complex video APIs."
Scraping SVT requires more than simple HTTP GET requests. SVT Play content is geo-restricted to Swedish IP addresses, and the frontend relies heavily on dynamic Next.js hydration. DataFlirt handles the Swedish proxy routing, state extraction, and video manifest parsing so you get clean, structured data without managing the infrastructure.
Everything supported by our svt.se scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and API pagination. Playwright handles JavaScript rendering and Next.js state interception for complex video pages.
We maintain dedicated pools of Swedish residential and mobile IPs to ensure consistent access to geo-restricted SVT Play content.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About svt.se scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available factual data, such as news headlines, broadcast schedules, and metadata from svt.se, is generally permissible. DataFlirt extracts only public, non-authenticated data. We do not distribute copyrighted video files or bypass DRM. Clients should consult legal counsel regarding their specific use of the extracted text and metadata.
We use Swedish residential and mobile ISP proxies to ensure our requests originate from within Sweden, granting access to the full catalogue of SVT Play metadata and manifest URLs.
No. We extract the metadata, descriptions, broadcast dates, and the HLS/DASH manifest URLs (.m3u8 files). We do not download or deliver the raw video payload bytes.
For real-time media monitoring, we can configure pipelines to poll the SVT homepage and RSS feeds at sub-5-minute intervals. Full historical archive exports run on scheduled batch processes.
Yes. We locate the VTT or SRT subtitle files associated with SVT Play videos and can deliver the raw subtitle text alongside the video metadata.
Yes. The SVT pipeline can be configured to target specific regional news sections (e.g., SVT Nyheter Skåne, SVT Nyheter Västerbotten) and extract local municipality tags.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive of Swedish news or a real-time feed of broadcast schedules, we scope, build, and operate the pipeline. Tell us what you need.