We extract news articles, author profiles, editorial categories, and tag metadata from liberation.fr. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Article Metadata objects from liberation.fr. All fields typed and schema-versioned.
"article_id": "847291A", "headline": "Réforme des retraites : les syndicats dans la rue", "author_name": "Jean Dupont", "published_at": "2023-10-12T08:30:00Z", "category": "Politique", "paywall_status": "premium", "word_count": 1240, "tags": "['Retraites', 'Syndicats', 'Manifestation']"
| # | article_id | url | headline | subheadline | author_name | published_at |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Author Profiles objects from liberation.fr. All fields typed and schema-versioned.
"author_id": "JD492", "name": "Jean Dupont", "role": "Journaliste Politique", "twitter_handle": "@jdupont_libe", "article_count": 342, "latest_article_date": "2023-10-12", "profile_url": "https://www.liberation.fr/auteur/jean-dupont/"
| # | author_id | name | bio | role | twitter_handle | article_count |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Article Content objects from liberation.fr. All fields typed and schema-versioned.
"url": "https://www.liberation.fr/politique/reforme-retraites", "headline": "Réforme des retraites : les syndicats dans la rue", "body_text": "Les rues de Paris étaient remplies ce jeudi...", "quotes": "['Nous ne céderons pas', 'Le gouvernement doit écouter']", "external_links": "['https://www.insee.fr/statistiques']", "paywall_cutoff": true, "scraped_at": "2023-10-12T09:15:22Z"
| # | url | headline | body_text | quotes | external_links | internal_links |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Homepage & Sections objects from liberation.fr. All fields typed and schema-versioned.
"section_name": "À la une", "rank_position": 1, "article_url": "https://www.liberation.fr/societe/climat", "headline": "Climat : le rapport alarmant", "is_featured": true, "layout_type": "hero_banner", "scraped_at": "2023-10-12T10:00:00Z"
| # | section_name | rank_position | article_url | headline | is_featured | scraped_at |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for CheckNews Data objects from liberation.fr. All fields typed and schema-versioned.
"question": "Est-il vrai que le gouvernement a annulé cette taxe ?", "answer_summary": "Non, la taxe a été reportée à 2025.", "verdict": "Faux", "article_url": "https://www.liberation.fr/checknews/taxe-annulee", "author": "Service CheckNews", "published_at": "2023-10-11T14:20:00Z", "tags": "['Economie', 'Gouvernement']"
| # | question | answer_summary | verdict | article_url | author | published_at |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our scraper handles every layer of the publication: front-page rankings, deep article archives, author profiles, and real-time news updates with anti-bot circumvention built in.
Headlines, publication dates, update timestamps, authorship, and category tags scraped at the URL level.
Capture author bios, social handles, and historical article lists to track editorial focus over time.
Automatically flag whether an article is open-access or requires a premium subscription.
Monitor the 'Direct' feed for real-time updates, timestamped per crawl to track breaking news evolution.
Extract specific structured fields from the CheckNews section, including questions, verdicts, and cited sources.
Track which articles are featured on the homepage and their rank position over time.
Extract internal tags and category structures to categorise content across Politique, Société, International, and more.
Capture main image URLs, embedded video links, and image captions associated with news reports.
Run continuous pipelines at hourly or daily cadences to maintain an up-to-date archive of published content.
Brief in. Clean data out.
Provide section URLs, author IDs, or keyword sets. We design the extraction schema together.
We configure Scrapy crawlers, proxy rotation, and session management for liberation.fr.
Schema validation, null-rate checks, and sample article data before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Media sites employ strict scraping limits and dynamic layouts. Here is how we stay resilient.
Media sites block data centre IPs aggressively. Our crawlers use French residential ISP proxies with realistic browser fingerprints and randomised request timing to maintain access.
News sites change article templates for special events. We use multi-layer fallback chains via CSS selectors and JSON-LD schema parsing to ensure data flows even when the DOM changes.
News stories evolve. We maintain a hash index of last-seen values per article. Subsequent runs push diffs when headlines change or new paragraphs are added.
We convert relative time formats and French locale date strings into standard ISO 8601 UTC timestamps for clean database insertion.
Every run emits structured logs. We alert on null-rate spikes and schema drift, responding immediately to keep your data feed active.
PR agencies and brands track mentions, sentiment, and editorial placement across major French publications.
Machine learning teams use high-quality French journalistic text to train language models and entity recognition systems.
Researchers analyse coverage bias, topic frequency, and author focus during election cycles.
Other media organisations monitor publication velocity, paywall strategies, and author output.
Analysts track tag frequency and category volume to identify emerging social and political trends in France.
Platforms ingest CheckNews data to build databases of verified claims and debunked misinformation.
"Journalistic data provides the most accurate reflection of societal trends, but extracting structured text from dynamic media layouts requires dedicated infrastructure."
Most teams underestimate the complexity of media scraping. Reliable extraction from liberation.fr requires handling paywall logic, dynamic editorial templates, French locale date parsing, and constant anti-bot updates. DataFlirt manages the entire pipeline so your data science team receives clean, normalised text ready for analysis.
Everything supported by our liberation.fr scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering and dynamic layout hydration.
We maintain pools of residential ISP proxies across France. Rotation happens per-request to avoid media anti-bot blocks.
Pipelines run on AWS ECS. Airflow handles scheduling, dependency management, and SLA alerting. State stored in Postgres.
Data delivered to where your team already works — no new tooling required.
About liberation.fr scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available news headlines, metadata, and open-access snippets is generally permissible. DataFlirt targets only public, non-authenticated data. We do not bypass paywalls to extract premium content without authorisation. Clients must review their specific use cases regarding copyright and text-and-data-mining (TDM) exemptions under EU law.
No. DataFlirt does not circumvent authentication walls or steal paid content. We extract the metadata, headlines, and the publicly visible preview text of premium articles, alongside a flag indicating the content is paywalled.
For live monitoring use cases, we configure pipelines to poll the 'Direct' feed and homepage at high frequencies, achieving sub-5-minute latency for new publication detection.
Our core extraction pipelines deliver the raw French text exactly as published. If translation is required, we can integrate third-party translation APIs into the processing layer before delivery.
Yes. We can run deep crawls through author pages and category archives to extract historical articles going back several years, depending on site availability.
Yes. We provide a sample run of up to 500 articles as part of the scoping process so you can validate the schema and text quality before committing.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive dump or a continuous live news feed — we scope, build, and operate the pipeline. Tell us what you need.