SYSTEM all green source nationalpost.com queue 11,842 URLs p99 latency 215ms dataflirt.com · scraper/nationalpost-com
RUN · 14 active pipelines · nationalpost.com live

National Post data,
at warehouse scale.

We extract news articles, Financial Post columns, author metadata, and comment threads from nationalpost.com. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Articles extracted
18.4K /day
Author updates
1.2K /24h
Comment records
54.3K /run
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from nationalpost.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Articles objects from nationalpost.com. All fields typed and schema-versioned.

urlheadlinesubheadlineauthorpublish_dateupdate_datecategorybody_textword_countpaywalled
articles
● 200 OK
"url": "https://nationalpost.com/news/politics/example-article",
"headline": "Parliament debates new fiscal policy measures",
"author": "John Ivison",
"publish_date": "2026-05-12T14:30:00Z",
"category": "Politics",
"paywalled": false
# urlheadlinesubheadlineauthorpublish_dateupdate_date
1
2
3

Complete list of extractable fields for Financial Post objects from nationalpost.com. All fields typed and schema-versioned.

urlticker_mentionsmarket_categoryheadlineauthorpublish_datebody_textrelated_companiessector
financial_post
● 200 OK
"ticker_mentions": "['TSX:RY', 'TSX:TD']",
"market_category": "Banking",
"headline": "Canadian banks report quarterly earnings beat",
"author": "Barbara Shecter",
"publish_date": "2026-05-12T09:15:00Z",
"related_companies": "['Royal Bank of Canada', 'TD Bank']"
# urlticker_mentionsmarket_categoryheadlineauthorpublish_date
1
2
3

Complete list of extractable fields for Authors objects from nationalpost.com. All fields typed and schema-versioned.

author_idnamerolebiotwitter_handlearticle_countlatest_article_urlprofile_image
authors
● 200 OK
"author_id": "auth_49281",
"name": "Kelly McParland",
"role": "Columnist",
"twitter_handle": "@KellyMcParland",
"article_count": 842,
"latest_article_url": "https://nationalpost.com/opinion/example"
# author_idnamerolebiotwitter_handlearticle_count
1
2
3

Complete list of extractable fields for Comments objects from nationalpost.com. All fields typed and schema-versioned.

comment_idarticle_urluser_nameuser_idcomment_texttimestampupvotesdownvotesreplies_count
comments
● 200 OK
"comment_id": "cmt_993821",
"user_name": "CanuckReader99",
"comment_text": "This policy will have significant impacts on the housing market.",
"timestamp": "2026-05-12T15:45:22Z",
"upvotes": 42,
"downvotes": 3
# comment_idarticle_urluser_nameuser_idcomment_texttimestamp
1
2
3

Complete list of extractable fields for Multimedia & Tags objects from nationalpost.com. All fields typed and schema-versioned.

article_urlimage_urlsvideo_idsprimary_tagsecondary_tagsseo_titlemeta_descriptionreading_time_mins
multimedia_& tags
● 200 OK
"article_url": "https://nationalpost.com/news/canada/example",
"primary_tag": "Canadian Politics",
"secondary_tags": "['Housing', 'Interest Rates', 'Bank of Canada']",
"seo_title": "Housing market reacts to Bank of Canada rate decision",
"meta_description": "A detailed look at how the latest interest rate announcement affects mortgages.",
"reading_time_mins": 4
# article_urlimage_urlsvideo_idsprimary_tagsecondary_tagsseo_title
1
2
3

Capabilities

Extract clean text from an ad-heavy DOM

News sites deploy complex layouts, programmatic advertising, and infinite scrolls. Our pipelines normalise this into clean, NLP-ready datasets.

Full Article Text

Extract body paragraphs while stripping programmatic ads, newsletter signups, and related-article injection modules.

Financial Post Sections

Target specific market, investing, and economy verticals with ticker mention extraction and sector classification.

Author Metadata

Capture contributor bios, roles, social links, and historical article counts across the Postmedia network.

Comment Extraction

Render third-party comment frames (like Viafoura) via Playwright to capture user discourse, upvotes, and reply threads.

Paywall Flagging

Detect and flag premium gated content versus free-to-read articles to maintain dataset integrity.

Metadata & SEO

Scrape primary tags, secondary topics, SEO titles, and meta descriptions used for internal taxonomy.

Multimedia Links

Extract hero image URLs, inline image captions, and embedded video IDs.

Infinite Scroll Handling

Paginate through category feeds and author pages that rely on dynamic JavaScript loading.

Scheduled Updates

Run hourly or daily diffs to capture newly published articles and updated timestamps on developing stories.

Historical Archives

Deep scrape past articles by year and month to build extensive NLP training corpora.

// engagement pipeline

From section URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide section URLs, author pages, or date ranges. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, session management, and CAPTCHA handling for nationalpost.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, ad-stripping verification, and sample datasets before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our Postmedia pipeline handles the hard parts

Modern news sites prioritise ad delivery and dynamic engagement over static HTML. Here is how we extract clean data.

pipeline-monitor · nationalpost.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
DOM cleaning
Removing programmatic ad injection

National Post articles interleave body text with dynamic ad containers, newsletter prompts, and read-more widgets. Our parsers target specific article-body classes and filter out injected nodes, ensuring your text data is contiguous and NLP-ready.

JavaScript execution
Rendering third-party comment frames

Comments are not present in the initial HTML payload. We use Playwright to execute page scripts, wait for the comment provider frame to load, and simulate scroll events to capture full discussion threads.

Pagination
Navigating infinite scroll feeds

Category pages and author profiles use infinite scroll rather than standard pagination. Our crawlers intercept the underlying GraphQL or REST API calls triggered by scroll events to paginate efficiently without rendering the full DOM.

Change detection
Tracking developing stories

News articles are frequently updated after initial publication. We hash the article content and track 'updated_at' timestamps, emitting a new record only when substantive changes occur to the headline or body text.

Paywall handling
Accurate premium content flagging

We detect paywall overlays and metadata flags to accurately categorise articles as premium. This prevents your dataset from being polluted with truncated article summaries.

Applications

Who uses National Post data — and how

Teams across industries use nationalpost.com data to build competitive products and smarter operations.

01
Media Monitoring

PR agencies and corporate communications teams track brand mentions, sentiment, and narrative development across national news.

02
Financial Sentiment Analysis

Quant funds parse Financial Post columns to gauge market sentiment on Canadian equities and macroeconomic policy.

03
NLP Training Data

AI researchers ingest decades of Canadian political discourse and journalistic text to train regional language models.

04
Political Discourse Tracking

Think tanks and academic researchers analyse opinion columns and comment threads to map ideological shifts.

05
Competitor Intelligence

Other media organisations track publication velocity, author output, and topic coverage to benchmark editorial strategy.

06
Topic Trend Analysis

Marketing teams analyse tag frequency and article volume to identify emerging cultural and economic trends.

Why DataFlirt

"The National Post archive represents decades of Canadian political discourse and financial reporting — but extracting clean text from an ad-heavy DOM requires precision engineering."

News publishers rely on complex, ad-injected DOM structures and third-party scripts that break standard HTTP parsers. DataFlirt executes full browser sessions to render comment frames, bypass infinite scrolls, and extract pure article text. We handle the infrastructure so your data science teams receive clean NLP-ready text.

Technical Spec

National Post scraper — technical capabilities

Everything supported by our nationalpost.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions — required for comment threads and dynamic content loading
Supported
Residential proxy rotation
ISP-grade residential IPs from CA/US pools to prevent rate limiting
Supported
Comment thread extraction
Captures user names, text, timestamps, and vote counts from third-party frames
Supported
Ad removal from body text
Strips programmatic ads, newsletter embeds, and related links from article bodies
Supported
Financial Post ticker linking
Extracts explicit stock ticker mentions embedded in financial reporting
Supported
Infinite scroll handling
Paginates through category feeds via API interception or scroll simulation
Supported
Change detection (diffs)
Emits updated records when an article's text or headline is modified post-publication
Supported
Premium paywalled full text
Requires active Postmedia subscription credentials; we do not bypass hard paywalls
Partial
User account details from commenters
PII and authenticated user data is strictly excluded from extraction
Partial
Infrastructure

Infrastructure powering the Postmedia pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, infinite scrolls, and interaction flows for complex news layouts.

DOM Cleaning & NLP Prep

Custom middleware strips non-content nodes — ads, tracking pixels, and inline widgets — ensuring the output text is contiguous and ready for natural language processing.

Cloud-Native Orchestration

Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — schema versioned per run
CSV
Flat file with typed columns — Excel/Sheets compatible
XLS
Legacy spreadsheet format for non-technical stakeholders
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoints to query historical article datasets
BigQuery
Streamed directly into your dataset with schema auto-detect
Snowflake
Stage + COPY INTO workflow — incremental or full-replace
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About nationalpost.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping National Post legal?

Scraping publicly available news articles and metadata is generally permissible. DataFlirt targets only public, non-authenticated text and metadata. We do not circumvent hard paywalls or extract personally identifiable information from user accounts. Clients should review Postmedia's ToS and consult legal counsel for specific use cases.

How do you handle paywalled articles?

We extract the publicly visible metadata (headline, author, tags, publication date) and flag the article as 'paywalled'. We do not bypass hard paywalls to extract premium body text without valid credentials.

Are advertisements included in the extracted text?

No. Our parsers use specific CSS and XPath selectors to target article body paragraphs while explicitly excluding programmatic ad containers, newsletter signups, and related-article injection modules.

Can you extract the comment sections?

Yes. We use headless browsers to execute the necessary JavaScript to load third-party comment frames, capturing the text, timestamps, and vote counts of the discussion.

Do you extract data from the Financial Post section?

Yes. The Financial Post is integrated into the nationalpost.com domain. We extract specific market categories, author data, and explicit stock ticker mentions embedded in the text.

Can I get historical data?

Yes. We can configure crawlers to traverse category archives and sitemaps to extract years of historical articles, subject to availability on the site.

How frequently can the data be updated?

Pipelines can run at hourly or daily cadences. We use change detection to emit updated records when a developing story's text or timestamp changes.

Is the text ready for NLP models?

Yes. By stripping ads, HTML tags, and inline widgets, we deliver contiguous strings of article text formatted specifically for ingestion into language models and sentiment classifiers.

$ dataflirt scope --new-project --source=nationalpost.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a historical archive dump or a continuous feed of Financial Post columns — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in news and media

Services

Data Extraction for Every Industry

View All Services →