SYSTEM all green source news.com.au queue 12,408 URLs p99 latency 218ms dataflirt.com · scraper/news-com.au
RUN 42 active pipelines news.com.au live

Australian media data,
at warehouse scale.

We extract breaking news, editorial content, author profiles, and public comment threads from news.com.au. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Articles extracted
14.2K /day
Comments processed
89.4K /24h
Author profiles
1.2K /run
Active pipelines
42
Uptime
99.98%
Data Dictionary

Every field we extract from news.com.au

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Article Data objects from news.com.au. All fields typed and schema-versioned.

article_idurlheadlinesubheadlineauthorpublish_dateupdate_datecontent_bodycategorytags
article_data
● 200 OK
"article_id": "nca-123456789",
"headline": "RBA holds interest rates steady at 4.35pc",
"author": "Finance Reporter",
"publish_date": "2026-05-12T14:30:00Z",
"category": "Finance",
"tags": "['RBA', 'Interest Rates', 'Economy']"
# article_idurlheadlinesubheadlineauthorpublish_date
1
2
3

Complete list of extractable fields for Author Profiles objects from news.com.au. All fields typed and schema-versioned.

author_idnamerolebiotwitter_handlearticle_countrecent_articlesprofile_url
author_profiles
● 200 OK
"author_id": "auth-9876",
"name": "Jane Doe",
"role": "Senior Political Reporter",
"twitter_handle": "@janedoe_news",
"article_count": 342,
"profile_url": "https://www.news.com.au/author/jane-doe"
# author_idnamerolebiotwitter_handlearticle_count
1
2
3

Complete list of extractable fields for Comment Threads objects from news.com.au. All fields typed and schema-versioned.

comment_idarticle_iduser_namecomment_texttimestampupvotesdownvotesreplies_count
comment_threads
● 200 OK
"comment_id": "cmt-55432",
"article_id": "nca-123456789",
"user_name": "AusReader99",
"comment_text": "Expected outcome given the inflation data.",
"upvotes": 45,
"timestamp": "2026-05-12T15:01:22Z"
# comment_idarticle_iduser_namecomment_texttimestampupvotes
1
2
3

Complete list of extractable fields for Category Feeds objects from news.com.au. All fields typed and schema-versioned.

category_namesub_categorytop_storiestrending_topicsfeed_urlarticle_countlast_updatedsection_id
category_feeds
● 200 OK
"category_name": "Sport",
"sub_category": "AFL",
"feed_url": "https://www.news.com.au/sport/afl",
"article_count": 50,
"last_updated": "2026-05-12T16:00:00Z",
"trending_topics": "['Collingwood', 'Trade Draft']"
# category_namesub_categorytop_storiestrending_topicsfeed_urlarticle_count
1
2
3

Complete list of extractable fields for Multimedia Metadata objects from news.com.au. All fields typed and schema-versioned.

media_idarticle_idmedia_typeurlcaptioncreditdurationformat
multimedia_metadata
● 200 OK
"media_id": "vid-88776",
"media_type": "video",
"url": "https://video.news.com.au/v/123.mp4",
"caption": "Press conference following RBA decision",
"duration": 124,
"format": "mp4"
# media_idarticle_idmedia_typeurlcaptioncredit
1
2
3

Capabilities

Extract the Australian news cycle

Our pipeline handles the News Corp network architecture: dynamic paywall detection, JavaScript rendered comment widgets, and infinite scroll category feeds.

Full Text Extraction

Clean text extraction from article bodies, stripping out inline advertisements, related story widgets, and newsletter signup forms.

Author Attribution

Map articles to specific journalists. Extract author bios, social handles, and historical publication counts.

Comment Mining

Extract user generated content from third party commenting widgets. Capture upvotes, downvotes, and nested reply structures.

Timestamp Normalisation

Convert relative publish times into absolute UTC timestamps. Track both initial publication and subsequent update times.

Tag and Category Mapping

Extract primary categories, subcategories, and keyword tags assigned to each article for precise content classification.

Multimedia Metadata

Capture image URLs, video embed links, captions, and photographer credits embedded within the article body.

Breaking News Detection

High frequency polling on homepage and category feeds to capture breaking news alerts within minutes of publication.

Paywall Identification

Automatically detect and flag articles locked behind the News+ premium subscription tier.

Change Detection

Monitor live blogs and developing stories. Emit diffs when headlines change or new paragraphs are added to existing URLs.

// engagement pipeline

From section URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target categories, author profiles, or keyword sets. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, JavaScript rendering for comment widgets, and proxy rotation for news.com.au.

Validation & QA
d 4–6

Schema validation, null rate checks, and text normalisation before full launch.

Delivery
ongoing

JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our media pipeline handles the hard parts

News publishers utilise aggressive caching and dynamic rendering. Here is how we maintain data integrity.

pipeline-monitor · news.com.au · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Dynamic rendering
JavaScript execution for third party widgets

Comments and embedded social media posts on news.com.au rely heavily on client side rendering. We use Playwright to execute JavaScript and wait for network idle states, ensuring we capture content that basic HTTP clients miss.

Text cleaning
Stripping editorial cruft

Article bodies are littered with inline ads, related link blocks, and newsletter embeds. Our parsers isolate the core editorial text, delivering clean paragraphs without the layout noise.

Infinite scroll
Pagination circumvention

Category feeds and search results use infinite scroll mechanisms. We simulate browser scroll events and intercept background API calls to extract the full historical feed.

Rate limiting
Geo targeted proxy rotation

News Corp employs edge protection to block aggressive scraping. We route requests through Australian residential proxies with randomised delays to maintain high throughput without triggering blocks.

Schema stability
Resilient DOM selectors

Media layouts change frequently for special events or breaking news. We use fallback selector chains targeting semantic HTML and JSON LD blocks to ensure continuous data flow during layout shifts.

Applications

Who uses Australian media data

Teams across industries use news.com.au data to build competitive products and smarter operations.

01
Media Monitoring

PR firms and corporate communications teams track brand mentions, sentiment, and share of voice across the News Corp network.

02
AI Training Data

Machine learning teams ingest clean, locally contextualised Australian English text corpora to train regional LLMs and NLP models.

03
Sentiment Analysis

Financial analysts process article tone and comment thread reactions to gauge public sentiment on economic policies and market events.

04
Competitor Intelligence

Rival publishers monitor publication velocity, author output, and trending topics to optimise their own editorial strategies.

05
Trend Forecasting

Researchers aggregate keyword frequencies and tag usage over time to identify emerging social and political trends.

06
Author Network Analysis

Investigative teams map relationships between journalists, topics, and cited sources to understand editorial bias and influence.

Why DataFlirt

"News.com.au drives the Australian daily narrative, but transforming unstructured editorial content into queryable time series data requires dedicated infrastructure."

Media scraping involves bypassing aggressive caching layers, handling dynamic paywalls, and normalising inconsistent editorial formats. DataFlirt manages the proxy rotation, JavaScript execution, and schema validation so your data science team receives clean text corpora instead of broken HTML.

Technical Spec

News.com.au scraper technical capabilities

Everything supported by our news.com.au scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Full text extraction
Clean paragraph extraction stripping inline ads and related links
Supported
Comment thread mining
Extraction of user comments, upvotes, and nested replies via JS rendering
Supported
Author metadata
Capture of journalist names, bios, and social media handles
Supported
Timestamp normalisation
Conversion of relative times to absolute UTC timestamps
Supported
Change detection
Hash based diffing for live blogs and updated articles
Supported
Multimedia links
Extraction of embedded image and video URLs with captions
Supported
News+ Premium content
Full text of articles locked behind the News+ subscription paywall
Partial
User account settings
Extraction of private user profiles and comment history requiring login
Partial
Infrastructure

Infrastructure powering the media pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering for dynamic comment widgets and infinite scroll feeds.

Residential Proxy Infrastructure

We route requests through Australian residential IPs to prevent edge blocking and ensure accurate regional content delivery.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling for high frequency breaking news polling. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline delimited or nested schema versioned per run
CSV
Flat file with typed columns Excel compatible
Parquet
Columnar format for BigQuery, Snowflake, Athena
S3
Direct bucket delivery compatible with any data lake
Webhook
HTTP POST per record for real time downstream processing
API
REST endpoints for querying historical article data
BigQuery
Streamed directly into your dataset with schema auto detect
Snowflake
Stage and COPY INTO workflow incremental or full replace
// faq

Common questions.

About news.com.au scraping, legality, and pipeline operations.

Ask us directly →
Do you extract articles behind the News+ paywall?

No. DataFlirt only extracts publicly available information. We detect and flag articles locked behind the News+ premium tier, but we do not bypass authentication walls or use subscriber credentials to access gated content.

How quickly can you detect breaking news?

For monitored category feeds and the homepage, we can configure polling intervals as low as 5 minutes. New articles are extracted and pushed via Webhook or S3 immediately upon detection.

Can you extract the comments section?

Yes. We use Playwright to execute the JavaScript required to load third party commenting widgets. We extract the commenter name, text, timestamp, and vote counts.

How do you handle live blogs that update frequently?

Our change detection system maintains a hash of the article body. When polling a live blog URL, we compare the current state against the previous run and only emit a new record if the content has changed.

Is the extracted text clean?

Yes. Our parsing logic specifically targets the editorial content blocks. We strip out inline advertisements, related story links, newsletter signups, and navigation elements, delivering clean paragraphs of text.

Can I get historical data?

We can execute backfills by crawling category archives and sitemaps. The depth of historical data depends on the publisher's site structure and archive availability.

$ dataflirt scope --new-project --source=news.com.au ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one off historical text corpus or a continuous feed of breaking news and comments we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in news and media

Services

Data Extraction for Every Industry

View All Services →