We extract articles, breaking news, political updates, economic reports, and historical archives from cna.com.tw. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for News Articles objects from cna.com.tw. All fields typed and schema-versioned.
"article_id": "202310240012", "headline": "Taiwan export orders decline narrows in September", "publish_timestamp": "2023-10-24T10:15:00Z", "update_timestamp": "2023-10-24T11:02:15Z", "author": "Pan Tzu-yu", "category": "Economics", "tags": "['exports', 'manufacturing', 'MOEA']", "url": "https://focustaiwan.tw/business/202310240012"
| # | article_id | url | headline | subheadline | content_body | publish_timestamp |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Breaking Alerts objects from cna.com.tw. All fields typed and schema-versioned.
"alert_id": "B-99281", "headline": "Central Bank announces rate decision", "timestamp": "2023-10-24T08:30:00Z", "priority_level": "high", "category": "Finance", "source_url": "https://www.cna.com.tw/news/afe/202310240001.aspx", "related_tags": "['Central Bank', 'Interest Rates']"
| # | alert_id | headline | summary | timestamp | priority_level | category |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Author Metadata objects from cna.com.tw. All fields typed and schema-versioned.
"author_id": "A-4492", "name": "Yeh Su-ping", "role": "Senior Political Reporter", "desk_location": "Taipei", "article_count": 1420, "recent_articles": "['202310240015', '202310230088']"
| # | author_id | name | role | article_count | recent_articles | desk_location |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Categories & Tags objects from cna.com.tw. All fields typed and schema-versioned.
"category_id": "C-12", "name": "Cross-Strait", "parent_category": "Politics", "article_count_24h": 45, "trending_score": 88.5, "last_updated": "2023-10-24T12:00:00Z", "feed_url": "https://www.cna.com.tw/list/acn.aspx"
| # | category_id | name | parent_category | article_count_24h | trending_score | top_article_url |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Media Extraction objects from cna.com.tw. All fields typed and schema-versioned.
"media_id": "IMG-20231024-001", "image_url": "https://imgcdn.cna.com.tw/www/WebPhotos/1024/20231024/1024x768.jpg", "caption": "TSMC facility in Hsinchu Science Park.", "photographer": "CNA Photo", "resolution": "1024x768", "format": "image/jpeg", "article_url": "https://www.cna.com.tw/news/afe/202310240012.aspx"
| # | media_id | image_url | caption | photographer | article_url | resolution |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our CNA scraper captures every layer of the publication: breaking news alerts, deep historical archives, author metadata, and categorised feeds - with proper UTF-8 encoding and anti-bot circumvention built in.
Headline, body text, author, publish date, and update timestamps extracted cleanly without boilerplate HTML.
Sub-minute polling on breaking feeds to capture high-priority alerts the moment they are published.
Pagination through years of historical content to build comprehensive datasets for backtesting and research.
Strict UTF-8 text normalisation ensures no mojibake or encoding corruption in your downstream models.
Extract internal CNA tags, categories, and related article links to map topic clusters accurately.
High-resolution image URLs, captions, and photographer credits captured and linked to the parent article.
Monitor specific reporters or regional desks to track editorial focus and coverage volume.
Targeted extraction of specific political and international relations feeds for geopolitical risk modeling.
Run one-off historical exports or configure continuous pipelines at hourly or real-time cadences.
Brief in. Clean data out.
Provide target categories, historical date ranges, or keyword sets. We design the extraction schema together.
We configure Scrapy crawlers, proxy rotation, and encoding normalisation specifically for cna.com.tw.
Schema validation, null-rate checks, timestamp parsing verification, and text encoding tests before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
News scraping requires high-frequency polling without triggering rate limits. Here is how we maintain reliable extraction.
To capture breaking alerts in real time, we distribute polling requests across a wide pool of Taiwanese residential IPs, preventing IP bans while maintaining sub-minute latency on critical feeds.
Older articles often require navigating through complex paginated structures and dynamic API endpoints. Our crawlers map these hidden endpoints to extract deep historical data efficiently.
News sites frequently contain mixed encodings or invisible control characters. We apply strict normalisation rules to ensure all Traditional Chinese text is perfectly formatted for NLP ingestion.
News articles are frequently updated after initial publication. We maintain a hash of article contents and emit diffs when headlines, bodies, or timestamps change.
Media sites update their CMS templates without warning. We monitor field null-rates in real time and automatically alert our engineers if a DOM change impacts data completeness.
Firms monitor cross-strait relations, defense updates, and diplomatic statements to model regional stability.
Quants ingest economic reports, central bank announcements, and export data to inform trading algorithms.
AI teams use clean, high-quality Traditional Chinese text corpora to train and fine-tune language models.
Agencies track brand mentions, executive coverage, and industry sentiment across Taiwan's primary news network.
Manufacturers monitor semiconductor industry news, energy policy changes, and infrastructure updates.
Researchers analyse historical publication trends, political discourse, and editorial shifts over time.
"CNA is the definitive source for Taiwanese geopolitical and economic developments - but raw HTML is useless for programmatic analysis."
Most teams underestimate the investment required: reliable news scraping requires high-frequency polling, strict encoding management, proxy rotation, and real-time diffing to catch article corrections. DataFlirt absorbs that complexity so your engineers can focus on the analysis - not the infrastructure.
Everything supported by our cna.com.tw scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles high-throughput crawl orchestration and deduplication. Playwright renders dynamic article components and infinite-scroll archive pages.
We maintain pools of Taiwanese residential IPs to ensure requests appear as local reader traffic, avoiding geo-blocks and rate limits.
Pipelines run on AWS Lambda for burst scaling during breaking news events. Airflow handles scheduling and SLA alerting.
Data delivered to where your team already works — no new tooling required.
About cna.com.tw scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available news articles is generally permissible for analysis and indexing. DataFlirt extracts only public, non-authenticated text and metadata. We do not bypass enterprise paywalls or extract proprietary internal feeds. Clients should review applicable copyright laws regarding the redistribution of scraped news content.
We distribute polling requests across a large pool of Taiwanese residential proxies. This ensures our aggregate request volume remains high while individual IP request rates stay well below CNA's blocking thresholds.
Yes. We can configure pipelines to traverse historical indexes and search endpoints to extract articles dating back several years, building a complete retrospective dataset.
For real-time pipelines, we poll specific category feeds at sub-minute intervals. Newly detected articles are processed and pushed via Webhook within seconds of publication.
Yes. Our pipelines enforce strict UTF-8 normalisation. We strip invisible control characters and ensure the text payload is perfectly formatted for ingestion by modern NLP and LLM training pipelines.
Yes. We can restrict the crawl scope to specific CNA categories, author IDs, or keyword sets, ensuring you only pay for and process the data relevant to your models.
News sites frequently correct or expand articles. Our change-detection system hashes article content on each pass. If an article changes, we emit a new record with the updated payload and the new modification timestamp.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off historical archive dump or a continuous real-time news feed - we scope, build, and operate the pipeline. Tell us what you need.