SYSTEM all green source icac.org queue 3,412 reports p99 latency 215ms dataflirt.com · scraper/icac-org
RUN . 14 active pipelines . icac.org live

Global cotton data,
at warehouse scale.

We extract cotton production statistics, global price indices, supply-demand forecasts, and trade metrics from the International Cotton Advisory Committee. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Data points extracted
412K /month
Price updates
14.2K /run
Reports parsed
890 /week
Active pipelines
14
Uptime
99.98%
Data Dictionary

Every field we extract from icac.org

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Supply & Demand objects from icac.org. All fields typed and schema-versioned.

countryseasonproduction_mtconsumption_mtimportsexportsending_stocksstock_to_use_ratioforecast_datecurrencyunit
supply_& demand
● 200 OK
"country": "India",
"season": "2025/26",
"production_mt": 5820000,
"consumption_mt": 5400000,
"exports": 350000,
"ending_stocks": 1850000,
"stock_to_use_ratio": 34.2
# countryseasonproduction_mtconsumption_mtimportsexports
1
2
3

Complete list of extractable fields for Price Indices objects from icac.org. All fields typed and schema-versioned.

index_namedatepricecurrencyunitpercent_changehistorical_avgsourcebasisdelivery_month
price_indices
● 200 OK
"index_name": "Cotlook A Index",
"date": "2026-04-12",
"price": 94.25,
"currency": "USD",
"unit": "cents/lb",
"percent_change": 1.2
# index_namedatepricecurrencyunitpercent_change
1
2
3

Complete list of extractable fields for Government Support objects from icac.org. All fields typed and schema-versioned.

countryyeardirect_subsidyborder_protectioncrop_insurancemin_support_pricetotal_assistanceassistance_per_kgpolicy_type
government_support
● 200 OK
"country": "USA",
"year": "2025",
"direct_subsidy": 450000000,
"crop_insurance": 820000000,
"total_assistance": 1270000000,
"assistance_per_kg": 0.32
# countryyeardirect_subsidyborder_protectioncrop_insurancemin_support_price
1
2
3

Complete list of extractable fields for Yield Data objects from icac.org. All fields typed and schema-versioned.

countryregionseasonarea_harvested_hayield_kg_halint_productionseed_productionirrigation_pctgm_cotton_pct
yield_data
● 200 OK
"country": "Brazil",
"season": "2025/26",
"area_harvested_ha": 1750000,
"yield_kg_ha": 1920,
"lint_production": 3360000,
"gm_cotton_pct": 98.5
# countryregionseasonarea_harvested_hayield_kg_halint_production
1
2
3

Complete list of extractable fields for Trade Matrices objects from icac.org. All fields typed and schema-versioned.

exporter_countryimporter_countryseasonvolume_mtvalue_usdtransport_modetariff_ratequota_limitactual_shipped
trade_matrices
● 200 OK
"exporter_country": "Australia",
"importer_country": "Vietnam",
"season": "2025/26",
"volume_mt": 420000,
"tariff_rate": 0.0,
"actual_shipped": 185000
# exporter_countryimporter_countryseasonvolume_mtvalue_usdtransport_mode
1
2
3

Capabilities

Extract global cotton metrics with precision

Our ICAC scraper handles the complexity of agricultural data extraction: parsing PDF statistical reports, reconstructing chart data, and normalising decades of historical time-series.

Production & Yield Data

Extract area harvested, yield per hectare, and total lint production by country and season, normalised into standard units.

Price Index Tracking

Capture the Cotlook A Index and other regional pricing metrics, maintaining a continuous time-series database.

PDF Data Parsing

Automated extraction of tabular data from ICAC monthly and annual PDF reports using advanced OCR and table boundary detection.

Chart Data Extraction

Reconstruct underlying datasets from embedded web charts and visualisations published on the ICAC portal.

Supply & Demand Forecasts

Monitor changes in consumption, ending stocks, and stock-to-use ratios across major textile manufacturing hubs.

Trade Matrices

Map bilateral cotton trade flows, tracking import and export volumes between producing and consuming nations.

Government Policy Data

Quantify state intervention, tracking direct subsidies, minimum support prices, and border protection measures.

Historical Archives

Backfill your database with decades of historical cotton statistics, mapped to a consistent modern schema.

Scheduled Updates

Configure pipelines to run immediately after ICAC publishes new monthly bulletins or quarterly forecasts.

// engagement pipeline

From ICAC reports to warehouse records

Brief in. Clean data out.

Define Scope
d 0

Select the datasets you need: production yields, trade matrices, price indices, or historical archives.

Pipeline Build
d 2–4

We configure web crawlers and PDF parsing pipelines to extract and normalise the target statistics.

Validation & QA
d 4–6

Unit normalisation checks, null-rate monitoring, and time-series continuity validation.

Delivery
ongoing

Structured data pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our ICAC pipeline handles unstructured data

Agricultural bodies often publish critical data in hostile formats. Here is how we convert static reports into queryable databases.

pipeline-monitor · icac.org · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Document extraction
PDF table parsing and OCR

Much of ICAC's historical data exists only in PDF format. We deploy Camelot and Tesseract OCR to detect table boundaries, extract cell values, and reconstruct the tabular structure into machine-readable JSON.

Data normalisation
Handling shifting statistical methodologies

Over decades, ICAC has changed its reporting columns, country names, and unit metrics. Our pipeline applies normalisation rules to map legacy data structures into a unified, consistent schema.

Visualisation scraping
Chart data reconstruction

For data presented purely as web charts, we intercept the underlying network requests or parse the JavaScript data objects to extract the raw coordinates and values before they are rendered.

Change detection
Only re-scrape what has changed

We maintain a hash index of last-seen values per statistical category. Subsequent runs only push diffs, reducing storage bloat and downstream processing load for your data engineering team.

Monitoring & alerting
24/7 pipeline health

Every run emits structured logs to our observability stack. We alert on null-rate spikes, unit anomalies, and missing reports, responding before you notice a gap in your data.

Applications

Who uses ICAC data

Teams across industries use icac.org data to build competitive products and smarter operations.

01
Commodity Trading

Quantitative funds feed ICAC supply and demand forecasts into pricing models to predict movements in global cotton futures.

02
Supply Chain Planning

Textile manufacturers track yield forecasts and ending stocks to optimise raw material procurement and hedge against price volatility.

03
Policy Research

Agricultural economists analyse government subsidy data and trade matrices to evaluate the impact of state intervention on global markets.

04
Agritech Models

Machine learning teams correlate historical ICAC yield data with satellite imagery and weather patterns to train predictive crop models.

05
Market Intelligence

Consultancies track the long-term shift in mill use and consumption metrics to advise clients on factory location strategy.

06
Investment Due Diligence

Private equity firms evaluate macroeconomic cotton trends before investing in regional spinning mills or apparel manufacturing hubs.

Why DataFlirt

"The ICAC holds the definitive dataset on global cotton economics, but accessing historical time-series often requires manually extracting tables from decades of PDF reports."

Most data teams underestimate the complexity of parsing unstructured statistical reports. Reliable ICAC extraction requires OCR, table-boundary detection, PDF parsing, and strict normalisation rules to handle shifting statistical methodologies over time. DataFlirt absorbs that complexity so your analysts can focus on forecasting.

Technical Spec

ICAC scraper - technical capabilities

Everything supported by our icac.org scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

PDF Table Extraction
Automated parsing of monthly and annual statistical reports
Supported
Chart data reconstruction
Extraction of raw data from Highcharts and D3 visualisations
Supported
Historical time-series
Backfilling of legacy data mapped to modern schemas
Supported
Unit normalisation
Standardising bales, metric tons, and hectares across reports
Supported
Change detection (diffs)
Hash-based diff: only emit records with changed fields
Supported
Webhook delivery
HTTP POST per record or batch for real-time workflows
Supported
Member-only publications
Requires ICAC member state credentials to access
Partial
Private working group minutes
Internal committee documents are strictly gated
Partial
Cotlook real-time feed
Requires direct commercial subscription with Cotlook
Partial
Infrastructure

Infrastructure powering the ICAC pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheusCamelotTesseract OCR
PDF Parsing Stack

We utilise Camelot and Tesseract OCR to accurately detect table structures within unstructured PDF documents, converting visual grids into structured JSON arrays.

Scrapy + Playwright Stack

Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering for interactive charts and dynamic portal navigation.

Cloud-Native Orchestration

Pipelines run on AWS ECS. Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested - schema versioned per run
CSV
Flat file with typed columns - Excel/Sheets compatible
Parquet
Columnar format for BigQuery, Snowflake, Athena
S3
Direct bucket delivery - compatible with any data lake
BigQuery
Streamed directly into your dataset with schema auto-detect
Webhook
HTTP POST per record for downstream processing
Postgres
Upsert into your existing schema with conflict resolution
Snowflake
Stage + COPY INTO workflow - incremental or full-replace
// faq

Common questions.

About icac.org scraping, legality, and pipeline operations.

Ask us directly →
Is scraping icac.org legal?

Scraping publicly available statistical data from government and intergovernmental portals is generally permissible. DataFlirt targets only public, non-authenticated agricultural reports and indices. We do not extract gated member-only content or circumvent authentication walls. Clients should consult legal counsel for specific use cases.

How do you handle data locked in PDFs?

We deploy specialised PDF parsing libraries like Camelot alongside Tesseract OCR to identify table boundaries, extract text, and reconstruct the data into structured JSON formats. We also apply strict validation rules to ensure accuracy.

How often is the data updated?

ICAC typically publishes major statistical updates on a monthly or quarterly basis. We configure pipelines to monitor the portal daily and trigger extraction immediately upon the publication of new reports.

Can you normalise historical data?

Yes. ICAC reporting formats have evolved over the years. We map legacy column headers, country names, and measurement units to a consistent modern schema, providing a clean time-series database.

Do you extract real-time cotton prices?

We extract the price indices published directly on the ICAC portal, such as the Cotlook A Index. However, live tick-by-tick market data requires direct integration with commodity exchanges.

What is the minimum viable engagement?

Our packages start at defined statistical sets (e.g., global production and mill use) with monthly delivery. For full historical archive extraction and custom normalisation, we price based on pipeline complexity.

$ dataflirt scope --new-project --source=icac.org ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off historical archive dump or a continuous feed of monthly cotton statistics - we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in textile and fabric

Services

Data Extraction for Every Industry

View All Services →