SYSTEM all green source panjiva.com queue 12,943 profiles p99 latency 812ms dataflirt.com · scraper/panjiva-com
RUN · 64 active pipelines · panjiva.com live

Global trade data,
at warehouse scale.

We extract importer directories, HS code classifications, supply chain networks, and shipment metadata from Panjiva. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Company profiles
1.2M /run
Shipment records
4.8M /24h
HS code mappings
412K /run
Active pipelines
64
Uptime
99.92%
Data Dictionary

Every field we extract from panjiva.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Company Profiles objects from panjiva.com. All fields typed and schema-versioned.

company_idcompany_namecompany_typecountryaddresstotal_shipmentstop_hs_codesprimary_trading_partnerslast_shipment_datepanjiva_rating
company_profiles
● 200 OK
"company_id": "PJ-9821445",
"company_name": "Foxconn Precision Electronics",
"company_type": "Exporter",
"country": "Taiwan",
"total_shipments": 48291,
"last_shipment_date": "2023-11-14",
"panjiva_rating": 84
# company_idcompany_namecompany_typecountryaddresstotal_shipments
1
2
3

Complete list of extractable fields for Shipment Metadata objects from panjiva.com. All fields typed and schema-versioned.

shipment_iddateshipper_nameconsignee_nameorigin_portdestination_portcarrierhs_code_summaryweight_kgteusproduct_description
shipment_metadata
● 200 OK
"shipment_id": "BOL-US-8839201",
"date": "2023-11-12",
"origin_port": "Shanghai, China",
"destination_port": "Los Angeles, CA",
"weight_kg": 14500.5,
"teus": 2,
"product_description": "Lithium-ion battery packs and accessories"
# shipment_iddateshipper_nameconsignee_nameorigin_portdestination_port
1
2
3

Complete list of extractable fields for HS Code Analytics objects from panjiva.com. All fields typed and schema-versioned.

hs_codedescriptioncategorytotal_shipmentstop_origin_countrytop_destination_countryaverage_weighttrend_indicatortrade_volume_usd
hs_code analytics
● 200 OK
"hs_code": "8507.60",
"description": "Lithium-ion accumulators",
"category": "Electrical Machinery",
"total_shipments": 104592,
"top_origin_country": "China",
"trend_indicator": "+14.2%"
# hs_codedescriptioncategorytotal_shipmentstop_origin_countrytop_destination_country
1
2
3

Complete list of extractable fields for Port & Carrier Data objects from panjiva.com. All fields typed and schema-versioned.

port_namecountrylocodetotal_vesselstotal_teus_handledtop_carriersport_authorityterminal_operatorsinbound_volumeoutbound_volume
port_& carrier data
● 200 OK
"port_name": "Port of Long Beach",
"country": "United States",
"locode": "USLGB",
"total_vessels": 4192,
"total_teus_handled": 8200000,
"inbound_volume": "6.1M TEUs"
# port_namecountrylocodetotal_vesselstotal_teus_handledtop_carriers
1
2
3

Complete list of extractable fields for Supply Chain Relationships objects from panjiva.com. All fields typed and schema-versioned.

relationship_idbuyer_namesupplier_nametransaction_countfirst_seenlast_seenprimary_commodityrisk_scorevolume_share
supply_chain relationships
● 200 OK
"relationship_id": "REL-4492-8110",
"buyer_name": "Tesla Inc",
"supplier_name": "Panasonic Energy Co",
"transaction_count": 842,
"primary_commodity": "Battery Cells",
"risk_score": "Low"
# relationship_idbuyer_namesupplier_nametransaction_countfirst_seenlast_seen
1
2
3

Capabilities

Extract the world's supply chain graph

Our Panjiva scraper targets public trade profiles, shipment metadata, and HS code directories, resolving complex supply chain relationships into structured, queryable datasets.

Importer & Exporter Discovery

Extract company profiles, operational locations, primary trading regions, and aggregate shipment volumes across global trade routes.

Shipment Metadata Extraction

Capture Bill of Lading (BoL) summaries, port of lading, port of unlading, carrier names, and product descriptions from public shipment records.

HS Code & Commodity Tracking

Map product descriptions to harmonised system (HS) codes. Track trade volumes and regional shifts by specific commodity categories.

Supply Chain Graphing

Identify buyer-supplier relationships, calculate transaction frequencies, and map multi-tier supply chain dependencies.

Port & Logistics Intelligence

Monitor port throughput, carrier market share, and shipping lane density based on aggregated bill of lading data.

Competitor Sourcing Analysis

Identify the suppliers of competitor brands, assess their supply chain concentration, and monitor their inbound freight volumes.

Macro Trade Volume Trends

Aggregate shipment data to model macroeconomic trends, regional export strength, and global supply chain bottlenecks.

Anti-Bot Circumvention

Bypass S&P Global's aggressive rate limiting and CAPTCHA walls using residential proxy pools and TLS fingerprint spoofing.

Scheduled Delta Updates

Run continuous pipelines that detect new shipments or supplier relationships, delivering only the incremental changes to your warehouse.

Multi-Region Coverage

Extract cross-border trade data spanning US customs, Latin American imports, and Asian export records available on the platform.

// engagement pipeline

From target entities to warehouse records

Brief in. Clean data out.

Define Scope
d 0

Provide target company names, HS codes, or competitor lists. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy / Playwright crawlers, proxy rotation, session management, and CAPTCHA handling for panjiva.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, and relationship mapping verification before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

Navigating S&P Global's data defences

Panjiva employs strict rate limits and bot mitigation to protect its proprietary trade data. Here is how we maintain extraction stability.

pipeline-monitor · panjiva.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Rate limiting
Distributed request pacing

Panjiva aggressively throttles IPs that exceed standard human browsing velocities. We distribute requests across a large pool of residential IPs, enforcing strict concurrency limits and randomised delays to stay beneath detection thresholds.

Session management
Sticky IP and cookie handling

Certain data views require consistent session state. We maintain sticky IP sessions tied to specific cookie jars, ensuring pagination sequences complete without triggering session-hijacking alarms.

Dynamic rendering
Playwright for heavily interactive DOMs

Supply chain graphs and complex trade charts rely on client-side JavaScript. We execute full Playwright browser sessions to hydrate the DOM, extract the underlying JSON payloads, and capture structured graph data.

Pagination limits
Search space partitioning

Panjiva restricts deep pagination on broad searches. We programmatically partition search queries by date ranges, HS sub-categories, and port codes to extract the full dataset without hitting arbitrary page limits.

Data normalisation
Entity resolution across disparate records

Company names in bill of lading data are notoriously messy (e.g., 'Foxconn' vs 'Hon Hai Precision'). We apply string distance algorithms and normalisation rules post-extraction to ensure consistent entity IDs in your warehouse.

Applications

Who uses Panjiva data — and how

Teams across industries use panjiva.com data to build competitive products and smarter operations.

01
Supplier Discovery & Diversification

Procurement teams identify alternative suppliers for critical components by analysing the export volumes and buyer networks of global manufacturers.

02
Competitor Supply Chain Mapping

Market intelligence teams reverse-engineer competitor supply chains, tracking their inbound freight volumes and identifying their primary manufacturing partners.

03
Private Equity Due Diligence

Investors validate target company growth claims by analysing their historical import/export volumes and assessing supply chain concentration risk.

04
Trade Finance & Risk Modeling

Financial institutions assess the health of trade finance portfolios by monitoring the shipment velocity and trading partner stability of corporate clients.

05
Logistics & Freight Planning

Carriers and NVOCCs analyse trade lane volumes, port congestion indicators, and competitor market share to optimise fleet deployment.

06
Macroeconomic Forecasting

Quant funds aggregate shipment metadata across specific commodity codes to predict macroeconomic shifts, inflation pressures, and industrial output.

Why DataFlirt

"Global trade runs on paper, but supply chain alpha requires structured data. Panjiva holds the world's shipping graph, but extracting it at scale demands industrial-grade infrastructure."

Extracting trade data from S&P Global platforms requires navigating strict rate limits, complex session flows, and aggressive bot mitigation. DataFlirt handles the proxy rotation, session management, and schema normalisation so your procurement and quant teams can focus on supplier risk — not pipeline maintenance.

Technical Spec

Panjiva scraper — technical capabilities

Everything supported by our panjiva.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Public company profiles
Extraction of importer/exporter metadata and aggregate trade volumes
Supported
HS code directories
Scraping of commodity classifications and regional trade summaries
Supported
Supplier-buyer graph previews
Mapping of primary trading partners based on public shipment records
Supported
JavaScript rendering
Execution of client-side code to capture dynamic charts and graphs
Supported
Change detection (diffs)
Hash-based diffing to emit only new shipments or updated profiles
Supported
Automated CAPTCHA bypass
Integration with CapSolver for handling S&P Global challenge pages
Supported
Full Bill of Lading (BoL) documents
Unredacted, line-item customs documents (requires Panjiva Enterprise login)
Partial
Private shipment financial values
Exact USD transaction values for private shipments (gated by S&P Global)
Partial
Real-time vessel tracking
Live AIS coordinates (Panjiva focuses on customs records, not live telemetry)
Partial
Infrastructure

Infrastructure powering the Panjiva pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across US/EU/APAC regions. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.

Cloud-Native Orchestration

Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested — schema versioned per run
CSV
Flat file with typed columns — Excel/Sheets compatible
XLS
Formatted spreadsheet for procurement analysts
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery — compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoint to query your extracted Panjiva dataset
PostgreSQL
Upsert into your existing schema with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About panjiva.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Panjiva legal?

Scraping publicly available information is generally permissible under applicable law, reinforced by the hiQ v. LinkedIn ruling. DataFlirt targets only public, non-authenticated trade data and company profiles on Panjiva. We do not circumvent authentication walls or extract proprietary S&P Global data reserved for paying enterprise clients. Clients should review terms of service and consult legal counsel for specific use cases.

How do you handle Panjiva's anti-bot systems?

We use residential ISP proxies, full Playwright browser sessions with realistic TLS fingerprints, and strict request pacing. We partition search queries to avoid triggering pagination traps and monitor for rate-limit responses in real time, triggering automatic pool rotation.

What data is gated vs public on Panjiva?

Public extraction yields company profiles, aggregate trade volumes, HS code summaries, and high-level supplier-buyer relationships. Full, unredacted Bill of Lading (BoL) documents and specific transaction values are gated behind Panjiva Enterprise authentication and cannot be scraped publicly.

How fresh is the data?

Customs data reporting inherently involves a lag of several days to weeks depending on the jurisdiction. We can configure pipelines to run weekly or monthly to capture the latest published shipments as soon as they appear in the public index.

Can you extract historical supply chain data?

Yes. We can traverse historical shipment records and company profiles available on the public platform, allowing you to model supply chain shifts over multiple years.

Can I request a sample dataset before committing?

Absolutely. We provide a sample run targeting specific HS codes or company names as part of the pre-engagement scoping process — so you can validate schema fit, field completeness, and data quality before signing any contract.

$ dataflirt scope --new-project --source=panjiva.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off extraction of specific HS codes or continuous monitoring of competitor supply chains — we scope, build, and operate the pipeline. Tell us what you need.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in industrial and mro

Services

Data Extraction for Every Industry

View All Services →