We extract importer directories, HS code classifications, supply chain networks, and shipment metadata from Panjiva. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Company Profiles objects from panjiva.com. All fields typed and schema-versioned.
"company_id": "PJ-9821445", "company_name": "Foxconn Precision Electronics", "company_type": "Exporter", "country": "Taiwan", "total_shipments": 48291, "last_shipment_date": "2023-11-14", "panjiva_rating": 84
| # | company_id | company_name | company_type | country | address | total_shipments |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Shipment Metadata objects from panjiva.com. All fields typed and schema-versioned.
"shipment_id": "BOL-US-8839201", "date": "2023-11-12", "origin_port": "Shanghai, China", "destination_port": "Los Angeles, CA", "weight_kg": 14500.5, "teus": 2, "product_description": "Lithium-ion battery packs and accessories"
| # | shipment_id | date | shipper_name | consignee_name | origin_port | destination_port |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for HS Code Analytics objects from panjiva.com. All fields typed and schema-versioned.
"hs_code": "8507.60", "description": "Lithium-ion accumulators", "category": "Electrical Machinery", "total_shipments": 104592, "top_origin_country": "China", "trend_indicator": "+14.2%"
| # | hs_code | description | category | total_shipments | top_origin_country | top_destination_country |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Port & Carrier Data objects from panjiva.com. All fields typed and schema-versioned.
"port_name": "Port of Long Beach", "country": "United States", "locode": "USLGB", "total_vessels": 4192, "total_teus_handled": 8200000, "inbound_volume": "6.1M TEUs"
| # | port_name | country | locode | total_vessels | total_teus_handled | top_carriers |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Supply Chain Relationships objects from panjiva.com. All fields typed and schema-versioned.
"relationship_id": "REL-4492-8110", "buyer_name": "Tesla Inc", "supplier_name": "Panasonic Energy Co", "transaction_count": 842, "primary_commodity": "Battery Cells", "risk_score": "Low"
| # | relationship_id | buyer_name | supplier_name | transaction_count | first_seen | last_seen |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our Panjiva scraper targets public trade profiles, shipment metadata, and HS code directories, resolving complex supply chain relationships into structured, queryable datasets.
Extract company profiles, operational locations, primary trading regions, and aggregate shipment volumes across global trade routes.
Capture Bill of Lading (BoL) summaries, port of lading, port of unlading, carrier names, and product descriptions from public shipment records.
Map product descriptions to harmonised system (HS) codes. Track trade volumes and regional shifts by specific commodity categories.
Identify buyer-supplier relationships, calculate transaction frequencies, and map multi-tier supply chain dependencies.
Monitor port throughput, carrier market share, and shipping lane density based on aggregated bill of lading data.
Identify the suppliers of competitor brands, assess their supply chain concentration, and monitor their inbound freight volumes.
Aggregate shipment data to model macroeconomic trends, regional export strength, and global supply chain bottlenecks.
Bypass S&P Global's aggressive rate limiting and CAPTCHA walls using residential proxy pools and TLS fingerprint spoofing.
Run continuous pipelines that detect new shipments or supplier relationships, delivering only the incremental changes to your warehouse.
Extract cross-border trade data spanning US customs, Latin American imports, and Asian export records available on the platform.
Brief in. Clean data out.
Provide target company names, HS codes, or competitor lists. We design the extraction schema together.
We configure Scrapy / Playwright crawlers, proxy rotation, session management, and CAPTCHA handling for panjiva.com.
Schema validation, null-rate checks, and relationship mapping verification before full launch.
JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Panjiva employs strict rate limits and bot mitigation to protect its proprietary trade data. Here is how we maintain extraction stability.
Panjiva aggressively throttles IPs that exceed standard human browsing velocities. We distribute requests across a large pool of residential IPs, enforcing strict concurrency limits and randomised delays to stay beneath detection thresholds.
Certain data views require consistent session state. We maintain sticky IP sessions tied to specific cookie jars, ensuring pagination sequences complete without triggering session-hijacking alarms.
Supply chain graphs and complex trade charts rely on client-side JavaScript. We execute full Playwright browser sessions to hydrate the DOM, extract the underlying JSON payloads, and capture structured graph data.
Panjiva restricts deep pagination on broad searches. We programmatically partition search queries by date ranges, HS sub-categories, and port codes to extract the full dataset without hitting arbitrary page limits.
Company names in bill of lading data are notoriously messy (e.g., 'Foxconn' vs 'Hon Hai Precision'). We apply string distance algorithms and normalisation rules post-extraction to ensure consistent entity IDs in your warehouse.
Procurement teams identify alternative suppliers for critical components by analysing the export volumes and buyer networks of global manufacturers.
Market intelligence teams reverse-engineer competitor supply chains, tracking their inbound freight volumes and identifying their primary manufacturing partners.
Investors validate target company growth claims by analysing their historical import/export volumes and assessing supply chain concentration risk.
Financial institutions assess the health of trade finance portfolios by monitoring the shipment velocity and trading partner stability of corporate clients.
Carriers and NVOCCs analyse trade lane volumes, port congestion indicators, and competitor market share to optimise fleet deployment.
Quant funds aggregate shipment metadata across specific commodity codes to predict macroeconomic shifts, inflation pressures, and industrial output.
"Global trade runs on paper, but supply chain alpha requires structured data. Panjiva holds the world's shipping graph, but extracting it at scale demands industrial-grade infrastructure."
Extracting trade data from S&P Global platforms requires navigating strict rate limits, complex session flows, and aggressive bot mitigation. DataFlirt handles the proxy rotation, session management, and schema normalisation so your procurement and quant teams can focus on supplier risk — not pipeline maintenance.
Everything supported by our panjiva.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.
We maintain pools of residential ISP proxies across US/EU/APAC regions. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.
Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About panjiva.com scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available information is generally permissible under applicable law, reinforced by the hiQ v. LinkedIn ruling. DataFlirt targets only public, non-authenticated trade data and company profiles on Panjiva. We do not circumvent authentication walls or extract proprietary S&P Global data reserved for paying enterprise clients. Clients should review terms of service and consult legal counsel for specific use cases.
We use residential ISP proxies, full Playwright browser sessions with realistic TLS fingerprints, and strict request pacing. We partition search queries to avoid triggering pagination traps and monitor for rate-limit responses in real time, triggering automatic pool rotation.
Public extraction yields company profiles, aggregate trade volumes, HS code summaries, and high-level supplier-buyer relationships. Full, unredacted Bill of Lading (BoL) documents and specific transaction values are gated behind Panjiva Enterprise authentication and cannot be scraped publicly.
Customs data reporting inherently involves a lag of several days to weeks depending on the jurisdiction. We can configure pipelines to run weekly or monthly to capture the latest published shipments as soon as they appear in the public index.
Yes. We can traverse historical shipment records and company profiles available on the public platform, allowing you to model supply chain shifts over multiple years.
Absolutely. We provide a sample run targeting specific HS codes or company names as part of the pre-engagement scoping process — so you can validate schema fit, field completeness, and data quality before signing any contract.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off extraction of specific HS codes or continuous monitoring of competitor supply chains — we scope, build, and operate the pipeline. Tell us what you need.