We extract public company profiles, contact directories, and professional metadata from Swordfish.Ai. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Company Profiles objects from swordfish.ai. All fields typed and schema-versioned.
"company_id": "sf_comp_8921", "name": "TechCorp Solutions", "industry": "Information Technology", "headcount": "501-1000", "hq_location": "San Francisco, CA", "website": "techcorpsolutions.com", "linkedin_url": "linkedin.com/company/techcorp"
| # | company_id | name | industry | website | headcount | founded_year |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Directory Listings objects from swordfish.ai. All fields typed and schema-versioned.
"directory_id": "dir_eng_09", "category": "Engineering", "sub_category": "Software Development", "total_entries": 14500, "page_number": 1, "listed_companies": 50, "update_timestamp": "2026-05-12T10:00:00Z"
| # | directory_id | category | sub_category | total_entries | page_number | listed_companies |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Professional Metadata objects from swordfish.ai. All fields typed and schema-versioned.
"profile_id": "prof_99281", "full_name": "Jane Doe", "job_title": "VP of Engineering", "company_name": "TechCorp Solutions", "location": "New York, NY", "linkedin_url": "linkedin.com/in/janedoe123", "public_email_domain": "techcorpsolutions.com"
| # | profile_id | full_name | job_title | company_name | location | public_email_domain |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Search Results objects from swordfish.ai. All fields typed and schema-versioned.
"keyword": "Chief Marketing Officer", "result_type": "professional", "rank": 1, "entity_name": "John Smith", "entity_url": "swordfish.ai/p/john-smith", "match_score": 98, "scraped_at": "2026-05-12T10:05:00Z"
| # | keyword | result_type | rank | entity_name | entity_url | snippet |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Industry Aggregates objects from swordfish.ai. All fields typed and schema-versioned.
"industry_name": "Healthcare", "total_companies": 8200, "region": "North America", "average_headcount": 250, "sub_sectors": "['Telehealth', 'Medical Devices']", "last_updated": "2026-05-12T10:10:00Z"
| # | industry_name | total_companies | top_companies | region | average_headcount | sub_sectors |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our Swordfish.Ai scraper handles directory traversal, company profile extraction, and professional metadata parsing — with JavaScript rendering and anti-bot circumvention built in.
Capture every public company and professional listing across Swordfish.Ai directory structures.
Extract headcount, industry classification, and HQ locations from public profiles.
Map job titles, current companies, and public social URLs from professional listings.
Automate search queries to extract targeted lists of companies or professionals based on specific parameters.
Traverse deep directory structures with automated pagination and infinite scroll resolution.
Execute full Playwright sessions to capture dynamically loaded contact modules and profile data.
Bypass rate limits and CAPTCHAs using residential proxies and humanised request patterns.
Standardise extracted data into clean, CRM-ready formats with consistent field typing.
Maintain state across runs to only emit updated profiles and new directory additions.
Brief in. Clean data out.
Provide target industries, company lists, or directory URLs. We design the extraction schema together.
We configure Scrapy crawlers, Playwright sessions, and proxy rotation for swordfish.ai.
Schema validation, null-rate checks, and sample data delivery before full production launch.
JSON / CSV / Parquet pushed to your S3 bucket or Snowflake stage on agreed cadence.
Swordfish.Ai uses strict rate limiting and bot detection. Here is how we stay resilient.
Directory sites monitor request velocity and IP reputation. Our crawlers use residential ISP proxies with realistic browser fingerprints and full cookie session management — trained on real user behaviour patterns.
Swordfish.Ai profile pages are heavily JavaScript-rendered. We run full Playwright browser sessions with JavaScript execution to capture data that headless HTTP clients miss entirely.
Our selector strategy uses multiple fallback chains per field — CSS selectors, XPath, and text-pattern matching — so a layout change doesn't break your data pipeline overnight.
We distribute requests across thousands of IPs with randomised timing delays, ensuring we stay well below Swordfish.Ai's threshold for blocking or CAPTCHA triggers.
For large directories, we maintain a hash index of last-seen values per field. Subsequent runs only push diffs — reducing compute cost and downstream processing load.
B2B sales teams enrich their existing CRM records with updated firmographics and professional metadata.
Marketing teams extract targeted company lists from public directories for outbound campaigns.
Analysts aggregate industry directories to map out sector landscapes and identify total addressable market size.
Strategy teams track competitor headcount growth and hiring patterns across specific departments.
Recruiters build candidate pipelines by extracting professional profiles from targeted company lists.
Other data providers use public directory extracts to cross-reference and validate their own entity resolution models.
"Swordfish.Ai aggregates massive volumes of professional contact data, but extracting it systematically requires infrastructure built for evasion and scale."
Most teams underestimate the investment required: reliable directory scraping requires residential proxies, full JavaScript rendering, CAPTCHA handling, daily selector maintenance, and anomaly monitoring. DataFlirt absorbs that complexity so your engineers can focus on data integration — not infrastructure.
Everything supported by our swordfish.ai scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows.
We maintain pools of residential ISP proxies. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.
Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About swordfish.ai scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available directory information is generally permissible under applicable law, reinforced by the hiQ v. LinkedIn ruling. DataFlirt targets only public, non-authenticated company profiles and directory metadata. We do not circumvent authentication walls or extract gated personal contact details. Clients should consult legal counsel for specific use cases.
No. Direct phone numbers and verified emails on Swordfish.Ai are gated behind paid credits and user authentication. We only extract the publicly available directory metadata, firmographics, and public profile links.
We use residential ISP proxies, full Playwright browser sessions with realistic fingerprints, and request timing modelled on human behaviour to avoid triggering rate limits or CAPTCHAs.
Full directory refreshes can be configured at weekly or monthly cadences depending on volume. Incremental updates run daily to capture new listings.
Our smallest packages start at a defined list of 10,000 company profiles with weekly delivery. For full directory extraction, we price based on volume and delivery frequency.
Absolutely. We provide a sample run of up to 500 profiles as part of the pre-engagement scoping process — so you can validate schema fit and data quality.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off company list extraction or a continuous directory monitoring feed — we scope, build, and operate the pipeline. Tell us what you need.