SYSTEM all green source cognism.com queue 18,492 profiles p99 latency 214ms dataflirt.com · scraper/cognism-com
RUN · 41 active pipelines · cognism.com live

B2B intelligence,
at warehouse scale.

We extract public company profiles, firmographics, and industry taxonomies from Cognism directories. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.

Companies extracted
450K /day
Profile updates
1.2M /week
Directory pages
85K /run
Active pipelines
41
Uptime
99.98%
Data Dictionary

Every field we extract from cognism.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Company Firmographics objects from cognism.com. All fields typed and schema-versioned.

company_namecognism_urlwebsiteindustrysub_industryemployee_count_rangehq_locationfounded_yeardescriptionlinkedin_url
company_firmographics
● 200 OK
"company_name": "Acme Corp",
"industry": "Software Development",
"employee_count_range": "501-1000",
"hq_location": "London, UK",
"founded_year": 2012,
"website": "acme.com"
# company_namecognism_urlwebsiteindustrysub_industryemployee_count_range
1
2
3

Complete list of extractable fields for Public Employee Roster objects from cognism.com. All fields typed and schema-versioned.

company_nameemployee_namejob_titledepartmentlocationprofile_urlseniority_levelstart_date
public_employee roster
● 200 OK
"employee_name": "Jane Doe",
"job_title": "VP of Engineering",
"department": "Engineering",
"seniority_level": "VP",
"location": "New York, USA",
"company_name": "Acme Corp"
# company_nameemployee_namejob_titledepartmentlocationprofile_url
1
2
3

Complete list of extractable fields for Location Data objects from cognism.com. All fields typed and schema-versioned.

company_namehq_addresshq_cityhq_statehq_countryhq_postal_coderegional_officesoffice_count
location_data
● 200 OK
"hq_city": "London",
"hq_country": "United Kingdom",
"office_count": 4,
"hq_postal_code": "EC1A 1BB",
"company_name": "Acme Corp",
"regional_offices": "['New York', 'Berlin']"
# company_namehq_addresshq_cityhq_statehq_countryhq_postal_code
1
2
3

Complete list of extractable fields for Industry Taxonomy objects from cognism.com. All fields typed and schema-versioned.

category_namecategory_urlparent_categorytotal_companiesrelated_categoriesdescriptiontop_companieslast_updated
industry_taxonomy
● 200 OK
"category_name": "Fintech",
"parent_category": "Financial Services",
"total_companies": 14205,
"related_categories": "['Payments', 'Insurtech']",
"top_companies": "['Stripe', 'Revolut']",
"last_updated": "2023-10-12T00:00:00Z"
# category_namecategory_urlparent_categorytotal_companiesrelated_categoriesdescription
1
2
3

Complete list of extractable fields for Technographics objects from cognism.com. All fields typed and schema-versioned.

company_nametech_categorytechnology_nameimplementation_statusdetected_dateconfidence_scoresource_urlrelated_tools
technographics
● 200 OK
"company_name": "Acme Corp",
"tech_category": "CRM",
"technology_name": "Salesforce",
"implementation_status": "Active",
"detected_date": "2023-09-15",
"confidence_score": 0.95
# company_nametech_categorytechnology_nameimplementation_statusdetected_dateconfidence_score
1
2
3

Capabilities

Extract firmographics at scale

Our Cognism scraper navigates public directories to extract company profiles, taxonomies, and firmographic indicators. We manage the proxies, rendering, and schema maintenance.

Public Profile Extraction

Extract company names, firmographics, descriptions, and metadata from public-facing directory pages.

Industry Taxonomy Mapping

Traverse industry categories and sub-categories to build comprehensive market maps.

Employee Count Tracking

Monitor headcount ranges and growth signals across target accounts over time.

Location Intelligence

Capture HQ locations, regional offices, and geographic footprint data for territory planning.

JavaScript Rendering

Execute Playwright sessions to render dynamic directory content and lazy-loaded elements.

Change Detection

Maintain hash indexes to emit only changed records, reducing downstream processing load.

Anti-Bot Circumvention

Utilise residential proxy pools and TLS fingerprint spoofing to bypass rate limits.

Multi-Region Support

Access localised directory views to capture region-specific company data.

Continuous Pipelines

Schedule daily or weekly syncs to keep your CRM or data warehouse updated.

// engagement pipeline

From directory URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target industries, company sizes, or directory URLs. We design the extraction schema together.

Pipeline Build
d 2–4

We configure Scrapy crawlers, proxy rotation, session management, and CAPTCHA handling for cognism.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, and sample data reviews before full launch.

Delivery
ongoing

JSON / CSV / Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.

Under the hood

How our directory pipeline handles the hard parts

Directory scraping requires navigating rate limits and dynamic pagination. Here is how we maintain reliable extraction.

pipeline-monitor · cognism.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Anti-bot layer
Residential proxy rotation

Directory sites implement strict rate limits. Our crawlers use residential ISP proxies with realistic browser fingerprints and randomised request timing to distribute load.

JavaScript rendering
Playwright execution for SPA content

Modern directories rely heavily on JavaScript. We run full Playwright browser sessions to trigger lazy-loading and capture data that headless HTTP clients miss.

Schema stability
Resilient selectors

DOM structures change. Our selector strategy uses fallback chains combining CSS, XPath, and text-pattern matching to ensure pipeline stability.

Change detection
Hash-based diffing

For large directories, we maintain a hash index of last-seen values. Subsequent runs only push diffs, providing a clean changelog.

Monitoring
24/7 pipeline health

Every run emits structured logs. We alert on null-rate spikes and schema drift, responding before data quality degrades.

Applications

Who uses B2B directory data — and how

Teams across industries use cognism.com data to build competitive products and smarter operations.

01
TAM Expansion

Revenue operations teams map total addressable markets by extracting companies within specific industry taxonomies.

02
Competitor Intelligence

Strategy teams monitor competitor growth signals, employee count changes, and geographic expansion.

03
Account Scoring

Data engineering teams enrich internal CRM records with fresh firmographics to improve lead scoring models.

04
Territory Planning

Sales leadership segments target accounts by HQ location and regional office presence.

05
Market Research

Analysts track industry categorisation trends to identify emerging sectors and whitespace.

06
Investment Sourcing

Venture capital firms identify fast-growing companies based on headcount expansion and firmographic signals.

Why DataFlirt

"Cognism's public directories hold massive firmographic value, but mapping millions of companies requires infrastructure built for scale and anti-bot resilience."

Most data teams underestimate the complexity of directory scraping. Rate limits, CAPTCHAs, and dynamic DOM structures break naive scripts. DataFlirt manages the proxies, browser rendering, and schema maintenance so you receive structured firmographics directly in your warehouse.

Technical Spec

Cognism scraper — technical capabilities

Everything supported by our cognism.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

JavaScript rendering
Full Playwright sessions for dynamic directory content
Supported
CAPTCHA bypass
Automated 2Captcha + CapSolver integration
Supported
Residential proxy rotation
ISP-grade residential IPs rotated per request
Supported
Public company firmographics
Extraction of names, industries, and locations from public pages
Supported
Industry taxonomy extraction
Mapping of category hierarchies
Supported
Change detection (diffs)
Hash-based diff to emit only changed records
Supported
Webhook delivery
HTTP POST per record or batch
Supported
Verified direct dials
Authenticated B2B contact phone numbers
Partial
Direct email addresses
Authenticated B2B contact email addresses
Partial
Authenticated intent data
Buyer intent signals behind the login wall
Partial
Infrastructure

Infrastructure powering the directory pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles orchestration and deduplication. Playwright handles JavaScript rendering and interaction flows.

Residential Proxy Infrastructure

Pools of residential ISP proxies rotate per request to bypass directory rate limits.

Cloud-Native Orchestration

Pipelines run on AWS Lambda and ECS. Airflow handles scheduling and SLA alerting.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested schema
CSV
Flat file with typed columns
XLS
Excel compatible format for business teams
Parquet
Columnar format for data warehouses
AWS S3
Direct bucket delivery
Webhook
HTTP POST per record
API
REST endpoint for querying extracted data
PostgreSQL
Upsert into your existing schema
BigQuery
Streamed directly into your dataset
Snowflake
Stage and COPY INTO workflow
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About cognism.com scraping, legality, and pipeline operations.

Ask us directly →
Is scraping Cognism public directories legal?

Scraping publicly available firmographic information is generally permissible under applicable law. DataFlirt targets only public, non-authenticated directory pages. We do not extract personal data behind paywalls or circumvent authentication mechanisms.

How do you handle rate limits?

We use residential ISP proxies and request timing modelled on human behaviour to distribute load and avoid IP bans.

Can I get direct dials and emails?

No. DataFlirt only extracts data available on public directory pages. We do not support extracting authenticated contact data such as direct dials or verified emails.

How fresh is the directory data?

We can configure pipelines to run at daily, weekly, or monthly cadences depending on your freshness requirements.

Do you support historical tracking?

Yes. Every pipeline run produces timestamped snapshots. We maintain a time-series record for changes in employee counts and firmographics from the date your pipeline starts.

What is the minimum viable engagement?

Our minimum engagement typically starts at a defined list of 10,000 target companies or specific industry categories. Contact us for a scoped quote.

Can I request a sample dataset?

Yes. We provide a sample run of up to 500 company profiles as part of the pre-engagement scoping process.

$ dataflirt scope --new-project --source=cognism.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off firmographic dump or continuous tracking across 500K companies — we scope, build, and operate the pipeline.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in business directories

Services

Data Extraction for Every Industry

View All Services →