SYSTEM all green source signalhire.com queue 18,492 pages p99 latency 214ms dataflirt.com · scraper/signalhire-com
RUN * 114 active pipelines * signalhire.com live

SignalHire data,
delivered at scale.

Extract candidate profiles, work histories, skill graphs, and company directories from SignalHire. Delivered as clean JSON, CSV, or Parquet to your warehouse.

Profiles extracted
1.2M /day
Company records
45,310 /24h
Updates processed
312K /run
Active pipelines
114
Uptime
99.94%
Data Dictionary

Every field we extract from signalhire.com

Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.

Complete list of extractable fields for Candidate Profiles objects from signalhire.com. All fields typed and schema-versioned.

profile_idfull_namecurrent_titlecurrent_companylocationindustrysummaryprofile_urlconnection_countsocial_links
candidate_profiles
● 200 OK
"profile_id": "sh_9823471",
"full_name": "Arjun Patel",
"current_title": "Senior Software Engineer",
"current_company": "TechCorp India",
"location": "Bengaluru, Karnataka",
"industry": "Information Technology",
"connection_count": 500,
"profile_url": "https://www.signalhire.com/profiles/arjun-patel-9823471"
# profile_idfull_namecurrent_titlecurrent_companylocationindustry
1
2
3

Complete list of extractable fields for Work History objects from signalhire.com. All fields typed and schema-versioned.

profile_idcompany_nametitlestart_dateend_dateduration_monthsdescriptionlocationcompany_url
work_history
● 200 OK
"profile_id": "sh_9823471",
"company_name": "TechCorp India",
"title": "Senior Software Engineer",
"start_date": "2021-04-01",
"end_date": "present",
"duration_months": 38,
"location": "Bengaluru"
# profile_idcompany_nametitlestart_dateend_dateduration_months
1
2
3

Complete list of extractable fields for Education objects from signalhire.com. All fields typed and schema-versioned.

profile_idinstitution_namedegreefield_of_studystart_yearend_yearactivitiesdescription
education
● 200 OK
"profile_id": "sh_9823471",
"institution_name": "Visvesvaraya Technological University",
"degree": "Bachelor of Engineering",
"field_of_study": "Computer Science",
"start_year": "2015",
"end_year": "2019",
"activities": "Coding Club President"
# profile_idinstitution_namedegreefield_of_studystart_yearend_year
1
2
3

Complete list of extractable fields for Company Data objects from signalhire.com. All fields typed and schema-versioned.

company_idcompany_namewebsiteindustryemployee_count_rangeheadquartersfounded_yeardescriptionsignalhire_url
company_data
● 200 OK
"company_id": "comp_44512",
"company_name": "TechCorp India",
"website": "techcorp.in",
"industry": "Information Technology",
"employee_count_range": "1001-5000",
"headquarters": "Bengaluru, Karnataka",
"founded_year": "2010"
# company_idcompany_namewebsiteindustryemployee_count_rangeheadquarters
1
2
3

Complete list of extractable fields for Skills objects from signalhire.com. All fields typed and schema-versioned.

profile_idskill_nameendorsement_countcategoryprimary_skillyears_experienceverifiedadded_date
skills
● 200 OK
"profile_id": "sh_9823471",
"skill_name": "Python",
"endorsement_count": 42,
"category": "Programming Languages",
"primary_skill": true,
"verified": false,
"years_experience": 5
# profile_idskill_nameendorsement_countcategoryprimary_skillyears_experience
1
2
3

Capabilities

Complete talent intelligence extraction

Our SignalHire scraper handles complex search pagination, profile rendering, and anti-bot mitigation to deliver structured professional data at volume.

Profile Extraction

Extract names, titles, locations, summaries, and social links from public candidate profiles.

Company Directories

Map corporate structures, employee counts, headquarters, and industry classifications.

Work History Parsing

Capture chronological employment records, titles, durations, and role descriptions.

Education Mapping

Extract degrees, institutions, study fields, and graduation years for candidate profiling.

Skill Graphs

Compile technical and soft skills listed on profiles to build searchable talent pools.

Social Cross-referencing

Collect associated public social URLs provided within the SignalHire profile.

Search SERP Scraping

Paginate through SignalHire search results based on specific job titles, locations, or companies.

Alumni Tracking

Extract lists of current and past employees for specific target organisations.

Delta Updates

Monitor specific profiles for job changes, promotions, or location updates over time.

// engagement pipeline

From search URL to warehouse record

Brief in. Clean data out.

Define Scope
d 0

Provide target companies, job titles, locations, or specific SignalHire profile URLs.

Pipeline Build
d 2–4

We configure extraction logic, proxy rotation, and CAPTCHA handling for signalhire.com.

Validation & QA
d 4–6

Schema validation, null-rate checks, and data normalisation before full launch.

Delivery
ongoing

Clean structured data pushed to your warehouse or object storage on schedule.

Under the hood

Bypassing directory rate limits

SignalHire restricts high-volume profile viewing. We manage session rotation and IP distribution to maintain continuous extraction.

pipeline-monitor · signalhire.com · live ● active
// fingerprinting
Identity rotation
TLS fingerprintrandomised
User-agentrotated
IP poolresidential
Challenges blocked0
// pagination
Page coverage
48,291 pages queued running
// observability
Pipeline health
99.9%
uptime
142ms
p99 lat
0.3%
null rate
2
alerts
Proxy rotation
Residential IP distribution

SignalHire limits profile views per IP. We distribute requests across a global pool of residential proxies to maintain low request rates per node.

Session management
Cookie and header spoofing

We simulate real browser sessions with accurate headers, user agents, and TLS fingerprints to avoid automated detection triggers.

DOM parsing
Resilient selector chains

SignalHire updates its frontend structure. We use multi-layered selectors to ensure data extraction continues even when CSS classes change.

Delta diffing
Efficient update tracking

For ongoing monitoring, we hash profile states and only extract full records when changes are detected, reducing bandwidth and processing time.

Alerting
Automated health checks

Pipelines emit telemetry to Grafana. We monitor success rates and trigger automatic interventions if extraction yields drop.

Applications

Who uses SignalHire data

Teams across industries use signalhire.com data to build competitive products and smarter operations.

01
Talent Sourcing

Recruitment agencies build internal databases of passive candidates mapped by skill, location, and experience.

02
HR Tech Platforms

Applicant tracking systems enrich partial candidate profiles with full work histories and educational backgrounds.

03
B2B Sales Prospecting

Sales teams map target accounts, identifying decision-makers and their career trajectories.

04
Market Mapping

Consultancies analyse talent migration between competitors and identify emerging skill clusters in specific regions.

05
Competitor Analysis

Organisations track competitor hiring velocity and department expansion through employee headcount changes.

06
Investment Due Diligence

Private equity firms evaluate executive team backgrounds and engineering talent quality before acquisitions.

Why DataFlirt

"SignalHire contains millions of professional profiles and company structures, but extracting this graph requires distributed infrastructure."

Building a reliable pipeline for SignalHire means handling aggressive rate limits, CAPTCHA walls, and frequent DOM changes. DataFlirt manages the proxies, sessions, and extraction logic so you receive clean data without maintaining infrastructure.

Technical Spec

SignalHire scraper technical specifications

Everything supported by our signalhire.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.

Public profile parsing
Extract all publicly visible fields from candidate profiles
Supported
Company directory scraping
Capture corporate metadata and employee lists
Supported
Search result pagination
Iterate through multi-page search results automatically
Supported
Skill extraction
Parse technical and soft skills lists per profile
Supported
Work history mapping
Chronological array of past and present roles
Supported
Delta extraction
Only output profiles that changed since the last run
Supported
Webhook delivery
Real-time HTTP POST for individual profile updates
Supported
Direct email extraction
Requires authenticated SignalHire credits to reveal gated contact info
Partial
Direct phone extraction
Requires authenticated SignalHire credits to reveal gated contact info
Partial
Infrastructure

Infrastructure powering the pipeline

Open-source tooling on proven cloud infra — no vendor lock-in, full observability.

ScrapyPlaywrightPython 3.12RedisPostgreSQLApache AirflowAWS LambdaS3CloudWatch2CaptchaCapSolverResidential ProxiesDockerKubernetesGrafanaPrometheus
Scrapy + Playwright Stack

Scrapy handles crawl orchestration, deduplication, and retry logic. Playwright handles JavaScript rendering, cookie sessions, and interaction flows. Combined via scrapy-playwright middleware.

Residential Proxy Infrastructure

We maintain pools of residential ISP proxies across global regions. Rotation happens per-request with sticky sessions where required. IP score monitoring prevents blacklisted pool contamination.

Cloud-Native Orchestration

Pipelines run on AWS Lambda (burst) and ECS (sustained). Airflow handles scheduling, dependency management, and SLA alerting. All state stored in managed Postgres.

Output & Delivery

Your data, your destination

Data delivered to where your team already works — no new tooling required.

JSON
Newline-delimited or nested - schema versioned per run
CSV
Flat file with typed columns - Excel/Sheets compatible
XLS
Excel format for direct business user consumption
Parquet
Columnar format for BigQuery, Snowflake, Athena
AWS S3
Direct bucket delivery - compatible with any data lake
Webhook
HTTP POST per record for real-time downstream processing
API
REST endpoint to query extracted datasets on demand
PostgreSQL
Direct database upserts with conflict resolution
S3
Direct bucket delivery — compatible with any data lake
// faq

Common questions.

About signalhire.com scraping, legality, and pipeline operations.

Ask us directly →
What data can be extracted from SignalHire?

We extract publicly visible data including candidate names, job titles, work histories, educational backgrounds, skills, and company metadata. We do not bypass authentication to extract gated contact information like private emails or phone numbers without client-provided API credits.

How do you handle SignalHire rate limits?

We use a distributed network of residential proxies and manage session cookies to distribute request volume. This prevents IP bans and ensures continuous data extraction at scale.

Can you track job changes over time?

Yes. We can configure delta pipelines that monitor a specific list of profiles or companies, returning data only when a candidate updates their current role or location.

Do you extract contact information?

We extract public social links visible on the profile. For gated emails and phone numbers, SignalHire requires paid credits. If you provide an authenticated account or API key, we can integrate that retrieval into the pipeline.

How is the data structured?

Data is highly normalised. Work histories and educational backgrounds are delivered as nested arrays in JSON, or flattened into relational tables for CSV and Parquet formats.

What is the typical turnaround time?

For standard profile extraction based on search URLs, pipelines can be deployed within 48 hours. Delivery cadence can be set to daily, weekly, or monthly depending on your requirements.

Is a sample available?

Yes. We offer a sample extraction of up to 500 profiles based on your target criteria to validate schema fit and data quality before formal engagement.

$ dataflirt scope --new-project --source=signalhire.com ready

Tell us what
to extract.
We do the rest.

20-minute scoping call. Pilot dataset within the week. Production within two. From targeted company alumni lists to bulk talent extraction, we build and operate the pipeline. Tell us your requirements.

hello@dataflirt.com · Bengaluru · IST · typical reply < 4h
Related Scrapers

More in business directories

Services

Data Extraction for Every Industry

View All Services →