We extract tutor profiles, real-time availability, teaching styles, and language combinations from Cambly. Delivered as clean JSON, CSV, or Parquet to S3, BigQuery, or Snowflake on your cadence.
Structured, schema-consistent data across all major object types — delivered clean, typed, and ready to query.
Complete list of extractable fields for Tutor Profiles objects from cambly.com. All fields typed and schema-versioned.
"tutor_id": "60a8f9b2e4b0a1d2", "name": "Sarah Jenkins", "accent": "British", "rating": 4.9, "super_tutor": true, "languages_spoken": "['English (Native)', 'French (Conversational)']", "total_chats": 3412
| # | tutor_id | name | accent | rating | super_tutor | intro_video_url |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Schedules & Availability objects from cambly.com. All fields typed and schema-versioned.
"tutor_id": "60a8f9b2e4b0a1d2", "date": "2026-10-15", "time_slot": "14:30:00", "timezone": "UTC", "is_booked": false, "duration_minutes": 30, "last_updated": "2026-10-14T08:12:00Z"
| # | tutor_id | date | time_slot | timezone | is_booked | reservation_type |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Course Catalogue objects from cambly.com. All fields typed and schema-versioned.
"course_id": "c_biz_eng_101", "title": "Business English Fundamentals", "level": "Intermediate", "category": "Professional Development", "lesson_count": 12, "tags": "['Business', 'Vocabulary', 'Speaking']", "syllabus_topics": "['Introductions', 'Email Etiquette', 'Meetings']"
| # | course_id | title | level | category | description | lesson_count |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Tutor Ratings objects from cambly.com. All fields typed and schema-versioned.
"tutor_id": "60a8f9b2e4b0a1d2", "average_rating": 4.9, "review_count": 1205, "top_tags": "['Patient', 'Good with beginners', 'Clear pronunciation']", "join_date": "2021-04-12", "response_rate": 98.5, "attendance_rate": 99.1
| # | tutor_id | average_rating | review_count | top_tags | student_feedback_summary | join_date |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Complete list of extractable fields for Pricing Plans objects from cambly.com. All fields typed and schema-versioned.
"plan_id": "p_30m_3d_12mo", "minutes_per_week": 90, "days_per_week": 3, "duration_months": 12, "monthly_price": 85.0, "currency": "USD", "discount_pct": 25
| # | plan_id | minutes_per_week | days_per_week | duration_months | total_price | monthly_price |
|---|---|---|---|---|---|---|
| 1 | ||||||
| 2 | ||||||
| 3 |
Our Cambly scraper handles the dynamic single-page application architecture, infinite scrolling tutor lists, and complex availability grids. We manage the session states and proxy rotation required to extract accurate EdTech data.
Extract bios, teaching specialties, certificate details, intro video URLs, and spoken language arrays for every active tutor on the platform.
Monitor tutor availability grids, booked slots, and open reservations across different timezones with high-frequency polling.
Categorise tutor supply by specific accent tags (British, American, Australian, South African) and native speaker status.
Capture aggregate ratings, review counts, and student-assigned qualitative tags (e.g. 'Patient', 'Grammar expert').
Extract the complete taxonomy of Cambly courses, including syllabus breakdowns, lesson counts, and target proficiency levels.
Track the distribution of SuperTutor badges to analyse platform quality metrics and top-tier tutor retention.
Monitor subscription costs across different tier combinations (minutes per day, days per week, commitment length) and regional currencies.
Use region-specific proxies to observe localised pricing, promotional banners, and region-locked tutor availability.
Receive isolated updates when a tutor changes their schedule, updates their bio, or gains a new certification.
Brief in. Clean data out.
Provide target languages, accent preferences, or schedule frequencies. We design the extraction schema together.
We configure Scrapy and Playwright crawlers, proxy rotation, and session management for cambly.com.
Schema validation, null-rate checks, and schedule accuracy verification before full launch.
JSON, CSV, or Parquet pushed to your S3 bucket, BigQuery dataset, or Snowflake stage on agreed cadence.
Extracting data from modern SPA EdTech platforms requires JavaScript rendering and session persistence. Here is how we maintain reliable extraction.
Cambly's tutor search and schedule grids are heavily JavaScript-rendered. We run full Playwright browser sessions with lazy-load triggering to capture complete availability data that basic HTTP clients miss.
We utilise residential ISP proxies with realistic browser fingerprints and randomised request timing to navigate rate limits and IP bans during high-frequency schedule polling.
Our selector strategy uses multiple fallback chains per field. If a UI update changes the CSS class for the SuperTutor badge or schedule grid, XPath and text-pattern fallbacks ensure continuous data flow.
For tracking tutor schedules, we maintain a hash index of last-seen availability. Subsequent runs only push diffs, reducing compute cost and downstream processing load.
Every run emits structured logs to our observability stack. We alert on null-rate spikes in critical fields like tutor ratings or schedule slots, addressing issues before delivery.
Language learning platforms monitor Cambly's tutor supply, pricing structures, and course offerings to benchmark their own services.
EdTech analysts track the geographic distribution, accent diversity, and certification levels of active tutors to understand supply-side dynamics.
Companies track regional subscription pricing and discount cadences to optimise their own promotional calendars.
Investors evaluate platform health by tracking active tutor counts, review velocity, and schedule density over time.
By analysing booked versus available slots across different timezones, researchers model student demand patterns.
Extract structured profiles and course descriptions to train educational recommendation engines and matching algorithms.
"Tutor availability and profile metadata are the core assets of any synchronous EdTech platform. Extracting this at scale reveals the exact supply-and-demand mechanics of the network."
Building a reliable scraper for Cambly requires handling complex JavaScript states, infinite scrolling, and aggressive rate limits on schedule API endpoints. DataFlirt manages the proxy rotation, session handling, and schema maintenance so your data engineering team receives clean, normalised records ready for analysis.
Everything supported by our cambly.com scraper — rendered SPA elements, auth walls, rate-limit evasion and beyond.
Open-source tooling on proven cloud infra — no vendor lock-in, full observability.
Scrapy handles crawl orchestration and deduplication. Playwright handles JavaScript rendering, cookie sessions, and SPA interaction flows.
We maintain pools of residential ISP proxies. Rotation happens per request with sticky sessions where required for complex navigation.
Pipelines run on AWS Lambda and ECS. Airflow handles scheduling, dependency management, and SLA alerting. All state is stored in managed Postgres.
Data delivered to where your team already works — no new tooling required.
About cambly.com scraping, legality, and pipeline operations.
Ask us directly →Scraping publicly available information from Cambly is generally permissible under applicable law. DataFlirt targets only public, non-authenticated tutor profiles, schedules, and course data. We do not extract personal student data, circumvent authentication walls for private lessons, or violate GDPR. Clients should consult legal counsel for specific use cases.
We use residential ISP proxies, full Playwright browser sessions, and request timing modelled on human behaviour. We distribute schedule polling across multiple IPs to stay within acceptable request thresholds.
Real-time streaming pipelines can achieve sub-15-minute latency for availability signals on a defined subset of tutors. Full platform refreshes are typically executed at daily or weekly cadences.
Yes. Every pipeline run produces timestamped snapshots. We maintain a time-series table per tutor ID for rating averages, review counts, and badge statuses.
Our smallest packages start at a defined subset of tutors (e.g., specific accents or languages) with weekly delivery. For full platform extraction, we price based on volume and delivery frequency.
Yes. We provide a sample run of up to 200 tutor profiles and their associated schedules as part of the pre-engagement scoping process to validate schema fit and data quality.
20-minute scoping call. Pilot dataset within the week. Production within two. Whether you need a one-off tutor directory export or a continuous schedule-monitoring feed, we scope, build, and operate the pipeline. Tell us your requirements.