Data Engineer — Creator Intelligence
Keep the data engine humming — ingest, enrich and score millions of creator profiles reliably and cheaply.
SocialDB is the open database and marketplace for the creator economy, and the data is the moat. We discover, ingest, refresh and enrich 1.4M+ creator profiles across a dozen platforms — plus 24,000+ brands — and every product feature is built on top of it. As a Data Engineer, you own the pipelines that keep that engine humming: comprehensive, fresh, deduplicated, and cheap to run.
What you'll do
- Build and maintain resilient ingestion and enrichment jobs across a dozen platforms — official APIs, RSS, and robust scraping — that run cheaply at scale
- Own data quality: deduplication, fake-follower detection, audience-overlap and engagement scoring, and email discovery
- Design schemas and indexes that keep queries fast as the dataset grows well into the millions
- Make the pipeline observable and self-healing: sane retries, rate-limit handling, sharding, and hard cost/compute controls
- Partner with product to turn raw data into features — rankings, benchmarks, recommendations
- Keep throughput high without ever taking the site down
What success looks like
- 90 days: you own several importers end-to-end and have measurably improved coverage or freshness
- 6 months: the pipeline is more comprehensive, cheaper and more reliable than when you arrived, with quality metrics to prove it
What we're looking for
- 3+ years in data engineering or backend-heavy roles
- Strong SQL and a pragmatic, cost-aware approach to scale — you optimize before you throw hardware at it
- Comfortable with messy real-world data, APIs, rate limits and scraping
- You care about correctness and dedup as much as raw volume
Nice to have
- Experience with social-platform data or large web-scraping systems
- PHP/MySQL familiarity (our stack), or quick to get productive in it
- Interest in ranking, scoring or ML on top of the data
Why join
- Own the data moat of the whole company — the foundation everything else is built on
- Full-time or contract, whichever fits you
- Fully remote, with a direct line to the founder; competitive salary or contract rate (plus equity for full-time)
- Genuinely hard, interesting data problems at real scale on a lean budget
How to apply
Send your CV or LinkedIn and a short note on a data pipeline you've built — the scale, the sources, and how you kept it reliable and cheap. Links to relevant work are welcome. We reply to every serious application.