Data pipelines at volume
Pipelines that handle millions of requests a day and keep going when sources change, slow down, or fail, with monitoring so you hear about a break before your users do.
Data engineer, founder of many ideas
One person, no agency layer. You talk to the engineer who writes the code.
Ten years in data engineering, including technical leadership at Crawlbase, running pipelines that moved billions of pages. These days I split my time between client work on pipelines and APIs, and products I use myself.
Client projects, my own products, and the experiments in between.
Pipelines that handle millions of requests a day and keep going when sources change, slow down, or fail, with monitoring so you hear about a break before your users do.
Fetchers and parsers that ship real data: scheduling, pagination, auth, retries, and rate limits handled properly. Not a demo script that stops at the first page.
For pipelines that fail quietly: retry and backoff policies, rate-limit budgets, clear failed states instead of empty files, and alerts on stale data, not just crashed jobs.
Models where they remove toil: cleaning messy text, structured extraction, classification, and multi-step research agents. Added after the data feeding them is trustworthy.
HTTP APIs, admin tools, and glue code in whatever stack fits, so your team can trigger a run, see what happened, and trust the result.
Ingest, clean, enrich, store, deliver. Data lands in your queue, warehouse, or S3 bucket in a shape people can actually use.

About
I’ve spent about ten years building data systems. The fetcher is the easy part. The real work is everything around it: retries, scheduling, monitoring, and delivery that still shows up on time when a source changes overnight.
At Crawlbase I led architecture for large-scale data pipelines. I’ve also shipped client projects in e-commerce, real estate, finance, and travel, end to end: integrations, scheduling, automations, and datasets that land where the team needs them.
I also turn my own ideas into SaaS and micro SaaS products, usually built around a problem I hit in my own work. AI agents are part of how I build every day, and I write about what they actually speed up and where they get in the way.
And I believe public data should stay public: open for anyone to use and build a business on, not only the platform that hosts it. Here’s why.