data engineering · distributed systems · infrastructure

I build high-throughput data engineering pipelines, distributed systems, and scalable AI infrastructure.

Building enterprise data architectures, resilient pipelines, and high-performance backend systems.

On the side of the desk: helping teams ship data pipelines that touch billions of pages: not only fetchers and parsers, but the full stack that tackles data resiliency: network performance, distributed queuing, routing optimization, and advanced error recovery infrastructure that turns a hard “no” into a reliable feed.

Billions of records, built to last

End-to-end work: I build data pipelines and the systems that ensure their resiliency: custom layers for complex APIs, pragmatic network paths, request orchestration, and clean structured delivery you can plug straight into prod.

  • ···Data records processed & parsed across client projects
  • ···Resilient HTTP journeys: retries, network optimization, redirects & proxies
  • ···AI extraction, clean-up & structured output passes

Cumulative numbers across engineering programs and client projects. Need data pipelines, systems architecture, or both? Get in touch.

What I’m usually doing

Data pipelines, distributed systems, and the glue in between: day job, consulting, and side experiments.

  • High-Volume Data Engineering

    High-volume extraction that stays up under load: millions of requests, shifting targets, fragile sites, and the chaos of the open web, with observability so you know when something breaks.

  • Data Integration Pipelines

    Resilient fetchers that ship real data: parsing, scheduling, retries, and the hard edge of fault-tolerant HTTP requests, not demo scripts that stop at hello world.

  • System Hardening & Request Orchestration

    The systems that ensure data resiliency: request retry frameworks, API rate-limit management, status-code recovery optimization, and resilient network layer headers when off-the-shelf tools are not enough.

  • AI & agents

    Models plus agentic flows: tools, planners, and retries that summarize messy pages, structure extraction, and run multi-step crawl and research tasks without a human in every loop.

  • Automations & APIs

    Glue code, internal tools, and HTTP APIs in whatever stack fits: wiring data pipelines, orchestration jobs, and backends so teams can trigger, observe, and trust the pipeline.

  • Data pipelines

    Ingest, clean, enrich, store, and deliver, so scraped and automated output lands in queues, warehouses, or products people actually use.

Jamal Awad, portrait sketch
Jamal AwadBuilding enterprise data architectures

About

Data pipelines, distributed systems, and resilient infrastructure

I’ve spent the better part of a decade building data engineering systems from the ground up, but the interesting part is rarely only the fetcher: it’s the layers that ensure resiliency against network failures, adapt when APIs change, and still hand you clean, structured data.

I’ve led architecture for large-scale data pipelines and network performance optimization: request orchestration, recovery paths, retry logic, and pipelines that tie fetch, process, and deliver together. I’ve also shipped complete projects for clients across e-commerce, real estate, finance, and travel, end-to-end: integration pipelines, custom orchestration work, automations, scheduling, and production-ready datasets.

These days I still experiment with AI agents and agentic workflows: where models genuinely help extraction and research, vs. where they’re noise on top of a broken fetch path.

  • 10+ years in data engineering & systems architecture
  • Billions of records and complex integration paths
  • Full stack: ingest, process, delivery
  • Remote, working with teams worldwide

What people sayCompany names are kept confidential. Most of this feedback came from sensitive data engineering projects where clients prefer not to disclose the collaboration publicly.

From teams across data pipelines, SaaS, databases, APIs, and consulting engagements.

Data Engineering

“We needed 40M product pages processed weekly across six markets. Jamal built the large-scale information pipelines, hardening layer for the most complex high-performance APIs, and delivered clean CSVs to our S3 bucket on schedule. Our data team finally stopped complaining.”

Head of Data, EU e-commerce platform40M+ pages/week project
Data Engineering

“Our stack was failing on 60% of targets after network updates. Jamal didn't just patch data ingestion systems; he rebuilt the data pipeline in two weeks: request retry frameworks, API rate-limit management, status-code recovery optimization, and custom logic where generic tools died. Failure rate dropped to under 3%.”

CTO, price intelligence startupData Pipeline Hardening & Optimization
Data Engineering

“Jamal ships the full loop: scheduling, monitoring, alerts, and API delivery. But the difference was the resilient network layer headers: when everyone else shrugged at status-code errors, he had a system. Fourteen months, zero missed deliveries.”

VP Engineering, real estate data companyEnd-to-end data & network infrastructure
Database & HA

“Our MySQL replicas were drifting 45 seconds behind under write-heavy load. Jamal restructured the replication topology, tuned InnoDB buffer pools, and moved the hottest tables to a dedicated cluster. Lag dropped to sub-second and stayed there through Black Friday.”

DBA Lead, fintech marketplaceMySQL replication & performance
SaaS & Platform

“We were building multi-tenant SaaS from a single-tenant Rails monolith. Jamal designed the tenant isolation layer, data partitioning scheme, and migration path. We onboarded 200 tenants without a single data leak or downtime window.”

CEO, HR tech startupMulti-tenant SaaS architecture
API & Performance

“P95 latency on our public API had crept to 1.8 seconds. Jamal profiled the hot paths, added Redis caching with smart invalidation, restructured three N+1 query patterns, and pushed us down to 180ms. Customers noticed before we even announced it.”

Product Lead, logistics API providerAPI latency reduction
Database & HA

“After a major outage took our PostgreSQL primary down for two hours, we brought Jamal in to build a proper HA setup. Patroni cluster, automated failover, WAL archiving to S3, and a runbook the on-call team actually follows. We've had three hardware failures since then. Zero downtime.”

Infrastructure Lead, media streaming platformPostgreSQL high availability
Data Engineering

“We needed real estate listings from 14 fragmented MLS sources, each with different auth, pagination, and complex rate limits. Jamal built adapters for all of them, a unified schema, and deduplication logic. Our agents got a single clean feed for the first time ever.”

CTO, proptech startupMulti-source data aggregation
SaaS & Platform

“Moving from a monolith to event-driven microservices felt impossible with our team size. Jamal carved the domain boundaries, set up Kafka topics with proper schemas, and migrated the first three services while keeping the monolith running. The rest of the team picked up the pattern and kept going.”

Engineering Manager, e-commerce companyEvent-driven architecture migration
Data Engineering

“We needed 40M product pages processed weekly across six markets. Jamal built the large-scale information pipelines, hardening layer for the most complex high-performance APIs, and delivered clean CSVs to our S3 bucket on schedule. Our data team finally stopped complaining.”

Head of Data, EU e-commerce platform40M+ pages/week project
Data Engineering

“Jamal ships the full loop: scheduling, monitoring, alerts, and API delivery. But the difference was the resilient network layer headers: when everyone else shrugged at status-code errors, he had a system. Fourteen months, zero missed deliveries.”

VP Engineering, real estate data companyEnd-to-end data & network infrastructure
SaaS & Platform

“We were building multi-tenant SaaS from a single-tenant Rails monolith. Jamal designed the tenant isolation layer, data partitioning scheme, and migration path. We onboarded 200 tenants without a single data leak or downtime window.”

CEO, HR tech startupMulti-tenant SaaS architecture
Database & HA

“After a major outage took our PostgreSQL primary down for two hours, we brought Jamal in to build a proper HA setup. Patroni cluster, automated failover, WAL archiving to S3, and a runbook the on-call team actually follows. We've had three hardware failures since then. Zero downtime.”

Infrastructure Lead, media streaming platformPostgreSQL high availability
SaaS & Platform

“Moving from a monolith to event-driven microservices felt impossible with our team size. Jamal carved the domain boundaries, set up Kafka topics with proper schemas, and migrated the first three services while keeping the monolith running. The rest of the team picked up the pattern and kept going.”

Engineering Manager, e-commerce companyEvent-driven architecture migration
Data Engineering

“Our stack was failing on 60% of targets after network updates. Jamal didn't just patch data ingestion systems; he rebuilt the data pipeline in two weeks: request retry frameworks, API rate-limit management, status-code recovery optimization, and custom logic where generic tools died. Failure rate dropped to under 3%.”

CTO, price intelligence startupData Pipeline Hardening & Optimization
Database & HA

“Our MySQL replicas were drifting 45 seconds behind under write-heavy load. Jamal restructured the replication topology, tuned InnoDB buffer pools, and moved the hottest tables to a dedicated cluster. Lag dropped to sub-second and stayed there through Black Friday.”

DBA Lead, fintech marketplaceMySQL replication & performance
API & Performance

“P95 latency on our public API had crept to 1.8 seconds. Jamal profiled the hot paths, added Redis caching with smart invalidation, restructured three N+1 query patterns, and pushed us down to 180ms. Customers noticed before we even announced it.”

Product Lead, logistics API providerAPI latency reduction
Data Engineering

“We needed real estate listings from 14 fragmented MLS sources, each with different auth, pagination, and complex rate limits. Jamal built adapters for all of them, a unified schema, and deduplication logic. Our agents got a single clean feed for the first time ever.”

CTO, proptech startupMulti-source data aggregation
Data Engineering

“We needed 40M product pages processed weekly across six markets. Jamal built the large-scale information pipelines, hardening layer for the most complex high-performance APIs, and delivered clean CSVs to our S3 bucket on schedule. Our data team finally stopped complaining.”

Head of Data, EU e-commerce platform40M+ pages/week project
Database & HA

“Our MySQL replicas were drifting 45 seconds behind under write-heavy load. Jamal restructured the replication topology, tuned InnoDB buffer pools, and moved the hottest tables to a dedicated cluster. Lag dropped to sub-second and stayed there through Black Friday.”

DBA Lead, fintech marketplaceMySQL replication & performance
Database & HA

“After a major outage took our PostgreSQL primary down for two hours, we brought Jamal in to build a proper HA setup. Patroni cluster, automated failover, WAL archiving to S3, and a runbook the on-call team actually follows. We've had three hardware failures since then. Zero downtime.”

Infrastructure Lead, media streaming platformPostgreSQL high availability
Data Engineering

“Our stack was failing on 60% of targets after network updates. Jamal didn't just patch data ingestion systems; he rebuilt the data pipeline in two weeks: request retry frameworks, API rate-limit management, status-code recovery optimization, and custom logic where generic tools died. Failure rate dropped to under 3%.”

CTO, price intelligence startupData Pipeline Hardening & Optimization
SaaS & Platform

“We were building multi-tenant SaaS from a single-tenant Rails monolith. Jamal designed the tenant isolation layer, data partitioning scheme, and migration path. We onboarded 200 tenants without a single data leak or downtime window.”

CEO, HR tech startupMulti-tenant SaaS architecture
Data Engineering

“We needed real estate listings from 14 fragmented MLS sources, each with different auth, pagination, and complex rate limits. Jamal built adapters for all of them, a unified schema, and deduplication logic. Our agents got a single clean feed for the first time ever.”

CTO, proptech startupMulti-source data aggregation
Data Engineering

“Jamal ships the full loop: scheduling, monitoring, alerts, and API delivery. But the difference was the resilient network layer headers: when everyone else shrugged at status-code errors, he had a system. Fourteen months, zero missed deliveries.”

VP Engineering, real estate data companyEnd-to-end data & network infrastructure
API & Performance

“P95 latency on our public API had crept to 1.8 seconds. Jamal profiled the hot paths, added Redis caching with smart invalidation, restructured three N+1 query patterns, and pushed us down to 180ms. Customers noticed before we even announced it.”

Product Lead, logistics API providerAPI latency reduction
SaaS & Platform

“Moving from a monolith to event-driven microservices felt impossible with our team size. Jamal carved the domain boundaries, set up Kafka topics with proper schemas, and migrated the first three services while keeping the monolith running. The rest of the team picked up the pattern and kept going.”

Engineering Manager, e-commerce companyEvent-driven architecture migration

Stack & focus areas

Systems architectureSaaS platform designTechnical leadershipDistributed systemsAPI design & integrationData pipelinesLLMs & AI agentsCloud infrastructure (AWS)Container orchestration (K8s)Self-hosted infrastructureEvent-driven architectureObservability & monitoringCI/CD & DevOpsRubyPythonPostgreSQLRedisKafkaElasticsearch

Activity & contributions

Issues, merge requests, commits, and code review across internal GitLab and private GitHub.

17,467 contributions in the last year