Engineering Efficient Data Pipelines
Data Engineer whose pipelines run on schedule, recover on their own, and raise a flag when they can't. I built a daily ELT pipeline (PostgreSQL, dbt, Airflow) that unifies 4 operational areas into 1 analytical warehouse for a metalworking manufacturer. Every run has 2x automatic retries, 30-minute task timeouts, and Telegram failure alerts, and the pipeline was stress-tested with a synthetic dataset containing ~2.5% deliberately dirty records. Behind it are 5+ years in data operations and governance: 50 institutions and 2,000+ users consolidated into one verifiable source of truth, and 50% less processing time after automated validation across 30 units. I have also modeled usage data from 20 countries in dbt and BigQuery (served via FastAPI) and built an AI-graded homework platform for a ~120-student school client using Gemini.
I spent the last five years in the messy middle of public-sector data: 50 institutions, 2,000+ users, and spreadsheets that never quite agreed with each other. Turning all of that into one source of truth people could actually trust is where I fell for data engineering.
Today I build the pipes behind the numbers. An AI-graded homework platform for a school of 120 students. An offline-first ledger for a fish-procurement business moving a ton of catch every week. And pipeline-lint, an open source tool that catches the mistakes AI-written pipelines make before they reach production.
My rule is simple: when a pipeline breaks at 2 a.m., it should either fix itself or tell someone. I treat data quality as a feature, not a cleanup job.
I'm looking for a Junior or Associate Data Engineer role, remote-friendly, working from UTC+7. If your team cares about data that holds up, let's talk.
REST APIs, automation scripts, and stateless workers fed by a Redis job queue.
Complex queries, data warehousing, and idempotent upserts that stay correct on rerun.
Transformation, orchestration, and reliable data modeling for scheduled pipelines.
Spark writes, Delta MERGE, and the rerun pitfalls my linter pipeline-lint checks for.
BigQuery, Cloud Storage, and containerized services that run the same anywhere.
Infrastructure as code, CI test runs, and automated releases to PyPI.
AI grading bound to a strict JSON schema, with a dead letter queue for failed answers.
Offline-first Android apps with local Hive storage, Excel export, and thermal receipts.
Data architecture and software systems that I have designed and built.
Watch overview · 1:13
An AI-integrated academic portal for students that automates grading and academic tracking. Built with FastAPI on an event-driven architecture.
Watch overview · 0:54
A fisheries supply-chain tracking application that displays net sales balance, inbound and outbound stock metrics, and activity history in real time.
Watch overview · 1:11
An open source linter for SQL, PySpark, Databricks and Airflow code. It catches non-idempotent writes, duplicate data on rerun, and other mistakes AI-generated pipelines make, and can fail CI before they reach production.
I am open to Junior/Associate Data Engineer opportunities. Feel free to reach out through any of the channels below.