Platform & Infrastructure
Designing cloud foundations, deployment workflows, and internal platforms that help product teams move safely.
30K+ Terraform-managed resources across cloud environments.
Engineer and architect with 11 years building cloud infrastructure, platform tooling, data systems, and backend services — including pipelines that stream 35M+ events/hour. My path has taken me from business operations to engineering leadership, with the same thread throughout: turning complex problems into working systems.
I own ambiguous problems end to end. Understanding the need, shaping the system, building the software, and making it reliable enough for others to depend on.
Designing cloud foundations, deployment workflows, and internal platforms that help product teams move safely.
30K+ Terraform-managed resources across cloud environments.
Building services, libraries, APIs, and integration layers for product and platform systems.
Python library ecosystem for secure clients, resource management, and workflow automation.
Connecting data pipelines, ML workflows, and model-serving infrastructure to practical product outcomes.
100+ technical users enabled on ML platforms and 20+ ML models deployed.
Owning the operational systems that keep production software observable, recoverable, and audit-ready.
35M+ events/hour through telemetry pipelines across accounts, regions, and clusters.
Translating business goals into technical direction, mentoring engineers, and aligning stakeholders around durable systems.
Led architecture across startup infrastructure and global Cisco data-science teams.
A few systems I’ve designed, built, or led across cloud infrastructure, platform engineering,data & ML systems, and backend software — focused on reliability, scale, and practical leverage.
Designed and built a CloudEvents-based messaging system on AWS to connect AI/ML inference services, backend workflows, and product-facing applications. Built Python libraries and middleware to standardize message construction, routing, and service integration across a rapidly evolving system.
Built centralized observability pipelines across AWS accounts, regions, and Kubernetes clusters to stream security, infrastructure, and application telemetry into a unified platform. Designed the system for high-volume ingestion, resiliency, operational visibility, and real-time alerting, supporting 35M+ events/hour.
Built a Python library ecosystem to deliver and support mission-critical products, including frameworks for secure HTTP/REST API clients, data masking and troubleshooting utilities, declarative resource management, API integrations, data exports, flexible configuration, credential management, and Kubernetes CRD-based workflow automation.
Designed and developed Terraform GitOps workflows and an AWS landing zone with account vending, security baselines, and standardized infrastructure patterns. Scaled the platform to manage 30,000+ resources across cloud environments.
Migrated 270+ Terraform workspaces between IaC GitOps platforms, reducing platform spend by a projected $120K annually while preserving deployment continuity across production infrastructure.
Owned technical security and compliance controls for SOC 2, including infrastructure hardening, access controls, observability, backup and recovery processes, and operational evidence collection.
Architected a GCP-based ML development platform and MLOps workflow for global data science teams deploying production services and pipelines.
Migrated Elasticsearch from a managed internal service to GKE, preserving critical search and analytics workflows while reducing cost.
Built an on-prem Jupyter and PySpark development platform on Hadoop, adopted by data scientists and engineers for production ML workflows.
Built a Hadoop job management CLI that improved developer workflows and scheduling reliability across data science, engineering, and analytics teams.
Self-directed product and systems work outside my primary role.
A Terraform/OpenTofu provider for consistent, convention-driven resource labeling — published to both the Terraform and OpenTofu registries. Go, provider-framework internals, and release automation.
Building a full-stack audit event platform for capturing, querying, and replaying compliance-grade product events. Focused on API design, data modeling, immutable storage, event mutation workflows, and developer-facing product ergonomics.
A compact career history, current work first. Expand any role for the scope, context, and decisions behind the systems I build today.
My instinct is to find where my capabilities and experience can have the biggest impact.
I started in business operations, automating spreadsheet-heavy workflows because manual processes were slow and holding the entire team back. I picked up data science and machine learning through coursework in 2014, then spent the next decade filling the infrastructure and platform gaps for technical teams: ML platforms, cloud infrastructure, data systems, and tooling that made other engineers faster.
Most of what I know is self-taught and proven in production. The titles changed — engineer, lead, director — but the purpose stayed the same: find the constraint, build the system, and make the work more valuable. Starting on the business side means I tend to ask what a system is for before diving into how I'll build it.
Every project carries its own risk tolerance and appetite for using AI. While AI has changed how fast I can work, it has not changed what I'm responsible for. I use it for leverage on the parts I already understand, and I don't ship anything I can't read, reason about, and defend.