400M-Document Serverless Re-Architecture
Re-platformed a high-GPU SageMaker Elasticsearch seeding workload as a Serverless solution. 400M+ searchable documents seeded in under half an hour.
Fifteen-plus years turning customer pain points into production systems across AWS, GCP, and Azure, from multi-account landing zones to GenAI inference at scale. I build MLOps pipelines with experiment tracking, orchestrate workloads on Kubernetes (GKE and self-hosted HA clusters), and embed DevSecOps gates — SonarQube, Fortify, Trivy, Secret Scan — inline in CI. I optimize for the problem in front of me, not the stack I already know.
Architecture is the set of decisions that are expensive to reverse. I make them deliberately, working from PreSales pain points to feasible designs, with cost and reliability weighed up front.
Re-platformed a high-GPU SageMaker Elasticsearch seeding workload as a Serverless solution. 400M+ searchable documents seeded in under half an hour.
Led Architecture and Development of a Cloud Management Platform
Organization-scale account (AWS) and project (GCP) architecture as the governance foundation every workload and guardrail is built on.
Re-architected a Node.js + Python and React product so a single deployment serves many customers at a fraction of the cost, UI decoupled to S3 + CDN.
Led modernization from PreSales to implementation, decoupling UI, REST platform, partner APIs, and authentication piece by piece.
Partnered with sales teams and customers to translate constraints and feasibility into custom multi-cloud solutions.
Multi-cloud isn't a buzzword here, it's training on TPUs in one cloud and serving on GPUs in another, because that's what the workload and the bill demanded.
Multi-GPU Llama2-based serving on GKE with vLLM and FastAPI on A100 and L4 GPUs, with GPU autoscaling driven by KubeRay.
SDXL text-to-image inference on GKE with a JAX pipeline on Cloud TPU and L4 GPUs.
Fine-tuned Gemma 3 model inference on Vertex AI using LoRA adapters for domain-specific tasks.
Custom-dataset fine-tuning of large language models using LoRA and full fine-tuning approaches.
PyTorch / TensorFlow on SageMaker, EC2 deep-learning VMs, GCP Compute and TPU v3 with automated provisioning.
Training-to-inference pipelines across AWS and GCP, with ML experiment tracking on Kubernetes.
Served pickle models as Lambda REST APIs with on-demand model loading and IAM-aware Python applications.
A pipeline should make the secure, repeatable path the easy one. Security gates and infrastructure-as-code belong inside the build.
Declarative Jenkinsfile on Azure with Ansible, SonarQube, Nexus IQ, Fortify, WebInspect, Nexus, Docker registry, and Kubernetes.
kubeadm-bootstrapped HA cluster with external etcd, HPA autoscaling, and Prometheus / Grafana monitoring.
Naming and tagging factory, network factory, Go unit tests, module docs, and AWS / GCP boilerplates.
CloudWatch, SNS, SQS, Python Lambda, DynamoDB scheduled jobs.
Containerized services on ECS with Bitbucket-to-Jenkins-to-ECR CI/CD and automated deployment.
A polyglot toolbox — Go, Python, Node, Rails, PHP — means the language is a choice that fits the problem.
iOS / Android application with API on AWS, Jenkins backend deployments, and Fastlane mobile releases.
Led architecture and delivery of the FullBridge Ruby on Rails platform.
Partner-facing API on API Gateway and Lambda with CodePipeline CI/CD, lowering AWS spend.
Reading-list conversion and reporting platform with notification-driven delivery and customizable workers.
Personalised recommendation engine providing data-driven suggestions based on user profiles and activity.
Full-featured web auction platform with multi-format bidding, web stores, and payment gateway integration.
Practitioner notes for engineers, where the argument is settled by measurement, not opinion.
Bifrost crushed LiteLLM in AI Gateway benchmark. But that’s not the whole story.
Go delivers 7x better latency under high load with half the resources. Numbers, not opinions.
Multi-stage builds, .dockerignore hygiene, removing build-tool hangovers. Under 20 MB.
The answer isn't binary. Like the IDE before it, how well it works depends on how well you use it.
vLLM, GPU autoscaling, and getting real ROI from every dollar spent on A100s.
LoRA, QLoRA, quantization, hardware selection. Choosing the right fine-tuning approach.
The fastest way to reach me is email.