Hamza Hassan

Machine Learning Engineer

Senior Machine Learning Engineer with 6 years focused on LLM tuning, evaluation and production deployment, grounded in end-to-end AI product delivery. Built and operated LLM inference and tuning pipelines for a customer-facing product and an RL evaluation engine, emphasizing PyTorch-based training, quantization and model serving. Experienced with prompt tuning, RAG and automation across distributed teams, and maintaining open-source infrastructure for reproducible deployments.

Experience

Aug 2025 — Present

Founder

ResumizeAI

  • Built and shipped an end-to-end LLM pipeline for resume tailoring and cover letter generation, using prompt engineering and RAG to improve relevance for 12,000+ users.
  • Designed model inference stack and production hosting on AWS, deploying quantized models and inference endpoints to reduce latency for user-facing features.
  • Implemented data collection and cleaning pipelines for prompt evaluation, curating training examples and human feedback loops to improve output quality.
  • Applied Prompt Engineering Fine-Tuning (PEFT) workflows with iterative evaluation, documenting prompts and metrics to guide model updates and product iterations.
  • Operated CI/CD for model updates and site deployment with Pulumi and GitHub Actions, ensuring reproducible releases and rollback for production LLMs.

Nov 2024 — Jul 2026

Senior Software Engineer, AI Systems

Mercor

  • Coordinated delivery of 35+ Model Context Protocol servers, integrating LLMs into production workflows and enabling standardized model context handling across teams.
  • Calibrated evaluation rubrics and reinforcement learning environments used for policy tuning, supporting PPO and Direct Parameter Optimization (DPO) style experiments.
  • Automated 90% of evaluation-to-deployment steps for model experiments, including dataset generation, metric calculation and MCP trajectory creation to speed iteration.
  • Trained RL models on 170+ open-source repositories, applying Transformer architectures and PyTorch training pipelines to improve code reasoning and multi-language completion.
  • Built multi-agent experiment orchestration with n8n integrations to reproduce evaluation flows and persist experiment state for model comparison.

Apr 2022 — Oct 2024

Full Stack Engineer

Morningstar Ventures

  • Decomposed a monolith into NestJS microservices and Serverless functions, creating infrastructure that supported future model-serving endpoints and scalable inference.
  • Designed event-driven pipelines with AWS EventBridge and SQS to stream features and labels, enabling reliable data ingestion for ML dataset preparation.
  • Optimized backend data pipelines and PostgreSQL queries to support low-latency retrieval, improving throughput for downstream model inference workloads.
  • Implemented GitHub Actions CI/CD, Docker and Kubernetes on EKS to standardize deployments and create reproducible environments for model packaging and serving.
  • Integrated Redis caching and Elasticsearch to speed retrieval for semantic search prototypes aligned with RAG experiments.

Aug 2021 — Mar 2022

Full Stack Engineer

Pairing

  • Built a Node.js backend with PostgreSQL and RabbitMQ for background processing, enabling reliable asynchronous jobs for model preprocessing and batch inference.
  • Implemented CI/CD and code quality tooling with GitLab, improving deployment reliability for services that feed ML pipelines.
  • Delivered end-to-end features with React and Ruby on Rails that integrated with backend services used to collect labeled data and user feedback for model improvement.

Jun 2019 — Jul 2021

Full Stack Developer

Pursue Today

  • Integrated Elasticsearch and Redis caching to improve search latency, supporting semantic retrieval experiments for downstream RAG workflows.
  • Built a GraphQL platform on Node.js that handled high-volume product data, providing structured inputs for feature extraction and dataset creation.
  • Implemented scalable data pipelines and storage on MSSQL to support large ingestion jobs used in dataset curation for ML experiments.

Projects

Jan 2026 — Present

Infra Foundry

Open Source Maintainer

A platform-agnostic cloud infrastructure components library for modern applications. Built with TypeScript and Pulumi, Infra Foundry provides reusable, composable infrastructure components that work across AWS, Cloudflare, and Vercel.

  • IaC
  • DevOps
  • Pulumi
  • AWS
  • GCP

Apr 2026 — Jul 2026

AI-Evaluation GTM Pipeline Engine

Lead Engineer & Architect

Built and engineered an end-to-end, 10-stage GTM engine for an AI-evaluation startup using TypeScript and GitHub Actions. Designed as an idempotent, Notion-backed state machine, the pipeline processes YC prospects via a provider-agnostic Claude client (Anthropic and AWS Bedrock via secure OIDC) with automatic model fallback and an automated, fail-closed IMAP reply-sweeper. To drive growth, I transformed pipeline artifacts into AI-generated evaluation rubrics that double as cold-outreach lead magnets, turning backend infrastructure into a high-leverage product-led growth loop.

  • TypeScript
  • Claude API
  • AWS Bedrock
  • GitHub Actions
  • Notion API
  • Resend
  • OIDC
  • n8n

Education

Jul 2016 — Jul 2020

Lahore University of Management Sciences (LUMS)

Bachelor of Computer Science · Computer Science

Skills

  • Machine Learning
  • Natural Language Processing (NLP)
  • Large Language Models (LLMs)
  • Sparse Fine-Tuning (SFT)
  • Prompt Engineering Fine-Tuning (PEFT)
  • Direct Parameter Optimization (DPO)
  • Proximal Policy Optimization (PPO)
  • Retriever-Augmented Generation (RAG)
  • Data collection and cleaning for ML (healthcare datasets)
  • Model conversion for production
  • Axolotl
  • vLLM
  • TGI (Text Generation Inference)
  • llama-cpp
  • Model quantization techniques and frameworks
  • Transformer architectures
  • PyTorch
  • Production-grade deep learning engineering
  • TypeScript
  • Python
  • AWS
  • Pulumi
  • GitHub Actions
  • n8n
  • Redis
  • Elasticsearch