Zulqarnain hiader

Zulqarnain hiader

Machine Learning Engineer & Generative AI Specialist

Based in: Lahore, PakistanMost recently: Machine Learning Engineer, CodeNinja Inc.Experience: 4 years

Hire me

Experienced Machine Learning Engineer specializing in architecting end-to-end AI platforms, multi-agent autonomous workflows, and production-grade RAG applications. Proven track record in optimizing model inference and building scalable MLOps pipelines across AWS and distributed computing environments.

Experience

Machine Learning Engineer

CodeNinja Inc.Lahore, PakistanworkDec 2024 – Apr 2026

Summary

Architected and shipped end-to-end machine learning platforms on AWS, streamlining data pipelines, model fine-tuning. Designed stateful, multi-agent autonomous workflows using LangGraph to automate complex enterprise reasoning tasks. Built production Retrieval-Augmented Generation (RAG) applications backed by PostgreSQL with pgvector and Amazon OpenSearch, incorporating dynamic re-ranking and prompt engineering to reduce hallucination rates by 35%. Integrated AWS data services (AWS Glue, Amazon Athena, Amazon Redshift) for scalable data processing, feature engineering, and automated ETL pipelines servicing ML endpoints. Fine-tuned pre-trained LLMs and domain-specific foundation models using PyTorch for enterprise clients. Optimized model inference using INT8 quantization, achieving a 42% reduction in deployment latency on AWS SageMaker.

What I did

  • Built React and TypeScript interfaces to allow enterprise clients to interact with machine learning models.
  • Owned the full stack from backend infrastructure to React and TypeScript interfaces for enterprise clients.

Results

  • Reduced hallucination rates by 35% through dynamic re-ranking and prompt engineering.
  • Achieved a 42% reduction in deployment latency on AWS SageMaker using INT8 quantization.
AWSLangGraphRAGPostgreSQLpgvectorAmazon OpenSearchAWS GlueAmazon AthenaAmazon RedshiftPyTorchINT8 quantizationAWS SageMakerReactTypeScriptPEFTLoRALlama 3 8B

Research Assistant

Parallel Computing NetworksIslamabad, PakistanworkNov 2023 – Nov 2024

Summary

Researched scalable deep learning architectures, implementing distributed training workflows across multi-GPU clusters using PyTorch Distributed and Ray on AWS infrastructure. Engineered distributed data ingestion pipelines utilizing Apache Spark and AWS Glue for feature extraction, chunking. Implemented hybrid search and similarity retrieval systems leveraging Amazon Kendra and vector databases for high-precision. Built reproducible MLOps workflows and scalable pipeline orchestration systems using Kubeflow and Apache Airflow to manage continuous model retraining, systematic evals, and reliable endpoint deployment. Documented technical architectures, conducted peer code reviews, and communicated complex AI strategies to both technical teams and business stakeholders.

Results

  • Reduced training runs from around 40 hours on a single node to 12 hours using distributed training infrastructure.
  • Scaled workloads from single-node multi-GPU baselines to multi-node AWS clusters.
PyTorch DistributedRayAWSApache SparkAWS GlueAmazon KendraVector DatabasesKubeflowApache AirflowMLOpsPyTorch DDP

Software Engineer (Rust)

ReplitSan Francisco, USA (Remote)workAug 2022 – Dec 2023

Summary

Engineered memory-efficient data processing modules in Rust for handling high-throughput streams. Applied performance profiling and optimization techniques to data-intensive applications, achieving a 30% reduction in compute overhead. Built data access layers with PostgreSQL, including indexing strategies and query tuning to support high-throughput. Architected resilient, asynchronous data pipelines and internal tooling to ensure high availability and fault tolerance. Designed and deployed highly concurrent microservices, facilitating seamless integration between low-level backend environments and TypeScript/React frontend clients.

Results

  • Achieved a 30% reduction in compute overhead through performance profiling and optimization.
RustPostgreSQLTypeScriptReactMicroservicesAsynchronous Pipelines

Skills

Technical

PythonPython
FastAPIFastAPI
LangGraph
Vector Embeddings
Transformers
RAG
Amazon SageMaker
PostgreSQL (pgvector)PostgreSQL (pgvector)
Prompt Engineering
SQLSQL
Vector Databases
Model Fine-Tuning
PyTorch DistributedPyTorch Distributed
AWS Glue
LangChain
PEFT/LoRA
Apache Spark
Data Warehousing
Keras
scikit-learn
AWS Bedrock
Amazon Athena
Amazon Redshift
Amazon OpenSearch
Amazon Kendra
Azure OpenAI
DockerDocker
Ray
Kubeflow
Apache Airflow
Quantization
RustRust
ONNX
Amazon Rekognition
CI/CD
OpenCVOpenCV
C++C++
Feature Stores
Cross-encoder
Computer Vision Pipeline
ReactReact
TypeScriptTypeScript
Independent project creation
Rapid prototyping
BM25

Projects

LLM Agentic Legal Information Retrieval System

personal
Built an end-to-end legal information retrieval system from scratch featuring a hybrid retrieval pipeline.
BM25Dense EmbeddingsCross-encoder

Solar Filament Segmentation

personal
Developed a complete computer vision pipeline for solar filament segmentation, including raw dataset pre-processing.
Computer Vision

Internal NLP Query Tool

personal
Built an internal tool that combines Microsoft Teams, Jira, and a local database to allow users to get results via NLP queries instead of checking multiple platforms.
JiraMicrosoft TeamsNLPLocal LLM

Education

Master of Science

University of Engineering and Technology (UET)Computer ScienceAug 2025

Bachelor of Science

FAST National University of Computer and Emerging SciencesComputer ScienceAug 2019 – Jun 2023
Hire ZulqarnainReach out about a role, a contract or a conversation.For recruiters

Zulqarnain's twin is AI, it can make mistakes.