Sarmad Khattak

Sarmad Khattak

AI/ML Engineer specializing in Agentic AI, Scalable RAG, and Computer Vision.

Based in: Lahore, Punjab, PakistanCurrently: AI/ML Engineer, muSharpExperience: 5 years

Hire me

Sarmad is an experienced AI/ML Engineer with a proven track record in architecting multi-agent systems, real-time voice agents, and advanced computer vision pipelines. He has authored research on hyperparameter optimization and excels at deploying high-performance AI microservices using frameworks like LangChain, TensorFlow, and PyTorch.

Experience

AI/ML Engineer

muSharpLahore, Pakistanwork2025 – Present

Summary

Agentic AI & Scalable RAG Architectures. Advanced Computer Vision & VLM Intelligence. Conversational AI & Voice Technology. Document Intelligence & Security. MLOps & Core Engineering.

What I did

  • Architected complex multi-agent systems using LangChain and LangGraph, focusing on agent orchestration, inter-agent communication, and scalability.
  • Built autonomous agent execution with safety rails: plan-then-execute flows, per-action
  • Implemented robust RAG pipelines featuring advanced document parsing, vector retrieval, and guardrailing to ensure response accuracy and safety.
  • Developed comprehensive audit logging mechanisms to monitor agent decision-making processes and ensure system transparency.
  • Implemented feedback and evaluation loops: tracks user acceptance/rejection/modification of AI.
  • Engineered real-time voice agents (Pipecat + WebRTC): STT → LLM → TTS with VAD and barge-in, low-latency turn handling, and Coturn/TURN credential minting (time-limited HMAC REST credentials) for reliable NAT traversal.
  • Leveraged and fine-tuned Vision Language Models (VLMs) for high-level activity detection, automated alerting, and complex scene understanding.
  • Engineered Person Re-Identification (ReID) systems with extensive gallery management and optimized algorithms for multi-camera tracking.
  • Implemented 3D landscape reconstruction and remote sensing pipelines for geospatial analysis and environmental monitoring.
  • Maintained real-time inference infrastructure using NVIDIA DeepStream and GStreamer, optimizing pipelines for state-of-the-art object detection and recognition.
  • Built end-to-end voice agents for automated interviewing and customer support, integrating Speech-to-Text (STT), Voice Activity Detection (VAD), and LLM-driven logic.
  • Designed low-latency conversational workflows that handle interruptions and maintain context for natural human-AI interaction.
  • Developed Intelligent Document Processing (IDP) solutions capable of parsing unstructured data and performing high-accuracy OCR on diverse international formats.
  • Engineered automated PII redaction systems ensuring data privacy compliance across sensitive documents.
  • Integrated Generative AI for synthetic data generation to augment training sets and deployed Virtual Try-On (VTON) applications for enhanced user engagement.
  • Established CI/CD pipelines using GitHub Actions, automating testing and deployment workflows across both local servers and cloud environments.
  • Optimized model deployment strategies to ensure high availability and scalability of AI microservices
  • Architected multi-agent AI systems, self-learning feedback loops, and agentic knowledge graphs connected to enterprise ERPs for yally.ai.
  • Deployed real-time facial recognition pipelines across all international airports for IBS.
  • Built automated real-time conversational voice agents using WebRTC and Pipecat for the Prime Minister Helpline.
  • Engineered agentic document extraction and OCR pipelines for Ocuco.
  • Deployed secure GenAI solutions and branch-level AI automation for Ubank and HBL.
  • Developed the Flare platform for PTCL.

Results

  • Used Arize Phoenix and Langsmith for agent observability.
  • Achieved end-to-end conversational latency of under 500 milliseconds for real-time voice agents.
  • Maintained sub-second query responses for high-volume enterprise document parsing and vector retrieval.
LangChainLangGraphRAGPipecatWebRTCVADCoturnTURNHMACRESTVLMReIDNVIDIA DeepStreamGStreamerSTTTTSLLMOCRIDPPII RedactionGenerative AIVTONGitHub ActionsCI/CDArize PhoenixLangsmith

Associate AI/ML Engineer

Swati TechnologiesLahore, PakistanworkNov 2023 – Dec 2024

Summary

Key Responsibilities: • Designed and deployed advanced machine learning models and algorithms to solve complex problems, with expertise in computer vision, natural language processing (NLP), and multimodal AI solutions. • Developed cutting-edge computer vision applications, including real-time face detection, pedestrian detection for autonomous systems, vehicle detection in parking lots, OCR-based document processing with Tesseract, iris recognition, and voice and facial recognition systems. • Fine-tuned large language models (LLMs) for building intelligent chatbots, including RAG-based Q&A systems and personalized virtual assistants, integrating text and vision-based knowledge to handle diverse user queries. • Pioneered multimodal AI systems, combining vision and language understanding to create innovative solutions for complex tasks in real-world applications. • Built intelligent agents for tasks like sales analysis, business strategy, and content generation, leveraging data-driven insights and automation tools to optimize workflows and improve decision-making. • Integrated automation workflows using APIs and tools to streamline operations, such as real estate calling agents and technical support systems, while minimizing overhead. • Proficient in TensorFlow, PyTorch, Hugging Face, and LangChain, with practical experience in fine-tuning and deploying AI models on both cloud platforms and local CPU environments.

Results

  • Achieved over 95% mAP for pedestrian and vehicle detection models in real-world lighting conditions.
  • Maintained real-time inference speeds above 30 FPS using optimized OpenCV and TensorRT pipelines.
  • Handled multiple concurrent high-definition RTSP camera streams per hardware node.
TensorFlowPyTorchHugging FaceLangChainNLPComputer VisionMultimodal AITesseractOCRLLMRAGOpenCVTensorRTRTSP

Outreach Manager

Avija DigitalRemotework2023 – Sep 2023
Key Responsibilities: • Created and implemented social media strategies to drive engagement and brand growth. • Analyzed metrics to adjust campaigns, improving reach and interaction. Engaged with audiences and worked with design teams to ensure consistent branding.

Skills

Technical

Machine Learning & AI
PythonPython
Scalable RAG Pipelines
Multi Agent Systems
PyTorchPyTorch
NLP
Computer Vision
LangChain
LangGraph
TensorFlowTensorFlow
GitHub ActionsGitHub Actions
NVIDIA DeepStream
FastAPIFastAPI
Voice/Conversation AI
Video Analytics
OCR Pipelines
MLOps
3D Reconstruction
Data Analysis & SQLData Analysis & SQL
Cloud Computing
AI Governance & Compliance
GStreamer
Arize Phoenix
Langsmith
Cybersecurity
Pipecat
WebRTC
OpenCVOpenCV
TensorRT
vLLM
Graph RAG
Agentic pipeline
Rapid prototyping
WSL2
DockerDocker

Languages

Pashto
English

Projects

Flare Platform (PTCL)

professional
Developed the Flare platform for PTCL.

GenAI Solutions (Ubank)

professional
Deployed secure GenAI solutions for Ubank.
Generative AILLM

Branch-level AI Automation (HBL)

professional
Deployed AI solutions and branch-level automation for HBL.

Voice Recognition System (KONE)

professional
Developed voice recognition solutions for KONE.
Voice Recognition

Iris Recognition

Developerpersonal
A biometric security system using iris scanning technology to authenticate individuals. Known for high accuracy, it enhances security in sensitive applications.
Computer VisionBiometricsPython

Pedestrian Detection for Autonomous Vehicles

Developerpersonal
Utilizes machine vision techniques to detect pedestrians and assess their movements. This project enhances autonomous vehicle safety by reducing the likelihood of collisions.
Machine VisionObject DetectionAutonomous Systems

Real-Time Face Detection System

Developerpersonal
A facial recognition system designed to detect faces accurately in real-time. Optimized for applications needing immediate feedback and high accuracy in varying lighting conditions.
Facial RecognitionReal-time processingComputer Vision

Vehicle Distance Measurement

Developerpersonal
A computer vision solution that calculates the distance between vehicles in real-time. It aims to improve road safety by facilitating safe following distances.
Computer VisionDistance Estimation

Voice Cloner

Developerpersonal
A voice synthesis system capable of replicating a user's vocal characteristics. Useful for creating digital assistants, voiceovers, and personalized customer service.
Voice SynthesisDeep Learning

Voice Recognition

Developerpersonal
A speech-processing system that identifies and authenticates users based on voice patterns. Designed for applications requiring seamless, hands-free interaction.
Speech ProcessingAuthentication

Agentic AI Systems (Yally.ai)

Lead Developerprofessional
Architected multi-agent AI systems with self-learning feedback loops and agentic knowledge graphs connected to enterprise ERPs for yally.ai.
LangChainLangGraphMulti-agent systemsKnowledge Graphs

Medical Insurance Intelligence System

personal
Built a zero-to-one system independently, fine-tuning a Laya System 1 model on medical insurance datasets and integrating a Graph RAG architecture to map complex entity relationships.
Laya System 1vLLMGraph RAGAgentic pipeline

Local DeepStream Windows Environment

personal
Configured NVIDIA DeepStream to run on Windows using Docker with a WSL2 backend after extensive trial and error to enable local development on a laptop.
NVIDIA DeepStreamDockerWSL2Windows

Facial Recognition & MRZ Pipeline (IBS Airports)

professional
Deployed real-time facial recognition and universal passport MRZ scanner pipelines across all international airports for IBS.
NVIDIA DeepStreamGStreamerComputer Vision

Conversational Voice Agent (Prime Minister Helpline)

professional
Built automated real-time conversational voice agents using WebRTC and Pipecat for the Prime Minister Helpline.
PipecatWebRTCSTTTTSVAD

Agentic Document Extraction (Ocuco)

professional
Engineered agentic document extraction and OCR pipelines for Ocuco.
OCRIDPAgentic Pipelines

Education

BS Software Engineering

Virtual University of Pakistan (VU)Software Engineering

Certificate in Accounting and Finance

The Institute of Chartered Accountants of PakistanAccounting and Finance2020 – 2023

F.Sc

Army Burnhall College for Boys2017 – 2019

Recognition

Certifications

MLOps

Duke University

Deep Learning Specialization

Stanford Online

Machine Learning Specialization

Stanford Online

Python

Edx

Patents & publications

Firefly Algorithm for Hyperparameter Optimization

publication2024

FA demonstrate a capability for exploring the hyperparameter space more effectively due to its adaptive nature, leading to better model performance.

Hire SarmadReach out about a role, a contract or a conversation.For recruiters

Sarmad's twin is AI, it can make mistakes.