Muhammad Humza

AI/ML Engineer specializing in LLM-powered voice agents and production RAG pipelines.

Based in: Lahore, PakistanCurrently: AI/ML Engineer, YAMSOLExperience: 3 years

Hire me

Experienced AI/ML Engineer with a proven track record of deploying real-time voice agents, scalable microservices, and predictive models that drive operational efficiency. He has successfully delivered high-impact solutions across NEMT dispatch, legal research, and retail forecasting industries using advanced LLM orchestration and RAG architectures.

Experience

AI/ML Engineer

YAMSOLworkJun 2025 – Present

Summary

Built and deployed a real-time AI voice agent using LLMs (Ollama) with dynamic tool orchestration (18+ tools) for a NEMT dispatch system, enabling trip booking, cancellation and trip queries, and driver assignment; handled 200+ daily requests with sub-second inference latency. Engineered a low-latency speech pipeline (Faster Whisper STT, TTS, VAD, WebSockets) and scalable microservices architecture (Docker, REST APIs), achieving 40%+ latency reduction and supporting real-time scheduling optimisation. Developed a full-stack dispatch and scheduling system (React, Python, Firebase) with real-time planning, filtering, and drag-and-drop reassignment; integrated 2 ML models (ETA + Trip-willcall-time prediction), improving dispatch accuracy by 15-20% and reducing manual scheduling effort by 60%. Built Twilio-integrated AI voice agents (inbound & outbound) using SLMs driven conversational decisioning to automate trip confirmation, cancellation, rescheduling, and real-time ETA/driver status inquiries in NEMT workflows, replacing manual dispatcher calls with end-to-end automation, including voicemail handling and SMS fallback

What I did

  • Integrated 2 ML models (ETA + Trip-willcall-time prediction) for dispatch accuracy
  • Implemented voicemail handling and SMS fallback for automated voice agents
  • Containerized services with Docker to ensure reproducible deployments.
  • Version-controlled LLM, STT, TTS, VAD, prompts, and tool definitions alongside the application code.
  • Monitored operational metrics including response latency, request volume, tool execution, and failures.
  • Logged conversations and errors to identify transcription failures, incorrect tool selection, and prompt regressions.
  • Helped build a voice pipeline using speech-to-text, an LLM agent, tool calling, text-to-speech, VAD, WebSockets, and backend services.

Results

  • Worked across the full lifecycle of a user-facing AI product, including deployment, monitoring, and iterating based on real user interactions and failures.
  • Achieved 40%+ latency reduction and supported real-time scheduling optimisation
  • Improved dispatch accuracy by 15-20%
  • Reduced manual scheduling effort by 60%
  • Replaced manual dispatcher calls with end-to-end automation
LLMsOllamaFaster Whisper STTTTSVADWebSocketsDockerREST APIsReactPythonFirebaseTwilioSLMsModel VersioningProduction MonitoringSpeech-to-TextTool calling

AI/ML Engineer

Vebtual LTDworkFeb 2024 – Apr 2025

Summary

Designed, built, and deployed 5+ production ML models across client projects using Python, TensorFlow, and Hugging Face, including a multi-store sales forecasting system (retail, 10+ locations) reducing overstock costs by ~18%, a medical record tagging model using fine-tuned BERT classifying 5,000+ patient records by condition and urgency with 89% accuracy, and a predictive maintenance model for a manufacturing client reducing unplanned downtime by 25%. Collaborated with the core product team to build and integrate a dual-modal emotion detection system for a real-time therapy chatbot, contributing to a CNN-based facial emotion recognition pipeline (FER+) on live video feed and a speech emotion recognition model extracting MFCCs, jitter, and shimmer features with a stacked ensemble (XGBoost + SVM), achieving 85% emotion classification accuracy and dynamically adapting chatbot tone at inference in real time. Engineered a production RAG chatbot for a US law firm, ingesting 500+ legal PDFs across multiple practice areas using hierarchical and semantic chunking to preserve clause-level structure, OpenAI embeddings stored in Pinecone, and hybrid retrieval (dense + BM25) for high-precision legal query responses, reducing manual legal research time by 40%.

What I did

  • Developed a multi-store sales forecasting system for retail with 10+ locations
  • Created a medical record tagging model using fine-tuned BERT for patient record classification
  • Built a predictive maintenance model for a manufacturing client
  • Contributed to a CNN-based facial emotion recognition pipeline (FER+) on live video feed
  • Developed a speech emotion recognition model extracting MFCCs, jitter, and shimmer features with a stacked ensemble (XGBoost + SVM)
  • Implemented hierarchical and semantic chunking, OpenAI embeddings, and hybrid retrieval (dense + BM25)
  • Owned the CNN-based facial emotion recognition pipeline, including image preprocessing and face-focused data pipeline development.
  • Managed training data preparation, normalization, and augmentation for the emotion classification model.
  • Trained, tuned, and evaluated the CNN model for performance across different emotion classes.
  • Developed the interface for integrating facial/visual modality predictions into the broader multimodal system.

Results

  • Reduced overstock costs by ~18%
  • Classified 5,000+ patient records with 89% accuracy
  • Reduced unplanned downtime by 25%
  • Achieved 85% emotion classification accuracy
  • Reduced manual legal research time by 40%
PythonTensorFlowHugging FaceBERTCNNFER+MFCCsXGBoostSVMRAGOpenAI embeddingsPineconeBM25

Skills

Technical

AI Agents
Voice Chatbots
RAG
LLMs
PythonPython
OpenAI SDK
NLP
LangGraph
Git/GitHubGit/GitHub
Machine Learning
Faster Whisper
Ollama
Semantic Search
Embeddings
Vector Databases
Text Chatbots
NumPy
Pandas
Scikit-learn
Deep Learning
LangChain
Hugging Face
Pinecone
FlaskFlask
Multimodal AI
Computer Vision
Predictive Modelling
Real-time Systems
Streaming Architectures
TensorFlowTensorFlow
PyTorchPyTorch
SQLSQL
PostgreSQLPostgreSQL
MySQLMySQL
FAISS
Weaviate
WebSockets
DockerDocker
FirebaseFirebase
Data Engineering
AWS
AWS Bedrock
AWS S3
NoSQLNoSQL
Statistical Analysis
YOLO
OpenCVOpenCV
MLops
NginxNginx
ETL/ELT
ReactReact
A/B Testing
Streamlit
Production Monitoring
Compliance guardrailsCompliance guardrails
Model Versioning
CNN
FastAPIFastAPI
Data Augmentation

Projects

AI-Powered Personal Finance Assistant

Lead Developerpersonal
Built an async 7-node LangGraph agent to handle transaction queries, budget insights, financial health scoring, and anomaly detection through natural conversation, using intent classification with LLM + heuristic fallback for reliable routing. Implemented a hybrid RAG pipeline (BM25 sparse search + Chroma vector search + Reciprocal Rank Fusion + cross-encoder reranking) over a 16-document US/UK finance knowledge base, with structured LLM output extraction (Pydantic) for reliable data parsing from bank APIs. Engineered a FastAPI microservice architecture with Redis session memory, SSE streaming, Prometheus observability, and a regex-based FCA compliance guardrail; validated system reliability with 59 passing regression tests.
LangGraphFastAPIRAGOllamaLLMBM25ChromaReciprocal Rank FusionCross-encoder rerankingPydanticRedisSSE streamingPrometheusRegexStreamlitSession memory

Education

Bachelor of Science

National University of Computer and Emerging SciencesBusiness AnalyticsGPA 3.08Sep 2021 – Jun 2025
Hire MuhammadReach out about a role, a contract or a conversation.For recruiters

Muhammad's twin is AI, it can make mistakes.