Muhammad Farhan Aslam

Muhammad Farhan Aslam

Senior Machine Learning Engineer specializing in Voice AI and Multimodal RAG Systems.

Based in: Lahore, PakistanCurrently: Senior Machine Learning Engineer, GTechSourcesExperience: 5 years

Hire me

Senior Machine Learning Engineer with over 5 years of experience building production-grade AI applications, including enterprise voice receptionists and multimodal retrieval systems for the US veterinary market. Expert in Python backend architecture, asynchronous services, and deploying containerized ML workflows across AWS and Azure.

Experience

Senior Machine Learning Engineer

GTechSourcesLahoreworkDec 2025 – Present

Summary

Enterprise Voice AI Receptionist for Veterinary Practices Enterprise Multimodal RAG Chatbot for Product Discovery and Customer Support Voice AI Sales Agent for Veterinary Instrument Ordering Veterinary AI Scribe and Clinical Documentation Platform

What I did

  • Built and deployed a multi-tenant voice AI receptionist for a US-based veterinary company using OpenAI Realtime, LiveKit Agents, Twilio SIP and FastAPI to handle inbound calls, new-client intake, bookings, rescheduling and cancellations.
  • On connection, mapped the dialed number to one clinic and location, then loaded its timezone, greeting, agent voice/instructions, enabled state, recording disclosure and after-hours rules. Rejected unknown numbers rather than falling back to another tenant.
  • Resolved the caller's number only within that clinic to retrieve the client, pets and upcoming appointments. Handled shared numbers and unmatched callers through pet-based verification and resumable registration; flagged uncertain pet details for staff review.
  • Designed backend-controlled conversation workflows so booking authority remained in application logic. Enforced clinic permissions, notice windows, identity checks and final approval with state-scoped tools; invalidated stale decisions when callers corrected earlier answers.
  • Matched visit needs to verified staff profiles, then offered slots only where the full visit fit clinic opening hours, a veterinarian's roster, time off and capacity. Rechecked these constraints at write time and scheduled in the location's timezone.
  • Supported in-clinic and home visits; used Google Routes to verify home-visit coverage and preserved existing visit duration and address when rescheduling unless the caller requested a valid change.
  • Answered clinic-specific FAQs with OpenAI embeddings and Pinecone tenant namespaces, grounding responses in the resolved clinic's knowledge without allowing retrieval or conversation state to cross tenants.
  • Protected client, pet and appointment writes with MongoDB validation, clinic/owner checks, atomic transactions and idempotent receipts. Reconciled uncertain outcomes and invalidated stale lookups before retries to prevent duplicate or wrong-client changes.
  • Triggered SMS/email confirmations and 30-minute reminders from committed appointment events, with consent, recipient and deduplication checks; suppressed stale notifications when visits were moved or cancelled.
  • Handled interruptions, silence and keypad staff transfers; applied emergency escalation and local-time after-hours booking, forwarding or voicemail. Persisted call outcomes and recording-linked summaries for clinic review.
  • Packaged API and LiveKit workers as isolated Docker services with health checks and versioned images; maintained rollback paths to make production releases recoverable.
  • Built a multimodal RAG chatbot for a US-based veterinary instrument company so customers could identify specialized products from descriptions or photos, inspect variants and product media, and ask business questions through one interface.
  • Exposed JSON and image-upload workflows through a JWT-protected Flask/Flask-RESTful API; used dependency injection to isolate authentication and validation from retrieval services, supporting focused testing and independent changes.
  • Used OpenAI and LangChain to classify each text request as product discovery or FAQ support, then route it to the appropriate knowledge source instead of searching unrelated content.
  • Combined SentenceTransformers and Redis vector search with lexical checks to retrieve products even when customer wording differed from catalog terminology. Ranked exact names, URLs, and variants ahead of similar instruments to resolve specific product requests.
  • Encoded uploaded photos with CLIP and searched Redis image vectors so customers could identify instruments they could not name. Similarity thresholds and question-aware selection returned a single best match or related products when requested.
  • Used Pinecone, OpenAI embeddings, and LangChain retrieval for business FAQs. Rewrote context-dependent follow-up questions and selected relevant, varied passages to answer from stored knowledge.
  • Validated generated responses against retrieved catalog records before returning product cards, variant relationships, images, and links. Replaced unsupported multi-product answers for exact searches with the verified match or clear no-match guidance.
  • Stored conversation history in MongoDB and short-lived search and pagination state in Redis, keyed by user and session. Reused saved results for next-page requests to preserve browsing context and avoid repeated retrieval or model calls.
  • Scheduled XML catalog synchronization with APScheduler; content hashes limited new embeddings to changed products, and stale records were removed from Redis to keep discovery aligned with the current catalog.
  • Built regression coverage around retrieval failure modes: session isolation, new-topic resets, exact and variant selection, empty results, mobile-photo orientation and import fallback, protecting behavior during catalog and application changes.
  • Built a voice-based purchasing assistant for a US veterinary instrument business, enabling callers to discover products, discuss catalog pricing and submit multi-item orders within one phone conversation.
  • Connected Twilio Media Streams to OpenAI Realtime through FastAPI WebSockets for bidirectional speech; handled voice activity and caller interruptions to support natural product-selection conversations.
  • Combined SentenceTransformers and Redis vector search with lexical and SKU-aware variant ranking to resolve product names, sizes and specifications; guided broad requests through categories and related catalog options.
  • Connected spoken purchase intent to backend order creation through function calling; enriched items with catalog SKUs and prices and validated quantities, email and shipping address before accepting an order.
  • Persisted pending orders and customer delivery details in MongoDB using asynchronous Motor operations; calculated order totals and generated itemized PDF receipts with ReportLab, providing a durable record for fulfillment.
  • Implemented an SMTP receipt-email service and order-ID status lookup, supporting confirmation delivery integration and post-order customer enquiries.
  • Built a veterinary documentation backend that converts recorded dictation into editable SOAP notes, free-form summaries and species-specific dental charts, helping clinicians structure patient records without manually reformatting each consultation.
  • Structured the backend into FastAPI gateway, authentication and clinical services to separate account/billing concerns from AI processing; routed HTTP requests and bidirectional WebSocket traffic through a common entry point.
  • Integrated Gemini and OpenAI for transcription and template-based note generation; added provider fallback for service failures and parallel processing for long transcripts. Selective SOAP-section updates preserved untouched clinical content.
  • Processed audio conversion, storage and transcription in background tasks with WebSocket progress notifications; retained multi-clip transcripts and note status in MongoDB and stored media in S3-compatible object storage.
  • Implemented registration, email verification, login and refresh-token workflows, using Redis for verification state; connected Stripe subscription lifecycle webhooks to plan records and voice-note usage limits.
  • Delivered reusable note templates, patient-linked records and editable generated outputs, allowing veterinarians to review and revise documentation within the same workflow.
  • Managed voice activity detection and background noise handling for real-time agents.

Results

  • Built a voice AI Receptionist for real-time conversations to book, cancel, or reschedule appointments.
  • Handled 100+ customers initially
  • Achieved latency of ~450ms - 650ms
OpenAI RealtimeLiveKit AgentsTwilio SIPFastAPIGoogle RoutesOpenAI embeddingsPineconeMongoDBDockerFlaskFlask-RESTfulJWTLangChainSentenceTransformersRedisCLIPAPSchedulerWebSocketsMotorReportLabGeminiS3Stripe

Backend Developer

Emerging Tech GridLahoreworkAug 2021 – Dec 2025

Summary

Decentralized Text-to-Music Generation Platform

What I did

  • Built and deployed a FastAPI backend for decentralized music generation, connecting user prompts to a blockchain-based network and returning generated audio through application APIs.
  • In the subnet, validators distribute prompts to independent miners; miners generate music, and validators score audio quality and prompt alignment to update on-chain weights and rewards.
  • Implemented authenticated asynchronous generation endpoints to support long-running inference and artifact delivery; persisted generated-file references and network data in PostgreSQL and MongoDB.
  • Used Kubernetes for autoscaling and orchestration of services
  • Implemented Kafka for real-time events tracking and handling

Results

  • Owned the full product lifecycle of the TTM platform from initial build through to public release.
  • Managed a network with almost 128 subnets, each containing hundreds of miners
  • Served thousands of requests gracefully across the decentralized network
  • Produced approximately 10,000 total text-to-music generations
  • Supported audio generation durations of 15s, 30s, 45s, and 60s
  • Processed thousands of events through Kafka during high-usage periods
  • Reached a target volume of thousands of daily audio generations.
FastAPIPostgreSQLMongoDBBlockchainKubernetesKafka

Skills

Technical

PythonPython
Docker ComposeDocker Compose
OpenAI SDK
RedisRedis
LiveKit Agents
OpenAI Realtime API
PyTorchPyTorch
MongoDBMongoDB
FastAPIFastAPI
TensorFlowTensorFlow
Pydantic
NumPy
REST APIs
DockerDocker
Pandas
AWS (EC2, S3, ECR, EKS, RDS, Lambda, SageMaker, Bedrock)
Pinecone
Hugging Face Transformers
LangChain
GitGit
Scikit-learn
TorchVision
SciPy
XGBoost
LightGBM
Keras
OpenCVOpenCV
YOLO
LlamaIndex
LangGraph
CrewAI
Twilio
FlaskFlask
Uvicorn
PostgreSQLPostgreSQL
MySQLMySQL
ChromaDB
FAISS
Celery
MLflow
Airflow
GitHub ActionsGitHub Actions
NginxNginx
Google CloudGoogle Cloud
Microsoft Azure (Azure Container Apps)
KubernetesKubernetes
Apache Spark
DVC
OpenTelemetry
Kubeflow
KServe
CUDA
Albumentations
Kafka
RabbitMQ
Prometheus
Grafana
Helm
TerraformTerraform
Argo CD
gRPC
JavaScriptJavaScript
Agent2Agent (A2A)
MCP (Model Context Protocol)
Elasticsearch
CSSCSS
HTMLHTML
CLIP
ViT
MLP

Projects

Misinformation Detection Research Project

personal
A private research project conducted for a KSA client focused on misinformation detection.
ViTCLIPMLP

Education

Bachelor's in Political Science

University of the Punjab, LahorePolitical Science2023

Recognition

Certifications

Artificial Intelligence Development

PIAIC

Professional Data Science

DataCamp

Backend Development

Udemy

Patents & publications

Multimodal Detection of Misleading Image-Caption Pairs

research paper

Investigated out-of-context misinformation using NewsCLIPpings benchmark and VisualNews source media (71,072 labeled pairs). Extracted VGG16, ViT, and CLIP visual embeddings. Encoded captions with CLIP and Sentence-BERT. Trained MLP classifiers and compared PCA-reduced SVM baselines. Regenerated images with Stable Diffusion and DALL-E 3 to probe visual context mismatch. Achieved 80% accuracy on 7,264 test pairs.

Hire MuhammadReach out about a role, a contract or a conversation.For recruiters

Muhammad's twin is AI, it can make mistakes.