Saad Anwar

Full-Stack ML Engineer specializing in production AI, Voice Systems, and MLOps.

Based in: Lahore, PakistanCurrently: Software Engineer, Xeven SolutionsExperience: 4 years

Hire me

Product-minded engineer with 4 years of experience building end-to-end AI systems, from fine-tuning YOLO models to deploying real-time ASR-to-TTS applications. He bridges the gap between complex ML infrastructure and user-facing SaaS products using Python, React, and FastAPI.

Experience

Software Engineer

Xeven SolutionsLahore, PakistanworkDec 2025 – Present

Summary

Own production features across FastAPI services, PostgreSQL data flows, secure API contracts, and React Native integration for a regulated digital-health product. Partner with frontend and product teams to scope workflows, ship iteratively, and support performance and reliability after release.

What I did

  • Owned features from API design and implementation through frontend integration, performance testing, and production support.

Results

  • Architected asynchronous FastAPI services for a HIPAA-compliant healthcare platform.
  • Built patient-management workflows that search and filter over 30,000 records while maintaining sub-200ms responses.
  • Reduced appointment-scheduling API response times by approximately 35% through PostgreSQL query optimization.
FastAPIPostgreSQLReact NativePython

Software Engineer

Evolve Innovative SolutionsLahore, Pakistanwork2024 – Nov 2025

Summary

Delivered AI products across LLM/RAG, computer vision, recommendation serving, real-time voice, and user-facing SaaS workflows. Built with React, Django, and FastAPI for platforms supporting 5,000+ concurrent users while maintaining sub-300ms API response times. Owned services from architecture and implementation through containerization, deployment, production debugging, and iterative optimization.

Results

  • Maintained sub-300ms API response times for platforms supporting 5,000+ concurrent users.
LLMRAGComputer VisionReactDjangoFastAPIDocker

Backend Engineer

CheetayLahore, PakistanworkJul 2022 – Dec 2023

Summary

Built and operated high-volume Django APIs, search pipelines, background jobs, and internal web tools for a production commerce platform. Worked end-to-end across API versioning, PostgreSQL-backed workflows, Elasticsearch, external integrations, and operational troubleshooting.

What I did

  • Built Celery pipelines for daily product record synchronization.

Results

  • Redesigned Elasticsearch indexing and search APIs, reducing search latency from roughly 800 ms to under 100 ms.
  • Reduced search latency by approximately 87%.
  • Built Celery pipelines to synchronize more than 70,000 product records daily.
DjangoPostgreSQLElasticsearchPythonCelery

Skills

Technical

PythonPython
YOLO
Model Inference
LLM Integration
Dataset Preparation
ReactReact
RAG
FastAPIFastAPI
Django / DRFDjango / DRF
REST APIs
DockerDocker
PostgreSQLPostgreSQL
Model Fine-Tuning
Object Detection
JavaScriptJavaScript
Bounding Boxes
Feature Extraction
NVIDIA Triton Server
NVIDIA ASR
NVIDIA TTS
Audio2Face
LangChain
Prompt Engineering
Embeddings
Vector Databases
Semantic Search
React NativeReact Native
Async IO
Elasticsearch
Celery
AWS (ECS, S3)
LinuxLinux
Git / CI/CDGit / CI/CD
RedisRedis
SSE
NGINXNGINX
Gunicorn
WebSockets
Dataset Annotation

Projects

MetaMe - Self-Service AI Character & Knowledge Platform

Full-Stack ML Engineerprofessional
Built a modular AI platform supporting Agent, Fictional, and Digital Twin characters with configurable memory, behavior, and response logic. Developed the full knowledge workflow: document ingestion, chunking, embeddings, vector storage, semantic retrieval, and grounded generation. Exposed asynchronous FastAPI services for product integration and containerized the platform with Docker for repeatable deployment. Optimized inference and Celery/Redis background processing, reducing average AI pipeline execution time by ~30%. Translated complex retrieval and generation components into configurable product workflows usable without direct ML infrastructure access.
ReactPythonFastAPIRAGEmbeddingsVector SearchDockerCeleryRedis

Recommendation System - Production Model Serving

Full-Stack ML Engineerprofessional
Integrated NVIDIA Triton Server to serve recommendation models behind a stable production inference boundary. Designed REST APIs that delivered personalized recommendations to product clients in real time. Combined multiple data sources into consistent model inputs and reliable inference results. Maintained separation between model serving and downstream consumers so model updates could evolve without breaking clients.
PythonNVIDIA Triton Inference ServerREST APIs

LivePortrait - Real-Time Voice-to-Avatar AI System

Full-Stack ML Engineerprofessional
Built a real-time conversational pipeline that transformed speech into text, generated an LLM response, synthesized speech, and animated a digital character. Integrated NVIDIA Omniverse ASR and TTS, then streamed generated audio to Audio2Face for blendshape extraction. Applied blendshapes to USD-based 3D characters to synchronize facial movement and spoken output. Debugged the system stage by stage across speech recognition, LLM generation, speech synthesis, and animation to improve real-time reliability.
PythonNVIDIA Omniverse ASRNVIDIA Omniverse TTSLLMsAudio2FaceUSD 3DLLMUSD character

Evolve-Search - Fine-Tuned Visual Search Platform

Full-Stack ML Engineerprofessional
Prepared training datasets and fine-tuned YOLO models for dress-type classification, creating practical ML workflows beyond hosted model integration. Produced bounding-box detections and extracted visual features for image-based product discovery. Indexed model-derived features in Elasticsearch to support image, text, and analytical search paths. Connected model outputs to a usable search product and collaborated with product and data partners on classification quality.
PythonYOLODataset PreparationFine-TuningObject DetectionElasticsearch

Education

Bachelor of Science

GC UniversityComputer Science2018 – 2022
Hire SaadReach out about a role, a contract or a conversation.For recruiters

Saad's twin is AI, it can make mistakes.