Munaza Ashraf

Munaza Ashraf

Machine Learning Engineer specializing in high-performance inference and multimodal AI systems.

Based in: PakistanCurrently: Machine Learning Engineer, RunwareExperience: 3 years

Hire me

A high-achieving Software Engineer with a 3.94 CGPA and extensive experience in optimizing production-grade ML pipelines for image, video, and LLM workloads. Proven track record in reducing latency and scaling AI infrastructure for enterprise-level applications across healthcare and industrial sectors.

Experience

Machine Learning Engineer

RunwareLondon, EnglandworkNov 2024 – Present

Summary

Design, integrate, optimize, evaluate, and deploy production image, video, audio, 3D, LLM/VLM, and multimodal AI models, achieving up to 2× inference improvement across selected workloads. Built scalable AI inference infrastructure supporting 1,000+ models, including model loading, GPU-memory management, batching, concurrency control, workload isolation, and GPU-aware scheduling. Train, benchmark, and optimize models using quantization, attention optimization, caching, compilation, distributed execution, and parallelism, evaluating latency, throughput, VRAM usage, and output-quality trade-offs. Profile production ML workloads to identify CUDA, memory, synchronization, and compute bottlenecks and translate results into model and system improvements. Built asynchronous production pipelines using RabbitMQ, Redis, Celery, Docker, Prometheus, Grafana, and Kibana for scalable execution, monitoring, and reliability. Supported high-volume enterprise AI workloads for customers including Higgsfield, Freepik, Wix, Vercel, PixVerse, and Mayflower.

What I did

  • Built production code for the AI inference platform.

Results

  • Built a self-service LoRA training platform, managing the full flow from training and GPU optimization to model deployment and inference.
  • Achieved up to 2× inference improvement across selected workloads.
  • Made training about 2x faster and cheaper.
LLMVLMMultimodal AIQuantizationCUDARabbitMQRedisCeleryDockerPrometheusGrafanaKibanaGPU-aware scheduling

Product Engineer

Faz AustraliaAustraliafounderSep 2024 – Jul 2025

Summary

Served as a founding engineer for MedTalk, building production ML, backend, MLOps, and security systems for a clinical AI platform used by 50,000+ clinicians. Fine-tuned and deployed speech, medical NLP, and LLM models for clinical transcription, structured generation, and domain-specific healthcare workflows. Reduced real-time speech-pipeline latency from approximately 150 ms P50 post-VAD to ~50 ms through model and inference optimization. Built production RAG, LLM, WebSocket, model-serving, and streaming inference pipelines for medical transcription and knowledge retrieval. Deployed and operated ML workloads across AWS and RunPod, supporting cloud-based model serving and low-latency production inference.

Results

  • Built systems for a platform used by 50,000+ clinicians.
MLOpsNLPLLMRAGWebSocketsAWSRunPod

Computer Vision Engineer

VisionRDPakistanworkJun 2024 – Nov 2024

Summary

Developed and fine-tuned computer vision and diffusion models using industrial image datasets for defect detection, predictive maintenance, and multimodal AI applications. Built production CV systems for Thal Engineering and KIA Motors, automating assembly-process validation and manufacturing defect detection. Worked across dataset preparation, model training, evaluation, and deployment for real-world industrial vision applications.

Results

  • Caught assembly issues in real time by checking if the right part was present and placed correctly.
  • Reached around 92 to 94% accuracy for assembly-process validation systems.
  • Improved productivity by about 23% according to the client team.
  • Reduced manual inspection time for assembly processes.
Computer VisionDiffusion ModelsIndustrial Image Datasets

Skills

Technical

LLMs
VLMs
Deep Learning
LoRA/QLoRA
PythonPython
PyTorchPyTorch
Model Fine-Tuning
Transformers
Diffusion Models
CUDA
TensorFlowTensorFlow
NLP
Computer Vision
Hyperparameter Experimentation
Pandas
NumPy
DockerDocker
FastAPIFastAPI
Model Serving
Production Monitoring
RunPod
RedisRedis
RabbitMQ
Celery
CUDA Profiling
Distributed Training/Inference
Mixed Precision
GitGit
REST APIs
WebSockets
Supervised Learning
Quantization
Data Preprocessing
Feature Engineering
Dataset Preparation
Data Augmentation
Model Evaluation
Benchmarking
CI/CD
Prometheus
Grafana
Kibana
PostgreSQLPostgreSQL
SQLSQL
GoGo
Experiment Tracking
Weights & Biases
Data Labeling
AWS
AWS Lambda
C++C++
Rapid prototyping
Independent project creation
First-principles problem solving
Agent Workflows

Projects

CarDealer AI

personal
A live AI product for automotive dealerships where I personally owned parts of the architecture and workflows.

Multimodal Model Inference & Optimization Platform

Lead Developerprofessional
Built a unified execution layer for heterogeneous image, video, audio, 3D, LLM/VLM, and multimodal architectures, supporting routing, batching, streaming, hot-swapping, and VRAM-aware scheduling. Integrated caching, sparse/parallel attention, compilation, and automated profiling to benchmark latency, throughput, VRAM usage, and output-quality trade-offs across models and hardware. Designed the platform to enable new model architectures and inference strategies to be integrated without changing the core orchestration layer.
LLMVLMMultimodal ArchitecturesVRAM-aware schedulingCachingParallel AttentionCompilation

Multi-Architecture LoRA/QLoRA Post-Training Platform

Lead Developerprofessional
Built a modular LoRA/QLoRA post-training framework supporting diffusion, LLM, and VLM architectures with configurable adapter injection, dataset preprocessing, captioning, checkpointing, validation, and evaluation. Implemented 4-bit/8-bit quantization, mixed-precision training, gradient checkpointing, gradient accumulation, and memory-efficient attention to reduce GPU-memory requirements during training. Ran controlled experiments across LoRA rank, alpha, target modules, learning rate, optimizers, schedulers, and training configurations to study quality and convergence trade-offs. Built adapter evaluation, dynamic LoRA loading, merging, weighted composition, and production hot-swapping.
LoRAQLoRADiffusionLLMVLMQuantizationMixed-precision trainingGradient checkpointingMemory-efficient attention

VLLM Microservice Platform in Go

Developerprofessional
Built a model-agnostic Go microservice layer exposing JSON APIs for prompt enhancement, image/video captioning, age detection, and multi-category content-safety scoring. Implemented model-specific inference configurations across FLUX, WAN, LTX, MiniMax, Klein, Anima, and related architectures using KV caching, compilation, configuration tuning, and GPU/host memory management. Containerized services with Docker and added production observability with Prometheus for errors, latency, endpoint health, and resource utilization.
GoJSON APIsFLUXWANLTXMiniMaxKV cachingDockerPrometheus

Content Safety Classification Model

ML Engineerprofessional
Trained a content-safety classification model that outputs confidence scores for NSFW detection and age-related content classification. Built the model to distinguish NSFW vs. safe content and identify whether detected content involved an underage category, enabling more granular safety decisions. Performed knowledge distillation on a generative AI model, focusing on model compression, post-training, and improving efficiency for production deployment.
Knowledge DistillationModel CompressionClassificationNSFW Detection

MedTalk | Real-Time Multimodal Clinical AI

Founding Engineerprofessional
Built and deployed low-latency speech and language-model pipelines for live clinical transcription and structured generation using WebSocket streaming and production model serving. Reduced speech-pipeline latency from approximately 150 ms P50 post-VAD to ~50 ms while integrating fine-tuned speech, medical NLP, and LLM components.
WebSocketsSpeech-to-TextMedical NLPLLMModel Serving

Automated Lumber Takeoff and Cost Estimation from Construction Drawings

ML Engineerprofessional
Built a multi stage computer vision pipeline using YOLO for object detection, SAM for segmentation, and Detectron2 based models for region level analysis, extracting lumber types, dimensions, quantities, and structural context from construction drawings. Integrated Gemma and Qwen vision language models for semantic interpretation of drawing annotations and material context, then deployed the pipeline as a FastAPI cloud service for automated bill of materials and cost estimation.
YOLOSAMDetectron2GemmaQwen VLMFastAPI

MiniMax Music 3 Reverse Distillation

personal
A research side project to recover a missing audio tokenizer from the generator itself using a reverse distillation pipeline where the model generates supervision for an encoder to learn hidden audio codes.
Reverse distillation

Education

Bachelor of Science (Honours)

University of Engineering & Technology (UET), TaxilaSoftware EngineeringGPA 3.942020 – 2024
Hire MunazaReach out about a role, a contract or a conversation.For recruiters

Munaza's twin is AI, it can make mistakes.