Muneeb Ahmad

Muneeb Ahmad

AI and Full Stack Software Engineer specializing in Speech AI and Scalable Backend Systems

Based in: Lahore, PKCurrently: Software Engineer (AI and Full Stack), Alethia AIExperience: 1 year

Hire me

A Computer Science graduate from LUMS with extensive experience in building AI-driven applications, including real-time voice role-play systems and OCR pipelines. He excels at optimizing backend architectures using Django, FastAPI, and AWS to ensure high performance and reliability in production environments.

Experience

Software Engineer (AI and Full Stack)

Alethia AIworkMar 2026 – Present

Summary

Fixed transcription jobs that failed whenever two ran at once on a single GPU. Jobs now queue behind FastAPI, and each WhisperX run frees its GPU memory before the next one starts. Replaced anonymous speaker labels in meeting transcripts with real names, by matching voice embeddings against 20 to 30 second enrollment recordings stored in pgvector. Recovered genuine speakers that far-field recordings were rejecting. Found their similarity scores sat between 0.63 and 0.75 and lowered the match threshold from 0.75 to 0.62. Built the scoring system for a real-time voice role-play training product, with weighted criteria, must-do and never-do rules and deal-breakers. Covered by 59 tests. Stopped failed AI grading calls from being recorded as a 0% score. Refused, truncated or malformed LLM responses are now rejected against a strict JSON schema. Rebuilt the backend of an AI video clipping platform on Django, Celery, Postgres and SQS, so long-running jobs survive restarts and deploys instead of being lost. Replaced free-form candidate status on an internal hiring platform with a defined workflow of 13 statuses and 25 valid transitions, backed by 241 API tests. Built a SaaS-spend tracker end to end in Next.js and MongoDB that alerts owners in Slack before each subscription renewal.

What I did

  • Tuned the speaker-match threshold from 0.75 to 0.62 for speech processing after analyzing real similarity scores.

Results

  • Rewrote the GPU Rental platform's log relay service to stream on demand using short-lived signed tickets.
  • Reduced upstream connections from one per running session to one per open log panel.
FastAPIWhisperXGPUpgvectorLLMJSON schemaDjangoCeleryPostgresSQSNext.jsMongoDBSlackGo

Associate Software Engineer

TkxelworkSep 2025 – Feb 2026
Added Stripe payments to a production Django codebase, then stopped duplicate charges under concurrent requests with thread locking and idempotency keys. Put Amazon CloudFront in front of backend image delivery, which cut image load times by about 50%.
StripeDjangoAmazon CloudFront

Engineering Intern

TkxelinternshipJul 2025 – Aug 2025
Automated AWS EC2 provisioning with Terraform and Ansible, which cut manual deployment effort by 60%.
AWS EC2TerraformAnsible

Skills

Technical

LLM API integration
PythonPython
DjangoDjango
WhisperX
pgvector
TypeScriptTypeScript
JavaScriptJavaScript
RAG
prompt engineering
Django RESTDjango REST
FastAPIFastAPI
Node.jsNode.js
ExpressExpress
Celery
PostgreSQLPostgreSQL
ReactReact
Next.jsNext.js
DockerDocker
GitGit
FlaskFlask
RedisRedis
MongoDBMongoDB
Stripe
pyannote
Deepgram
Twilio
LangChain
AWS
GCP Compute EngineGCP Compute Engine
TerraformTerraform
GoGo
React NativeReact Native
WebSockets
C++C++
CC
Fine-tuning
OCR
Data Preparation
Model Evaluation
Model Serving
Model Training

Projects

Toxicity Classifier

personal
Trained and evaluated a toxicity classifier on a dataset of tweets.
Machine Learningbert-base-uncasedLSTMNaive BayesSGD

Migraine Risk Model

personal
Built a migraine risk model using a national health survey dataset and served it through an API.
Machine LearningAPI

AI Call Automation Platform

Lead Developerpersonal
Built a phone-based AI agent that handles inbound and outbound calls. It converts speech to text in real time, passes the conversation to an LLM and answers with generated speech. Added call transfers and automated responses to both the inbound and outbound workflows. Split inventory, pricing validation and order processing into separate services the LLM uses to take orders mid-call. Containerized the platform with Docker and deployed it through a CI/CD pipeline.
Node.jsExpressTwilioDeepgramWhisperOpenAI APIDockerSpeech-to-textLLMTTS

Knowledge Graph RAG for Clinical Medicine

Researcher/Developeracademic
Combined knowledge graph reasoning with semantic vector search in a RAG framework for clinical questions, which scored 12 to 15% above the benchmarks it was tested against.
PythonKnowledge GraphSemantic Vector SearchRAG

Fault-tolerant Key-Value Store

personal
Built a five-node cluster key-value store based on the Raft consensus algorithm for a distributed systems course.
Go

Education

Bachelor of Science

Lahore University of Management SciencesComputer Science2021 – 2025
Hire MuneebReach out about a role, a contract or a conversation.For recruiters

Muneeb's twin is AI, it can make mistakes.