
Mohammad Umair Farooqui
Currently: Software Engineer, Netsol TechnologiesExperience: 3 years
Experience
Software Engineer
Summary
What I did
- Engineered the backend for an internal AI query chatbot using a vector store with hybrid semantic and keyword search
- Paired BullMQ and Kubernetes to decouple and scale the document ingestion and embedding pipeline.
- Used BullMQ backed by Redis to manage a multi-stage ingestion queue including document parsing, chunking, embedding generation, and vector upsert.
- Implemented KEDA to scale worker pods horizontally based on Redis queue depth during daily delta syncs or bulk re-indexing.
- Fine-tuned Llama-3-8B-Instruct using QLoRA to serve as a fast, deterministic intent-classification and structured query-generation router.
- Built an automated evaluation pipeline to curate and clean roughly 12,000 prompt-completion pairs from production logs and synthetic edge cases.
- Applied deduplication, strict JSON schema validation, and human-in-the-loop verification for ambiguous queries during data preparation.
- Evaluated model checkpoints against holdout sets using JSON schema conformity rates, exact-match tool invocation, and LLM-as-a-judge comparisons.
- Implemented Direct Preference Optimization (DPO) using paired 'chosen' and 'rejected' responses to optimize implicit rewards and reduce hallucinations.
- Utilized bfloat16 base weights with DeepSpeed ZeRO-2 across multi-GPU setups to preserve logical reasoning capabilities by removing quantization loss.
- Performed synthetic negative mining and data filtering using automated heuristics and an ensemble judge to over-index on hard negatives.
Results
- Achieved a 98.4% schema valid rate on structured output generation, up from approximately 72% zero-shot.
- Reduced inference latency by 65% compared to external API-based models by fine-tuning and self-hosting a smaller model.
- Significantly dropped per-query compute cost while gaining complete control over deployment and latency targets.
- Reduced edge-case failure rates by roughly 38% compared to the SFT QLoRA baseline by shifting to Direct Preference Optimization (DPO).
- Improved model calibration through DPO alignment, enabling the model to reliably defer or trigger fallback tools when uncertain.
- Serves around 1,200 to 1,500 internal employees across global teams
- Handles 8,000 to 10,000 queries daily
- Peak concurrency of roughly 60–80 requests per second during business hours
- Indexes and chunks over 500,000 internal documents, technical manuals, and policy records
- Processes scheduled daily delta syncs of tens of thousands of modified records
- Sub-second context retrieval of under 200 ms before hitting the LLM pipeline
Software Engineer
Summary
What I did
- Optimized order book performance by switching from array traversals to flat Map lookups for O(1) state updates.
Results
- Reduced UI latency from ~450 ms to under 45 ms for the trading platform's order book.
- Decreased client CPU usage by ~65% through virtualization and optimized state management.
- Throttled incoming WebSocket ticks into 50 ms windows using requestAnimationFrame.
Associate Software Engineer
Summary
What I did
- Built a tokenized styling system using CSS custom properties and Tailwind CSS to enable runtime tenant switching without flash of unstyled content.
- Architected a modular grid engine using CSS Grid and @container queries to allow dashboard cards to adapt layouts based on container dimensions.
- Integrated Framer Motion and custom CSS transitions for layout morphing and state transitions using composite layers to avoid repaints.
Results
- Owned the styling architecture and frontend layout system from scratch for a complex, multi-tenant analytics dashboard.
- Maintained a Cumulative Layout Shift (CLS) score of 0.00 while streaming asynchronous metrics hydrated in the background.
- Reduced the initial CSS payload by ~55% and significantly improved First Contentful Paint by replacing runtime CSS-in-JS with compiled, tokenized Tailwind CSS.
- Halved development time for new feature pages by providing a strict, reusable primitive component library.
Skills
Technical
Projects
Audio Labeling & Scrubbing Prototypes
Snowgum
Instantly Legal
Success AI
Ciities
Education
BS Computer Science
Recognition
Awards
Solutions Contributor Leetcode
Active solutions contributor with over 400 problems solved.
Top 100 in Biweekly Leetcode cup x2
Finished in the top 100 twice in Biweekly Leetcode cup competitions.
Certifications
Complete Web Development Bootcamp
Node.js, Express, MongoDB & More: The Complete Bootcamp
Advanced NodeJS: Scalable Backends
DevOps Bootcamp cohort
Mohammad's twin is AI, it can make mistakes.
More people on Kanvis
Abdullah Farid
Lahore
AI/ML Engineer specializing in LLM Fine-Tuning and Full-Stack Voice AI Solutions
English · Python · Urdu · Hindi
kanvis.me/abdullah-farid
Syed Irtaza
Multan
AI Automation Engineer | n8n & RAG Specialist | React Native Developer
n8n · React Native · REST APIs · Git
kanvis.me/syed-irtaza
Ali Rana
Lahore
Full Stack Developer & AI Voice Agent Engineer
Urdu · JavaScript (ES6+) · React.js · Next.js 15
kanvis.me/ali-rana
Dayyan Saeed
Lahore
AI & ML Engineer specializing in Autonomous Multi-Agent Systems and RAG Pipelines.
Python · LangChain · LangGraph · OpenAI API
kanvis.me/dayyan-saeed



