Muhammad Haseeb

Muhammad Haseeb

Founding Product Engineer & Machine Learning Specialist

Based in: PakistanCurrently: Founding Product Engineer (Team Lead), NEXAExperience: 2 years

Hire me

Muhammad Haseeb is a versatile engineer with extensive experience in architecting scalable AI-powered systems and optimizing machine learning models for production. He has a proven track record of leading technical product decisions and contributing to high-impact open-source projects in the ML ecosystem.

Experience

Founding Product Engineer (Team Lead)

NEXARemotefounderMay 2026 – Present

Summary

Owning end-to-end product engineering, architecting scalable backend systems, building responsive frontend experiences, and integrating AI capabilities into production workflows. Leading technical and product decisions as a founding engineer, translating early customer needs into reliable AI-powered features while driving rapid product iteration.

What I did

  • Maintained a feature ship velocity of one major release every 10 days.

Results

  • Released a major feature to Main every 10 days while handling daily bug fixes and customer issues.
AIBackend SystemsFrontend

ML Research Assistant

CITY Lab (Efficient ML & Interpretability)work2025 – May 2026

Summary

Built a contrastive-learning model compression framework that aligns compressed models with pretrained, fine-tuned, and historical checkpoints, matching original accuracy at up to 99% compression across ViTs and LLMs. Compressed diffusion models to 4-bit precision using weight-sensitivity statistics, quantizing dynamic activations via signal-to-noise ratio to retain generation quality (55% improvement over prior methods). Stabilized distributed training on highly unbalanced datasets by diagnosing and mitigating weight-interference issues during server averaging.

Results

  • 55% improvement over prior methods in generation quality retention
Contrastive-learningViTsLLMsDiffusion modelsQuantization

Open Source Software Engineer

DatacurveAI (YC S24)workDec 2025 – Mar 2026

Summary

Engineered features for open-source ML repositories and designed expert-level coding challenges used to train and evaluate code-generation LLMs.

What I did

  • Designed expert-level coding challenges to train LLMs by architecting complex problems involving high-level reasoning.
  • Reviewed reasoning traces to ensure models were not gaming the system or hacking the reward mechanism.

Results

  • Engineered features for open-source ML repositories, including work on group query attention.
MLLLMsGroup Query AttentionLinear Algebra

Skills

Technical

PythonPython
PyTorchPyTorch
Transformers
LLMs
ReactReact
RAG
FastAPIFastAPI
REST APIs
Node.jsNode.js
LoRA
AI Agents
Vector Databases
TypeScriptTypeScript
JavaScriptJavaScript
SQLSQL
Next.jsNext.js
PostgreSQLPostgreSQL
MERN Stack
AWS (EC2, ECR, S3, IAM)
DockerDocker
LinuxLinux
BashBash
System Optimization
CI/CD
C++C++
TensorRT
NestJS
ONNX
LangChain
RedisRedis
RustRust
KubernetesKubernetes
Fine-tuning
Diffusion Models
Group Query Attention
Linear Algebra

Projects

Multi-Agent Full-Stack ML Platform

Lead Developerpersonal
Built a full-stack application that automates ML workflows: a multi-agent pipeline handles feature selection, data cleansing, and model selection, while an integrated RAG system answers natural-language questions. Deployed on AWS EC2, Lambda, and S3.
ReactFastAPIDockerGeminiRAGAWS (EC2, Lambda, S3)

Autonomous Dev Agent

Lead Developerpersonal
Built an AI system that generates full-stack apps from a prompt and Figma file. Powered by Google ADK, A2A, and MCP protocols, a system of agents scaffolds REST APIs, writes code, and configures deployment.
TypeScriptGoogle ADKModel Context ProtocolAgent-2-Agent Protocol

RAG-Powered Code Search & Understanding

Lead Developerpersonal
Built a semantic search tool for developers to analyze large codebases, including a custom GitHub MCP server that vectorizes repositories so users can query and debug software in natural language via a Streamlit interface.
StreamlitMCPGitHub APIsLLMs

Adapter-hub/adapters Contribution

personal
Contributed a pull request to the adapter-hub/adapters repository focused on group query attention.
Group Query Attention

Apple Native ViT Architecture Optimization

personal
Worked on Apple's native ViT architecture repository using reductions and linear algebra operations to reduce tensor operations and memory reads. Focused on reducing multiple layers while maintaining activation output semantics by handling skip connections with matrix operations.
Linear Algebra

Qwen 3B Misalignment Research

personal
Used LoRA fine-tuning to instill emergent misalignment behavior in a Qwen 3B model, evaluating results using an LLM-as-a-judge approach.
LoRAGradient accumulationCheckpointingQwen 3B

Manim Code Generation Models

personal
Fine-tuned Qwen and Llama models using QLoRA to produce Manim code for a personal teaching assistant startup concept.
QLoRAQwenLlamaManim

Education

B.S. Computer Science

Lahore University of Management Sciences (LUMS)Computer ScienceSep 2022 – May 2026

Coursework

Software EngineeringData StructuresAlgorithmsDatabase SystemsLLM SystemsAdvanced Machine Learning

Social impact & activities

Open Source Contributor

pytorch-image-models (timm)volunteer

Integrated F1, precision, and recall metrics directly into distributed training.

Open Source Contributor

adapters (Hugging Face)volunteer

Implemented PEFT support for Group Query Attention models to fine-tune with LoRA.

Hire MuhammadReach out about a role, a contract or a conversation.For recruiters

Muhammad's twin is AI, it can make mistakes.