Muhammad.
Kanvis
Muhammad Saram Hassan

Muhammad Saram Hassan

AI/ML Researcher | Mechanistic Interpretability & Model Compression | BS CS @ LUMS
Hire Me

Overview

Final-year CS student at LUMS with a research focus on LLM evaluation, mechanistic interpretability, and model compression. Published author (EMNLP Findings 2025) and SRI International research intern developing frameworks for conceptual understanding in LLMs. Expert in the full ML stack, from fine-tuning and compression to distributed systems and data pipelines.
Lahore, Pakistan

Experience

Aug 2025 – Present

INTERNSHIP

Research Intern

SRI International
Menlo Park, CA (Remote)
Developing a model-agnostic evaluation framework combining behavioral consistency analysis and representational geometry.
Conducting large-scale reproduction of 'Potemkin Understanding' across 7 LLMs, identifying critical methodological fragilities.
Implementing activation steering strategies (definition-addition vs. difference-based) on Llama-3.2-3B.
Analyzing cross-lingual contradiction hypotheses using CKA-based geometric stability.

Key Achievements

Developed the Concept Coherence Metric (CCM) to unify five fragmented evaluation silos: behavioral consistency, internal representation, causal intervention, robustness, and domain-specific semantics.
Introduced the Latent Direction Score (LDS) using Sparse Autoencoders to causally improve concept application via activation steering (+4.38pp in Classify tasks).
Identified up to 31% variance in Potemkin scores across runs, surfacing significant stochastic instability in current conceptual understanding benchmarks.
Verified computational reproducibility of Procedure 1 from Mancoridis et al. (ICML 2025) with max absolute deviation of 0.006.
Manuscript submitted to ACM REP 2026.
Identified judge sensitivity as a major confounder in LLM evaluation pipelines.
LLMsCentered Kernel Alignment (CKA)Sparse AutoencodersActivation SteeringMechanistic InterpretabilityTransformer Layers

Jul 2025 – Present

WORK

Research Assistant

Security & Privacy Lab, LUMS
Lahore, Pakistan
Robustness of Web Agents in the Wild: Adversarial Influence on Autonomous Agents. Supervised by: Dr. Fareed Zaffar (Associate Professor, LUMS) and Dr. Taha Khan (CMU CyLab)
Leading an SoK on vulnerabilities in autonomous web agents, distinguishing agentic risks from traditional crawler exploits.
Architecting a synthetic adversarial HTML testbed with attack primitives like MCP poisoning and CSS obfuscation.
Developing a behavioral fingerprinting framework analyzing HTTP headers and DOM traversal to detect LLM agent traffic.

Key Achievements

Authored a full Systematization of Knowledge (SoK) paper on prompt injection threats in LLM-powered agents.
Developed a novel 12-class attack taxonomy based on semantic channels (e.g., RAG poisoning, tool-mediated hijacking, obfuscated HTML).
Identified the 'reasoning gap': stronger models like Claude 3.7 are immune to direct injection but highly vulnerable to memory injection (55.1% ASR).
Discovered that LLM browsers bypass robots.txt while masquerading as human users, triggering programmatic ad-auctions.
Synthesized empirical findings across 12+ major benchmarks including WASP, Mind the Web, and AgentDojo.
LLM AgentsAdversarial HTMLDOM ManipulationHTTP FingerprintingPrompt InjectionAI Safety

Aug 2023 – Present

WORK

Teaching Assistant

LUMS
Assisting students and leading tutorials for core CS courses.
Conducted weekly office hours for 300+ students in Discrete Mathematics (CS-210).
Led tutorials and lab sessions for Computational Problem Solving (CS-100).
Developed supplementary problem sets and graded research projects.
Discrete MathematicsComputational ThinkingMentorship

2024

WORK

Instructor

LUMS Science School & UpliftAI
Designing and delivering technical curricula to large student cohorts.
Designed and delivered a Machine Learning curriculum to 100+ students.
Developed an interdisciplinary 'CS for Astronomy' course for 100+ students at FAST University.
Created hands-on coding exercises using PyTorch, scikit-learn, and Python.
Machine LearningPythonCurriculum DevelopmentPublic Speaking

Aug 2021

WORK

Peer Tutor — Mathematics

International School Lahore (ISL)
Teaching 500+ O/A-Level students.
Led the Students’ Mathematics Department.
Attended COMSATS Mathematics Summer Camp and earned a place on a faculty research team.
MathematicsTeaching
PART_TIME

College Counselor & SAT Tutor

MR Consultants
Assisting applicants and teaching SAT.
Assisted 60+ undergrad/master applicants with essays and college applications.
Taught SAT to 50+ students.
SAT
PART_TIME

Math Expert

Photomath
EdTech platform math expert.
Solved 200+ math problems and reviewed 500+ solutions for ML training data quality.
Mathematics

2025 – Jul 2025

WORK

Research Assistant

TPI Lab, LUMS & CommNets Lab, NYUAD
Lahore, Pakistan
Knowledge Retention and Transferability in Fine-Tuned LLMs [EMNLP Findings 2025]. Supervised by: Dr. Fareed Zaffar (LUMS) and Dr. Yasir Zaki (New York University Abu Dhabi)
Designed controlled experiments isolating task format as a predictor of knowledge integration robustness during SFT.
Curated a novel evaluation dataset of 126 atomic facts from 2024 world events to eliminate prior model knowledge.
Fine-tuned 5 model families across architectures including Gemini-1.5-Flash and Llama-3.2-3B.

Key Achievements

Demonstrated that comprehension-intensive tasks (QA) yield 48% knowledge retention vs. 17% for token-mapping tasks (translation).
Confirmed scaling laws for knowledge retention across Qwen2.5 models from 1.5B to 72B parameters.
Revealed that even high-retention models fail to generalize injected knowledge to indirect semantic queries.
Published at EMNLP Findings 2025 (Acceptance Rate: 17.35%).
LLMsSupervised Fine-TuningKnowledge Retention/IntegrationScaling Laws

Sep 2024 – May 2025

WORK

Research Assistant

CITY Lab @ LUMS
Lahore, Pakistan
Bias-Aware Interpretable Pruning for Robust Vision Models. Supervised by: Dr. Muhammad Tahir (Associate Professor, LUMS)
Developed a bias-aware pruning framework for CNNs (VGG-19) and ViTs, testing against Scrambled and Color Jittered variants.
Implemented a filter-locking mechanism to immunize top 40% of property-relevant filters from pruning.
Analyzed model confidence and representation geometry using entropy and CKA.

Key Achievements

Achieved a 1.8% increase in ID accuracy (62.47%) while compressing to 80% sparsity using PAP-Taylor with filter-locking.
Reduced average model entropy by ~47% (0.5097 vs 0.9636 baseline), signaling more confident and structured predictions.
Validated that filter-locking reduces representational drift in task-critical deep layers via CKA alignment.
Introduced a filter-locking mechanism to preserve domain-invariant representations for improved OOD robustness.
CNNsViTsPruningOOD RobustnessTaylor ApproximationsFisher Information

Jun 2024 – Jul 2024

WORK

Integration Lead / Partner

Uplift AI (YC Startup)
Integration of LLMs in Mobile Banking — Uplift AI (US-Based AI Startup)
Collaborated to integrate multilingual LLMs into mobile banking platforms with regional accent recognition.
Designed quality control and preprocessing pipelines including noise filtering and speaker diarization validation.
Led a team of 50+ students for data curation and localization.

Key Achievements

Led a cross-functional team of 50+ students to collect and curate 35,000+ regional language speech samples.
Developed modules for interaction tracking, user feedback analysis, and localization for Urdu, Punjabi, and Siraiki.
Dataset directly contributed to UpliftAI’s successful seed funding round from Y Combinator.
Achieved 95%+ transcription accuracy across dialect variants.
LLMsMobile BankingLocalizationUrduPunjabiSiraikiData Curation

2024 – 2024

WORK

Research Assistant

Human-Computer Interaction Lab, LUMS
Supervised by: Dr. Suleman Shahid (LUMS)
Applying HCI research methodologies to unify fragmented workflows (WhatsApp, Notion, Trello) for university societies.
Conducting structured problem analysis across four distinct problem areas.

Key Achievements

Designed a society-focused communication app prototype through iterative user studies and brainwriting/round-robin ideation.
Evaluated design alternatives against user needs and HCI principles to reduce cognitive load in university society workflows.
HCIPrototypingUsability TestingUser Research

Project Portfolio

PROFESSIONAL

Present

Voice AI for Regional Languages

Team Lead

Led a cross-functional team of 100+ students in partnership with Uplift AI to collect and curate 35,000+ regional language speech samples (Urdu, Punjabi, Siraiki) with regional accent recognition. Designed quality control and preprocessing pipelines ensuring 95%+ transcription accuracy.
Voice AIData CurationPreprocessing Pipelines
ACADEMIC

Present

Evaluating Architectural Bias and Generalization in Vision Models

Analyzed inductive biases (shape, texture, color, locality) across ResNet-50, ViT, and CLIP. Found ViT dominates on scrambled images due to global attention, while CLIP shows best shape generalization. Found ResNet-50 has the strongest color bias.
PyTorchCNNsViTsCLIPInductive Bias
PERSONAL

Present

Chrono Savior — 2D Time-Travel Adventure Game

Lead Developer

Developed a time-travel-themed 2D game combining shooter, action, and adventure genres with a team of 4; designed core mechanics including player movement, space combat, and time-based abilities. Implemented game modes (Infinite Mode, Space Travel, Ground Combat) and advanced UI for seamless navigation, pause/resume, and performance optimization.
UnityC#
ACADEMIC

Present

LLM Password Generation: Security, Usability & Policy Compliance

Researcher (Collaboration with CMU CyLab)

Evaluating LLM-generated passwords across security (entropy, dictionary resistance), usability (memorability), and policy compliance. Analyzing model size, temperature, and prompting strategies on generation quality. Investigating failure modes: policy violation patterns and hallucination rates.
LLMsSecurity AnalysisPrompt EngineeringPython
PERSONAL

Present

Raft-Based Fault-Tolerant Distributed Key-Value Store

Developer

Implemented a complete Raft consensus module in Go featuring leader election, log replication, and log compaction. Validated system safety and liveness across 3-node and 5-node clusters, achieving a 14/15 test pass rate (184.6s total runtime). Identified a specific safety edge case in the 'Figure 8 Unreliable' chaos test while successfully passing high-stress scenarios like 'Backup' and 'Persist'.
GoRaft ConsensusDistributed SystemsRPC
ACADEMIC

Present

Reproducing "Potemkin Understanding in Large Language Models"

Lead Researcher (Reproduction)

Large-scale reproduction and extension of benchmark-based claims about conceptual understanding across 7 LLMs and 32 concepts. Verified computational reproducibility of Procedure 1 from Mancoridis et al. (ICML 2025) and identified four critical sources of methodological fragility, including stochastic instability (up to 31% variance) and self-judging circularity.
PythonOpenAI APIAnthropic APIAutomated Evaluation PipelinesStatistical Robustness Testing
ACADEMIC

Present

Federated Learning Under Data Heterogeneity

Analyzed Federated Learning under data heterogeneity using Dirichlet distribution (alpha in {0.1, 0.5, 2.0}) to control label skew. Implemented and compared FedAvg, SCAFFOLD, and FedGH algorithms, identifying convergence-speed vs. accuracy tradeoffs.
PythonFederated LearningSCAFFOLDFedAvgFedGH
ACADEMIC

Present

Knowledge Distillation: Feature Alignment & Contrastive Representations

Compared four distillation methods (Logit Matching, CRD, Hints/FitNets) for a VGG-11 student. Hints (FitNets) achieved 61.3% accuracy, approaching the teacher ceiling. Established KL divergence as a reliable proxy for distillation quality.
PyTorchKnowledge DistillationFitNetsCRDGradCAM
ACADEMIC

Present

Model Compression via Structured and Unstructured Pruning

Developer

Applied L2-norm unstructured and channel-wise structured pruning to VGG-16, achieving 65.9% post-fine-tuning accuracy (+7.5% lift). Conducted layer-wise sensitivity analysis and implemented PTQ/QAT across multiple bit-widths (FP16, INT8, INT4) using torchao.
PyTorchVGG16PruningQuantization (PTQ/QAT)torchao
ACADEMIC

Present

Deep Unfolding for Robust Neural Network Design

Applied algorithm unrolling (ISTA to LISTA) for sparse signal recovery, accelerating convergence while maintaining theoretical guarantees. Integrated Robust PCA (RPCA) for foreground-background segmentation in noisy datasets.
PythonDeep UnfoldingISTA/LISTASparse RecoveryRPCA
ACADEMIC

Present

Systems Programming: Shell, Malloc & GC

Developed core systems programming components in C, including a command-line interpreter (Shell) with pipes and I/O redirection, a custom memory allocator with First/Best/Worst Fit strategies, and a mark-and-sweep garbage collector. Implemented IPC using fork/exec/wait.
CSystems ProgrammingIPCGarbage CollectionMultithreading

Education

Aug 2022 – May 2026

Bachelor of Science

Lahore University of Management Sciences (LUMS)Specialization in Computer Science
GPA: 3.67 / 4

Relevant Coursework

Advanced Topics in MLFoundations of Generative AIDeep LearningHuman-Computer InteractionDistributed SystemsOperating Systems

Societies & Activities

Dean's Honor List (2022-2025)

Aug 2021 – May 2022

A-Levels

International School Lahore (ISL)Specialization in Mathematics, Further Mathematics, Physics, Chemistry, Computer Science
Thesis: SAT: 1530 (800 Math, 730 EBRW) | 100% Merit Scholarship

Societies & Activities

President, ISL Robotics SocietyVice President, Student Council

Aug 2017 – May 2021

O-Levels

Lahore Grammar School Johar Town Senior Boys CampusSpecialization in 14 Subjects (10 A*s, 4 As)

Societies & Activities

School ValedictorianStudent Council PresidentHead Boy & Senior Prefect

Impact & Recognition

Awards

International School Lahore · 2021

100% Merit Scholarship

Awarded for A-Levels

NSS International Space Settlement Contest · 2021

4th Place Globally — NSS International Space Settlement Contest

Out of 1600+ international teams.

International Kangaroo Linguistics Contest · 2021

1st Place (Gold Medal) — International Kangaroo Linguistics Contest

Gold Medal among 3000+ competitors worldwide.

International Youth Mathematical Championship (IYMC) · 2020

Silver Medal — International Youth Mathematical Championship (IYMC)

Among 12,500+ participants across 98 countries.

International Astronomy Astrophysicists Contest

Bronze Medal — International Astronomy Astrophysicists Contest

Among 4700+ students.

CERN

Honorable Mention (Top 30) — CERN Beamline for Schools (BL4S)

Only Pakistani Team in Top 30 with Honorable Mention out of 1600+ international teams.

International Academic Marathon

15th Globally — International Academic Marathon

15th globally in the Mathematics category.

Pakistan National Teams

Shortlisted — IOI & International Philosophy Olympiad

Shortlisted for the national teams in Informatics and Philosophy.

LUMS

Dean's Honor List

Academic excellence award for years 2022–2025

HEC

Top 25 — Pakistan National Mathematics Talent Contest (NMTC)

HEC Pakistan.

Professional Certifications

ETS

TOEFL iBT: 116/120

Patents & Publications

2026 · PUBLICATION

Understanding or Imitation? Auditing Conceptual Understanding and Reasoning in Large Language Models

Large-scale reproduction and extension of Mancoridis et al. (ICML 2025); identifies four sources of methodological fragility and demonstrates that reasoning models reduce Potemkin rates.

arXiv:2505.17140 · 2025 · PUBLICATION

Data Doping or True Intelligence? Evaluating the Transferability of Injected Knowledge in LLMs

Demonstrates that comprehension-intensive SFT tasks yield substantially higher knowledge retention than token-mapping tasks across 5 model families. Confirmed scaling laws for knowledge retention.

Team:Essa Jan · Moiz Ali · Fareed Zaffar · Yasir Zaki

Governance & Advisory

LUMS Entrepreneurial Society (LES)2024

Project Head, Social Outreach Program

OTHER

Built an Excel-based Strategic Business Management simulation deployed to 500+ students at the LUMS Young Leaders Entrepreneurs Summit.

Key Impact

Updated LES role with business simulation impact.

LUMS SPADES2024

Research & Design Executive

OTHER

Team lead for 6 students on Pakistan’s first student-made propellant rocket targeting 1 km altitude. Overseeing Avionics, Propellants, and Design subsystems.

Key Impact

Updated SPADES R&D role with propellant rocket project.

LUMS SPADES2023

Director Observatory

OTHER

Revamped the LUMS observatory by repairing a Celestron 6SE AVX mount; organized 5+ interactive astronomy workshops.

Key Impact

Updated SPADES Observatory role with repair and workshop details.

Nergis Mavalvala Astronomy Society — LGS Johar Town

Founder

OTHER

Established Pakistan’s first high school observatory; fundraised and led acquisition of telescopes valued at PKR 1.5 million.

Key Impact

Added Nergis Mavalvala Astronomy Society role.

ISL Robotics Society2021 – 2022

Founder & President

OTHER

Founded and led the society from scratch; introduced LEGO and Arduino programs to 250+ students; coached teams to 12+ national competitions, winning 7+ contests.

Key Impact

Updated ISL Robotics role with student count and win metrics.

Social Impact & Activities

The Citizens’ Foundation Outreach

Volunteer Teacher

Developed STEM curriculum; taught web development and Arduino to 60+ underprivileged female students.

EngagementVOLUNTEER
Key Impact

Outreach reached 2000+ students.

Harsukh School of Knowledge and Arts

Lead Teacher & Project Coordinator

Developed a robotics curriculum for the Aga Khan Middle School Programme; taught 50+ students Arduino and LEGO EV3.

EngagementVOLUNTEER

2020 – 2022

Next Generation Pakistan

Core Team Member & Director IT

Fundraised for Trans Health Club with Project Pehchaan; led COVID-19 response reaching 400+ families.

EngagementVOLUNTEER

2019 – 2022

Climate Impact PK

Co-Lead, Climate Classrooms

Devised mobilization campaign with 11+ educational seminars; delivered Smog Action demands to Pakistan’s Ministry of Climate.

EngagementVOLUNTEER
LUMS Science School & UpliftAIWORKSHOP

Fundamentals of Machine Learning

Designed and taught a curriculum covering supervised/unsupervised learning, neural networks, and practical ML applications to 100+ students.

Vision & Goals

Core Drivers

What are my motivations

Advancing AI safety through interpretability and robustness research.

Bridging the gap between theoretical ML and production-quality systems.

Mentoring the next generation of engineers and researchers.

Objectives

Goals and Aspirations

01

Seeking global AI/ML research or engineering roles focusing on mechanistic interpretability, model robustness, and systems thinking.

Principles

What I value

Empirical Rigor

Systems Thinking

Collaborative Research

Educational Outreach

Skills & Interests

Pandas

Matplotlib

LLMs

NumPy

Python

PyTorch

Seaborn

Scikit-Learn

C/C++

HuggingFace

Git

Systems Programming

Model Compression

Mechanistic Interpretability

Knowledge Distillation

WordPress

Linux/Bash

OpenCV

LangChain

Unity (C#)

TensorFlow

Figma

Go

Federated Learning

Node.js

Docker

JAX

React

Communication

Working across multiple languages, enabling global collaboration and clear technical outreach.

Punjabi
English
Urdu
French
Persian

Let's Connect

I'd love to hear from you. Feel free to reach out.

These unlock once you finish onboarding and publish your profile.