Muhammad.
Kanvis
M

Muhammad Aarash Abro

Research Engineer | AI Infrastructure | High-Throughput RAG & Distributed Systems
Hire Me

Overview

Research Engineer with a track record of bridging the gap between complex algorithmic theory and massive-scale production. Specialized in high-throughput RAG pipelines (10M tokens/sec), multi-agent regret learning for finance, and efficient vision backbones. Thrives in high-pressure, fast-paced environments where technical complexity meets real-world scale. Experience spans from publishing at ICAIF to consulting for global giants like ExxonMobil and Warner Bros. Discovery.

Experience

2023 – Present

WORK

Research Lead (ML for Chemical Reaction Modeling)

Auxilart (Industry-Academic Collaboration)
Key Responsibilities: • Organic reaction mechanism identification under real-world data heterogeneity (Under Review, ESCAPE 36): Built Transformer-based encoder for irregular chemical trajectories and sparse-autoencoder based domain adaptation module; achieved 99.8% accuracy (20% masking) and 93.4% (40% masking); implemented full ODE-based data-generation pipeline. Achievements: • Achieved 99.8% accuracy (20% masking) and 93.4% (40% masking) in organic reaction mechanism identification.
Organic reaction mechanism identification under real-world data heterogeneity (Under Review, ESCAPE 36): Built Transformer-based encoder for irregular chemical trajectories and sparse-autoencoder based domain adaptation module; achieved 99.8% accuracy (20% masking) and 93.4% (40% masking); implemented full ODE-based data-generation pipeline.

Key Achievements

Achieved 99.8% accuracy (20% masking) and 93.4% (40% masking) in organic reaction mechanism identification.
TransformerSparse-autoencoderDomain adaptationODE-based data-generation

Jul 2022 – Present

WORK

Machine Learning Consultant

Zeta Solutions
Key Responsibilities: • ATR SmartProcedures — AI-Driven Factory Documentation: Automated digitization of factory equipment manuals into proprietary formats using LLMs, enabling less integration with internal software, reducing turnaround time per-table by 80%. • UHY Prime HK — Audit Automation: Automated auditing by leveraging LLMs and VLMs to match invoices and documents with transactions, verify accuracy, and flag discrepancies, reducing number of man-hours required by up to 90%. • Warner Bros. Discovery — Post Production Video Assistant Automation: Built computer vision models for automated text detection in videos, streamlining translation and production in the entertainment industry. • Whichdraft — AI Legal Assistant for Contract Generation: Developed agentic workflows for AI-generated wizards to assist novice lawyers in drafting standard contracts to reduce time spent on boilerplate work. • Enviro AI — Environmental Compliance Chatbot: Built an agentic reasoning system using TCEQ knowledge bases to automate environmental compliance reviews and permit applications for clients such as ExxonMobil and Dow Chemicals. • BDO Global — Efficiency Modeling and Routing Optimization: Extracted oil-pump efficiency curves from legacy documents using computer vision and curve fitting, tating optimized oil flow routing. Achievements: • Reduced turnaround time per-table by 80% for factory documentation. • Reduced number of man-hours required for auditing by up to 90%. • Streamlined translation and production for Warner Bros. Discovery. • Automated environmental compliance reviews for clients like ExxonMobil and Dow Chemicals.
ATR SmartProcedures — AI-Driven Factory Documentation: Automated digitization of factory equipment manuals into proprietary formats using LLMs, enabling less integration with internal software, reducing turnaround time per-table by 80%.
UHY Prime HK — Audit Automation: Automated auditing by leveraging LLMs and VLMs to match invoices and documents with transactions, verify accuracy, and flag discrepancies, reducing number of man-hours required by up to 90%.
Warner Bros. Discovery — Post Production Video Assistant Automation: Built computer vision models for automated text detection in videos, streamlining translation and production in the entertainment industry.
Whichdraft — AI Legal Assistant for Contract Generation: Developed agentic workflows for AI-generated wizards to assist novice lawyers in drafting standard contracts to reduce time spent on boilerplate work.
Enviro AI — Environmental Compliance Chatbot: Built an agentic reasoning system using TCEQ knowledge bases to automate environmental compliance reviews and permit applications for clients such as ExxonMobil and Dow Chemicals.
BDO Global — Efficiency Modeling and Routing Optimization: Extracted oil-pump efficiency curves from legacy documents using computer vision and curve fitting, tating optimized oil flow routing.

Key Achievements

Reduced turnaround time per-table by 80% for factory documentation.
Reduced number of man-hours required for auditing by up to 90%.
Streamlined translation and production for Warner Bros. Discovery.
Automated environmental compliance reviews for clients like ExxonMobil and Dow Chemicals.
LLMsVLMsComputer VisionAgentic workflowsTCEQ knowledge basesCurve fitting

Sep 2024 – Aug 2025

WORK

Research Assistant

Intelligent Machines & Sociotechnical Systems Lab, LUMS
Supervised by Dr. Hassan Jaleel
No-Regret Portfolio Optimization (Accepted, AI4DF Workshop @ ICAIF 2025): Developed a multi-agent regret learning framework for financial decision making and regime shifts; outperformed S&P 500 and gold on risk-adjusted returns; implemented large-scale simulation and stress-testing pipeline.

Key Achievements

Developed a multi-agent regret learning framework for financial decision making and regime shifts.
Outperformed S&P 500 and gold on risk-adjusted returns using a greedy variant of the no-regret approach.
Discovered that hindsight-based investing in constrained search spaces significantly outperforms industry benchmarks.
Implemented large-scale simulation and stress-testing pipeline.
Outperformed S&P 500 and gold on risk-adjusted returns.
Multi-agent regret learningSimulation pipelineStress-testing

Jun 2024 – Jul 2025

WORK

Research Assistant

Electrical Engineering Department, LUMS
Supervised by Dr. Muhammad Tahir
Self-Supervised Financial Modeling with RL: Developed multiresolution time-series representation model using transformer feature generator and PPO-guided pretext tasks; processed 20+ TB of data with Python–C++ pipelines on Cloud Run, BigQuery, and Cloud Storage.
Unsupervised Domain Adaptation with Sparse Autoencoders: Designed SAE-based domain-invariant encoder improving cross-domain and few-shot performance for vision backbones.

Key Achievements

Designed an SAE-based domain-invariant encoder for Unsupervised Domain Adaptation, outperforming DAN and DANN by over 20pp on PACS, CIFAR100, and Office datasets.
Achieved 7x faster convergence compared to state-of-the-art methods like SSRT and FFTAT, optimizing for efficiency-focused deployments.
Developed a multiresolution time-series representation model using a transformer feature generator and PPO-guided pretext tasks.
Processed 20+ TB of data using optimized Python–C++ pipelines on Google Cloud (Cloud Run, BigQuery, Cloud Storage).
Outperformed DAN/DANN by 20pp+ on vision benchmarks.
Achieved 7x faster convergence than SOTA methods.
Processed 20+ TB of data with Python–C++ pipelines.
RLTransformerPPOPythonC++Cloud RunBigQueryCloud StorageSparse AutoencodersSAE

Project Portfolio

PERSONAL

Present

Job Posting Analytics

Developer

Designed a data pipeline to scrape job postings and extract key skill requirements using Small Language Models (SLMs). Built a job board to aggregate frequently demanded skills by position, simplifying career progression and skill acquisition pathways.
Gemma 2 1BIndeed API/ScrapingData pipelineWeb scrapingSLMs
ACADEMIC

Present

Distributed Fault-Tolerant State Machine (Raft)

Lead Developer

Implemented the Raft consensus algorithm from scratch to manage a replicated state machine, ensuring fault tolerance and consistency across a distributed cluster. Used Redis for state persistence and Go's concurrency primitives for managing cluster communication and leader election.
GoRedisDistributed SystemsRaft Consensus
PERSONAL

Present

Scalable RAG Pipeline for Secure Knowledge Processing

Developer

Built a high-throughput, enterprise-grade RAG pipeline designed to embed massive data repositories in hours rather than months. Optimized for cost-efficiency, achieving 10M tokens/sec throughput on a $30 budget by leveraging maximal threading for I/O and parallel serverless processing for compute-bound tasks.
RAGEmbeddingsThreadingServerless processing

Education

Sep 2021 – Jul 2025

B.S. Computer Science

Lahore University of Management Sciences (LUMS)Specialization in Computer Science
GPA: 3.48

Relevant Coursework

Reinforcement LearningAdvanced Topics in MLMultiagent SystemsRoboticsDistributed SystemsData ScienceMachine LearningDeep Learning

Impact & Recognition

Patents & Publications

2025 · PUBLICATION

No-Regret Portfolio Optimization

Accepted, AI4DF Workshop @ ICAIF 2025

2025 · RESEARCH_PAPER

No-Regret Portfolio Optimization (AI4DF Workshop @ ICAIF 2025)

PUBLICATION

Organic reaction mechanism identification under real-world data heterogeneity

Under Review, ESCAPE 36

RESEARCH_PAPER

No-Regret Portfolio Optimization (AI4DF Workshop @ ICAIF 2025)

Developed a multi-agent regret learning framework for financial decision making and regime shifts; outperformed S&P 500 and gold on risk-adjusted returns; implemented large-scale simulation and stress-testing pipeline.

RESEARCH_PAPER

Organic reaction mechanism identification under real-world data heterogeneity (ESCAPE 36)

Built Transformer-based encoder for irregular chemical trajectories and sparse-autoencoder based domain adaptation module; achieved 99.8% accuracy (20% masking) and 93.4% (40% masking); implemented full ODE-based data-generation pipeline.

RESEARCH_PAPER

Organic reaction mechanism identification under real-world data heterogeneity (ESCAPE 36 / PSE)

Built Transformer-based encoder for irregular chemical trajectories and sparse-autoencoder based domain adaptation module; achieved 99.8% accuracy (20% masking) and 93.4% (40% masking).

RESEARCH_PAPER

No-Regret Portfolio Optimization (AI4DF Workshop @ ICAIF 2025)

Developed a multi-agent regret learning framework for financial decision making and regime shifts. Outperformed S&P 500 and gold on risk-adjusted returns using a greedy variant of the no-regret approach.

PUBLICATION

No-Regret Portfolio Optimization (AI4DF Workshop @ ICAIF 2025)

Vision & Goals

Core Drivers

What are my motivations

Solving highly challenging and stimulating problems

High-pressure environments

Fast-paced innovation

Principles

What I value

Performance under pressure

Algorithmic efficiency

Practical R&D

High-throughput engineering

Skills & Interests

Python

Hugging Face Transformers

NumPy

Pandas

Retrieval-Augmented Generation

Large Language Models

scikit-learn

PyTorch

Large-Scale Time-Series

C/C++

SQL

TensorFlow

BigQuery

Cloud Storage

Cloud Run

PostgreSQL

Online and Reinforcement Learning

No-Regret Learning

Self-Supervised Representation Learning

Domain Adaptation

Docker

Redis

Bash/Shell

JavaScript/TypeScript

Go

Let's Connect

I'd love to hear from you. Feel free to reach out.

These unlock once you finish onboarding and publish your profile.