Muqarrab Rahman

Muqarrab Rahman

Senior AI & Data Engineer specializing in LLM Agents, RAG, and Cloud-Native Data Platforms.

Based in: Lahore, PakistanCurrently: Senior AI & Data Engineer, IBHC (In Human Business Capital), Aslase Group - powered by Inception (G42)Experience: 10 years

Hire me

Muqarrab is an experienced engineer with over 10 years of technical expertise, including 4 years focused on building production-grade AI applications and agentic workflows. He has a proven track record of architecting complex data pipelines and ML orchestration engines for enterprise-scale talent and HCM platforms.

Experience

Senior AI & Data Engineer

IBHC (In Human Business Capital), Aslase Group - powered by Inception (G42)Dubai, UAEworkJul 2025 – Present

Summary

Takafo+ (takafo.ai / mubadala.com): Architected the ML orchestration engine behind Mubadala Investment Company's talent platform, ranking a corpus of 20,000+ active applicants and returning top-k shortlists per requisition using business-defined scoring weights. Humentra (humentra.ai): Engineered the AI layer of an AI-native HCM platform spanning 9 lifecycle modules, using hierarchical agents to cross-map and automate recruiting, assessment, onboarding, succession, and related workforce workflows.

What I did

  • Designed the scoring stack from AI signals across CV-to-JD skill matching, job-responsibility alignment, and functional-experience depth, emitting confidence values and reasoning traces at each stage while maintaining PII-compliant processing.
  • Extended the same matching engine to reverse AI job sourcing and adapted the architecture for the Department of Government Enablement (DGE) and Remote Work Emirates Foundation use cases.
  • Built bilingual English/Arabic resume-to-JD scoring across all four language permutations, combining duration-weighted functional experience, skill matching, and career-level distance; generalized seniority into a six-tier taxonomy and mapped enterprise skill banks to ESCO.
  • Standardized retrieval using BGE bi-encoders, cross-encoder reranking, and FAISS; reduced LLM spend in the CV-parsing pipeline by moving structured extraction to smaller models with JSON tool calls, reusing parsed fields downstream, and adding prompt/tool caching.
  • Architected reusable multi-agent workflows using LangGraph and LangChain with tool calling, state management, memory, conditional routing, context retention, inter-agent communication, and multi-step orchestration for enterprise document-processing and conversational-AI use cases.
  • Built an LLM-based interview agent that conducts structured, multi-turn candidate interviews, generates contextual follow-up questions, and evaluates responses in real time.
  • Developed a workspace-connected AI agent integrated with internal APIs, enabling users to query dashboard and workspace data in natural language without manual report lookup.
  • Introduced a staged orchestration approach using deterministic/rule-based filtering and candidate scoring for initial screening to reserve LLM calls for steps requiring deeper reasoning.
  • Designed a hierarchical agent architecture using LangGraph where a supervisor/router interpreted user intent to identify relevant HCM lifecycle modules.
  • Implemented a shared state as a control plane between agents to carry conversation context, identified entities, and results from downstream agents.
  • Managed cross-module queries by routing state sequentially through specialized agents, such as recruitment and interview-scheduling agents.
  • Used explicit state transitions and checkpoints with structured outputs to ensure predictable handoffs and allow for retries or requests for missing information.

Results

  • Ranked a corpus of 20,000+ active applicants for Mubadala Investment Company.
  • Reduced LLM spend in the CV-parsing pipeline by moving structured extraction to smaller models with JSON tool calls and prompt/tool caching.
ML OrchestrationBGE bi-encodersCross-encoder rerankingFAISSJSON tool callsLangGraphLangChainConversational AIMulti-agent workflows

AI & Data Engineer

DataPrismworkFeb 2024 – Dec 2025

Summary

Designed and implemented an Apache Airflow ETL pipeline migrating millions of records from PostgreSQL to Amazon Redshift through an S3 staging layer, with fault-tolerant scheduling, retries, and monitoring. Built a personal-data aggregation platform integrating Gmail, Calendar, iDrive, and bank transaction data via Plaid into Supabase; developed an MCP-connected retrieval agent and conversational interface for natural-language queries over the consolidated data. Built API connectors and cloud data pipelines for AccuLynx, HubSpot, and other third-party systems, loading data through Google Cloud Run into BigQuery and orchestrating workflows with Cloud Scheduler, Workflows, Firestore, and Secret Manager. Developed an influencer-content data pipeline using Python and dbt to collect, transform, and warehouse Instagram data in Snowflake for downstream analytics and dashboarding. Built a Selenium-based product-ranking automation solution to monitor rankings, simulate user interactions, analyze ranking factors, and support data-driven product visibility optimization. Developed an end-to-end job scraping and application automation system using Scrapy, Selenium, and Python, deployed with AWS Lambda, DynamoDB, and S3 for serverless processing and storage. Developed a Facebook data pipeline using APIs and Python, processed the collected data with the OpenAI API to generate insights, deployed the workflow on a GCP VM, and stored processed data in BigQuery for analysis.

Results

  • Migrated millions of records from PostgreSQL to Amazon Redshift.
  • Managed a throughput of 10-15 GBs of data per sync
  • Handled around 10-15 million records per sync
  • Scheduled data movement 2 times per day as a continuous process
Apache AirflowPostgreSQLAmazon RedshiftS3PlaidSupabaseMCPAccuLynxHubSpotGoogle Cloud RunBigQueryCloud SchedulerWorkflowsFirestoreSecret ManagerPythondbtSnowflakeSeleniumScrapyAWS LambdaDynamoDBOpenAI APIGCP VM

Data Engineer

Get Gaari Technologies (Pvt) LtdLahore, Pakistanwork2023 – 2024
Prepared and cleaned operational datasets for downstream analytics, reporting, and business decision support. Wrote complex OLAP queries and developed Power BI dashboards for cross-functional stakeholder reporting.
OLAPPower BIData Cleaning

Skills

Technical

PythonPython
Pandas
RAG
SQLSQL
Prompt Engineering
REST APIs
PostgreSQLPostgreSQL
LangGraph
BigQuery
LangChain
OpenAI API
AWS (S3, Lambda, Redshift, DynamoDB)
GCP (Cloud Run, BigQuery, Firestore, Secret Manager, VM)GCP (Cloud Run, BigQuery, Firestore, Secret Manager, VM)
Apache Airflow
LLM Agents
ETL/ELT
DockerDocker
LinuxLinux
Gemini API
MCP
BGE Embeddings
Cross-Encoder Reranking
FAISS
Vector Search
dbt
Google Cloud WorkflowsGoogle Cloud Workflows
Cloud Scheduler
Data Modeling
Amazon Redshift
Snowflake
DynamoDB
SupabaseSupabase
FastAPIFastAPI
Selenium
BeautifulSoup
Requests
Power BI
Dashboarding
Scrapy
Data Marts
cron
C++C++
Regression Modeling
Independent project creation

Projects

Marketing Campaign Analytics - Big Data, AWS

Lead Developerpersonal
Built a real-time analytics pipeline for monitoring prospect engagement using AWS S3, Glue, Athena, Redshift, Databricks, VPC, IAM, and SageMaker.
AWS S3GlueAthenaRedshiftDatabricksVPCIAMSageMaker

Term-Recency for TF-IDF and BM25 Weighting - Information Retrieval

Researcheracademic
Developed a method for incorporating term age and recency into document-term importance calculations for information-retrieval ranking.
TF-IDFBM25Information Retrieval

Affective Computing - Deep Learning

Developeracademic
Implemented a CycleGAN-based approach for transferring emotional expression between image domains.
CycleGANDeep LearningComputer Vision

Job-application automation pipeline

Independent Creatorpersonal
Built an independent end-to-end system from scratch to automate finding job openings, extracting information, and tracking applications.
PythonScrapyAWS LambdaDynamoDBS3

Education

Master's in Data Science (In Progress)

Information Technology UniversityData ScienceGPA 3.892020

Coursework

Big DataTools and Techniques for Data ScienceDeep LearningInformation Retrieval

Bachelor's in Electrical (Computer) Engineering

FAST-NUCES, LahoreElectrical (Computer) EngineeringGPA 3.152012 – 2016

Thesis. IoT-based Health Monitoring System

Recognition

Awards

Dean's Honor List

FAST-NUCES, Lahore2014

Fall 2013, Fall 2014

Certifications

SQL (Intermediate)

HackerRank

SQL (Basic)

HackerRank

ETL and Data Pipelines with Shell, Airflow, and Kafka

Coursera

SQL for Data Science

Coursera

Data Engineering for Everyone

DataCamp

Intermediate SQL and Joining Data in SQL

DataCamp

Snowflake SnowPro Core certification preparation

Snowflake

Talend Data Integration certification preparation

Talend

Getting Started with Data Warehousing and BI Analytics

Coursera
Hire MuqarrabReach out about a role, a contract or a conversation.For recruiters

Muqarrab's twin is AI, it can make mistakes.