Zeeshan Jamal

Zeeshan Jamal

Data Science Graduate | ML Engineer | Building Scalable Data Pipelines with Snowflake & Airflow

Based in: Lahore, PakistanMost recently: Freelance Web Developer, Domino's Pizza (Freelance)

Hire me

Data Science graduate from Riphah International University with a strong foundation in building end-to-end ML pipelines and scalable data architectures. Proven experience in automating complex data workflows using GitHub Actions, Snowflake, and PySpark. Passionate about bridging the gap between raw data engineering and actionable business insights, from medical image classification to customer churn prediction.

Experience

Freelance Web Developer

Domino's Pizza (Freelance)freelance

Research Assistant / Collaborator

PostDoc Research (China Collaboration)work
Collaborated on a data-driven research project with a PostDoc researcher in China.

Data Scientist Intern

Kinetic Intelligencework
  • Scraped data for 24,000 US schools across 35 different sports teams.
  • Automated scraping workflows using GitHub Actions, eliminating the need for local VPNs and manual execution.
  • Implemented data caching to optimize future data retrieval and prevent redundant scraping.
  • Processed 24,000 school records
  • Automated 100% of the scraping execution via GitHub Actions
PythonWeb ScrapingGitHub Actions

Skills

Technical

PythonPython
Pandas
NumPy
Exploratory Data Analysis (EDA)
Data Cleaning
Jupyter Notebook
Apache Airflow
ETL Pipelines
Data Warehousing
Feature Engineering
Data Visualization
GitGit
GitHubGitHub
Google Colab
SQLSQL
PySpark
Snowflake
Scikit-learn
LangChain
Retrieval-Augmented Generation (RAG)
Matplotlib
Statistical Analysis
MySQLMySQL
Medallion Architecture
Classification
Predictive Modeling
Model Evaluation
DockerDocker
PyTorchPyTorch
AI Agents

Areas of expertise

Continuous Learning
Team Collaboration
Time Management
Adaptability
Analytical Thinking
Critical Thinking
Technical Documentation
Research

Projects

AI Student Assistant using Retrieval-Augmented Generation (RAG)

academic
Developed an intelligent university assistant using Retrieval-Augmented Generation (RAG) to answer questions related to admissions, academics, and university policies. Implemented document retrieval, vector search, and LangChain to generate context-aware responses grounded in official university documents. Improved response accuracy by retrieving relevant context before LLM inference, reducing hallucinations compared to relying solely on the language model's internal knowledge.
PythonLangChainRAGVector Database

Customer Churn Prediction using Machine Learning

personal
Built a Customer Churn Prediction model using transactional retail data collected from ECS (Ehsan Chappal Store). Performed data preprocessing, feature engineering, Exploratory Data Analysis (EDA), and model evaluation to identify customers likely to stop purchasing. Generated actionable business insights to support customer retention strategies, targeted marketing campaigns, and data-driven business decisions.
PythonPandasScikit-learnSQL

Content-Based Movie Recommendation Engine

academic
Developed a Content-Based Recommendation System using Natural Language Processing (NLP) techniques to generate personalized movie recommendations. Applied Bag-of-Words, Count Vectorization, and Cosine Similarity to compute similarity scores between movies. Built a lightweight recommendation engine capable of suggesting relevant content based on user preferences.
PythonNLPScikit-learn

Lung Cancer Detection using EfficientNet

personal
Developed a deep learning-based image classification pipeline using EfficientNet to detect lung cancer from CT scan images. Applied image preprocessing, transfer learning, feature extraction, and model optimization techniques to improve classification performance. Evaluated model performance using Accuracy, Precision, Recall, F1-Score, and Confusion Matrix for medical image classification.
PythonTensorFlowEfficientNetComputer Vision

Modern Data Warehouse & Analytics Pipeline

personal
Designed and implemented an end-to-end ETL Pipeline using PySpark to ingest, transform, and load data into Snowflake. Built an Apache Airflow DAG to automate workflow orchestration, enabling scheduled, reliable, and repeatable data processing. Implemented the Medallion Architecture (Bronze, Silver, and Gold layers) to progressively refine raw data into analytics-ready datasets. Designed analytics-ready Gold Layer warehouse tables optimized for Power BI dashboards and Business Intelligence (BI) reporting. Containerized the complete data pipeline using Docker to ensure reproducible deployment and environment consistency.
SnowflakePySparkApache AirflowSQLDockerPower BI

Media

Education

Bachelor of Science in Data Science

Riphah International University, LahoreData ScienceGPA 3.92027

Recognition

Awards

Hackathon Winner

Superior University, Lahore

Winner of an AI hackathon at Superior University, Lahore

Hire ZeeshanReach out about a role, a contract or a conversation.For recruiters

Zeeshan's twin is AI, it can make mistakes.