Amara Nasir

Amara Nasir

Machine Learning Engineer specializing in Generative AI & LLM Orchestration

Based in: Redwood City, CACurrently: Machine Learning Engineer, SQUARE63Experience: 7 years

Hire me

Machine Learning Engineer with over 5 years of experience building production-grade Generative AI systems and fraud detection models across fintech and document intelligence sectors. Expert in orchestrating RAG pipelines on Azure Databricks and developing high-throughput asynchronous data flows for large-scale enterprise applications.

Experience

Machine Learning Engineer

SQUARE63work2026 – Present

Summary

Design, build, and deploy end-to-end production LLM/GenAI pipelines on Azure Databricks for large-scale document understanding, data extraction, and classification. Orchestrate Databricks workloads with Azure Data Factory: parameterized, schedule- and event-triggered pipelines that pass runtime parameters into notebooks for repeatable batch runs across environments. Architect multi-model LLM workflows that route each stage to the best-fit model (lightweight models for high-volume classification, stronger reasoning and vision models for complex parsing), balancing accuracy, latency, and cost. Engineer high-throughput asynchronous pipelines (asyncio, semaphores, rate-limiting) with structured-output validation (Pydantic) and automatic retries for schema-conformant results at scale. Build scalable PySpark data flows that read from and write to relational databases over JDBC, applying cleaning, normalization, and fuzzy-matching across large datasets. Integrate Azure Data Lake Storage (ADLS Gen2) for fault-tolerant concurrent I/O and manage credentials securely through Azure Key Vault. Develop ground-truth evaluation frameworks (accuracy metrics, confusion matrices, confidence-threshold analysis) to benchmark models and continuously refine prompts and routing logic.

What I did

  • Built document extraction and classification pipelines on Azure Databricks for scheduled batch processing of large document sets.
  • Build and run production pipelines on Azure Databricks including Data Factory orchestration, PySpark data flows, ADLS Gen2 storage, and Key Vault for credentials.
  • Created a labeled ground-truth set to benchmark prompt and routing changes before production deployment.

Results

  • Reduced batch processing time by switching LLM calls to asynchronous execution.
  • Maintained accuracy while reducing costs by implementing a document routing system that assigns tasks to models based on complexity.
  • Established a ground-truth evaluation set to measure accuracy for all prompt and routing changes before production deployment.
  • Designed and built an evaluation framework that reports accuracy, confusion matrices, and confidence-threshold analysis for LLM pipelines.
LLMGenAIAzure DatabricksAzure Data FactoryasyncioPydanticPySparkJDBCADLS Gen2Azure Key VaultRAGAzure OpenAIEvaluation FrameworksConfusion MatricesConfidence-threshold Analysis

Data Scientist / Machine Learning Engineer

i2c IncRedwood City, CAwork2023 – 2026

Summary

Built and deployed ML and deep-learning fraud-detection models over millions of card transactions, improving model performance by 50% and increasing fraud detection by 6%. Engineered statistical and behavioral features from transaction records (spending habits, card types, and demographic segments) to surface fraudulent activity. Developed end-to-end ML pipelines in Python (Gradient Boosting Trees, LightGBM, Random Forest, Decision Trees), addressing class imbalance and feature scaling with cross-validation for robustness. Tuned hyperparameters and applied advanced algorithms to raise precision and reduce false positives. Conducted exploratory data analysis across card programs, merchants, and cardholders to uncover fraud patterns and translate hypotheses into production features. Delivered a data-driven proof of concept for a credit-limit reassessment system through feature engineering and model selection. Partnered with cross-functional teams to integrate models into existing fraud-detection systems for reliable, scalable deployment.

What I did

  • Built and deployed fraud-detection models on real transaction data.
  • Built a proof of concept for a credit-limit reassessment system using existing transaction data and behavioral features to justify a full build.
PythonGradient Boosting TreesLightGBMRandom ForestDecision TreesDeep LearningStatistical Modeling

Software Engineer

Commit LabsFrisco, TXwork2020 – 2023
Led end-to-end development of 3 web applications, building Python services with modern front-end technologies. Implemented RESTful APIs in Python and FastAPI for seamless data exchange between back-end and front-end layers. Built responsive interfaces backed by Python and FastAPI, reducing bounce rates by 20%. Contributed to architectural decisions ensuring cohesion between front-end and back-end components. Cut annual software maintenance costs by 15% through resource-efficient allocation.
PythonFastAPIRESTful APIsJavaScriptHTML5CSS3

Skills

Technical

Machine Learning
Azure Databricks
Feature Engineering
LLM/RAG Pipelines
PythonPython
Generative AI
GitGit
Predictive Modeling
Prompt Engineering
Statistical Modeling
SQLSQL
Jupyter
Azure OpenAI
scikit-learn
NumPy
Pandas
PySpark
FastAPIFastAPI
Azure Data Factory
Deep Learning
NLP
ADLS Gen2
JavaScriptJavaScript
HTML5/CSS3HTML5/CSS3
Neural Networks
PyTorchPyTorch
TensorFlowTensorFlow
Azure Key Vault
MySQLMySQL
PostgreSQLPostgreSQL
Visual Studio
C/C++C/C++
Keras
Power BI
AWS (S3, EC2, Lambda)
PHPPHP
Pydantic
Independent project creation
LangGraph
Vision Models
Data Normalization
LangChain
Rapid prototyping

Projects

Receipt Extraction Pipeline

personal
A pipeline built to measure vision LLM performance on structured extraction using the CORD benchmark. It features an evaluation harness for field-level accuracy and a routing layer using LangGraph to send failed validations to stronger models or human review.
OpenAI vision modelLangChainPydanticLangGraphHugging Face

Public Receipt Extraction Project

personal
A document processing project using model routing and async calls to extract data from receipts.
OpenAI vision modelLangChainPydanticLangGraphHugging Face

Education

MS

National University of Computer & Emerging SciencesData Science2019 – 2022

Coursework

Big DataMachine LearningDeep LearningNatural Language ProcessingComputer VisionComputational IntelligenceStatistics & Probability

BS

National University of Computer & Emerging SciencesComputer Engineering2004 – 2008

Coursework

Design & Analysis of AlgorithmsDatabase SystemsData StructuresObject-Oriented ProgrammingIntroduction to Computing
Hire AmaraReach out about a role, a contract or a conversation.For recruiters

Amara's twin is AI, it can make mistakes.