WORK
AI/ML Engineer
Verdant Soft
Fine-tuned LLMs (GPT, BERT, Falcon) using LoRA and PEFT, improving contextual accuracy by 30%. Developed Retrieval-Augmented Generation (RAG) systems with FAISS and LangChain, supporting enterprise knowledge-grounded assistants. Built modular ML data pipelines for preprocessing, model training, evaluation, and inference optimization. Applied quantization (GPTQ, ONNX, Llama.cpp) and TensorRT acceleration to achieve 40% faster inference. Implemented RLHF (Reinforcement Learning from Human Feedback) for conversational model alignment. Collaborated with product managers to integrate ML logic into business workflows and create documentation for production models.
•Fine-tuned LLMs (GPT, BERT, Falcon) using LoRA and PEFT, improving contextual accuracy by 30%.
•Developed Retrieval-Augmented Generation (RAG) systems with FAISS and LangChain, supporting enterprise knowledge-grounded assistants.
•Built modular ML data pipelines for preprocessing, model training, evaluation, and inference optimization.
•Applied quantization (GPTQ, ONNX, Llama.cpp) and TensorRT acceleration to achieve 40% faster inference.
•Implemented RLHF (Reinforcement Learning from Human Feedback) for conversational model alignment.
•Collaborated with product managers to integrate ML logic into business workflows and create documentation for production models.
Key Achievements
→Built scalable ingestion and preprocessing pipelines for large text and multimodal datasets using PySpark on Databricks clusters, eliminating single-node memory bottlenecks.
→Parallelized schema validation, deduplication, and metadata normalization to ensure high-quality data for downstream LLM fine-tuning.
→Optimized the data lifecycle by staging raw payloads in S3 and utilizing Parquet and Delta formats for efficient access and storage.
→Improved contextual accuracy by 30% through LLM fine-tuning and achieved 40% faster inference using quantization and TensorRT acceleration.
→Improved contextual accuracy by 30% through LLM fine-tuning.
→Achieved 40% faster inference using quantization and TensorRT acceleration.
GPTBERTFalconLoRAPEFTRAGFAISSLangChainGPTQONNXLlama.cppTensorRTRLHFPythonAWSDatabricksPySparkS3Delta Lake