~/aamir-iqbal/portfolio (main)
AVAILABLE FOR ROLES
stochastic_modelling.py

DATA SCIENTIST/AI ENGINEER

Data Scientist/AI Engineer with 4+ years of experience building and deploying intelligent solutions across machine learning, generative AI, natural language processing, and data science. Experienced in solving complex business problems through applied AI, automation, predictive modeling, and document intelligence, with a strong focus on turning research and experimentation into practical, scalable systems.

M.Tech in Data Science
GITAM University // 2019 – 2021
Quick Actions & Direct Connections
4+ Years
ML & Data Science
<5 Mins
20 Doc Legal Batch
90% Auto
Ticket Resolution
Core Production Stack
Claude API / OpenAILangGraph & LangChainRAG & PineconeDocument AIFastAPIPySpark & DatabricksXGBoostPyTorch
Certifications:DeepLearning.AI MLOpsAWS GenAI AgentsGoogle AI
MD Aamir Iqbal
MD AAMIR IQBAL
Data Science/AI Engineer
Remote / India•Contact me →
GALTON QUINCUNX
CLT: Binomial(10, 0.5)
SAMPLE N0
MEAN μ0.00
VAR σ²2.50
// 01.PUBLICATIONS & ARTICLES

Articles & Technical Writings

Research articles, technical whitepapers, and production architecture specifications.

Paper .pdfOct 06, 2026•LLM Agents & Reinforcement Learning•19 pages · Paper Analysis

RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning

A detailed analysis of RLEF and how reinforcement learning teaches code LLMs to condition future generations on execution feedback rather than simply generating another independent answer.

Paper .pdfOct 06, 2026•LLM Systems & Inference•7 pages · Technical Note

KV Cache: Understanding Prefill, Decode, and Attention Caching

A technical walkthrough of KV caching, explaining how K and V are stored across decoder blocks, why Q is computed only for the current token, and how prefill and decode differ.

Article .mdSep 29, 2021•MLOps & Systems•8 min read

ML Data Is A First Class Citizen in Production

Why ML code in production is just a drop in the ocean compared to dynamic data pipelines, schema skew, and covariate drift.

Article .mdJun 25, 2021•Computer Vision & NLP•8 min read

Hierarchical Co-Attention for Visual Question Answering

Jointly reasoning about visual attention ("where to look") and question attention ("which words to listen to") across word, phrase, and question hierarchies.

Article .mdJul 02, 2021•Data Quality & Labeling•6 min read

Why Data Definition Is Hard in Real-World ML

The hidden trap of label inconsistency in unstructured datasets, resolving annotator disagreement, and establishing Human Level Performance baselines.

Article .mdAug 20, 2021•Model Evaluation•7 min read

Why Low Average Error Isn't Good Enough

A 99.2% aggregate test accuracy can disguise complete failure on critical query cohorts, catastrophic edge cases, and high-value customer segments.

Paper .pdfAug 2025•Document AI & GenAI•14 pages · Research Whitepaper

Automated Legal Document Intelligence: US Court Case Law PDFs to Compliant XML

Engineered a production-scale document intelligence pipeline for Thomson Reuters converting complex U.S. legal PDFs into structured XML/JAXML with zero content loss.

Spec .docxDec 2024•Enterprise AI & Databases•8 pages · Technical Architecture Spec

Multi-Dialect Text-to-SQL Pipeline with Vectorized Schema Caching & Pruning

Architected an enterprise Text-to-SQL system reducing generation latency by 75% via dynamic schema pruning, dialect-aware routing, and vector caching.

Article .mdMar 2024•Mathematical Foundations•9 min read · Mathematical Derivation

The Mathematics of Transformers: Attention, Projections & Softmax Geometry

A rigorous mathematical deconstruction of Scaled Dot-Product Attention, Query-Key-Value projection spaces, and temperature scaling proofs.

// 02.CURRICULUM_VITAE

Curriculum Vitae & Experience

Download PDF
MD AAMIR IQBAL|Data Scientist & ML Engineer

Data Scientist with effective over 4+ years of past experience specializing in Machine Learning, Deep Learning, and Natural Language Processing. Proficient in developing and deploying advanced analytical models to drive business insights and optimize processes.

Technical Skills
Machine LearningGenerative AIRAGDocument AIPineconeLangChainLangGraphLangSmithFastAPIFlaskClaude APIOpenAI APIAWSPythonStatisticsDatabases (MySQL)PySparkDatabricks
Professional Work Experience

ALGOLEAP TECHNOLOGIES

ML EngineerClient: Thomson Reuters (AI Engineer)
MAY 2025 – PRESENT
  • Engineered a production-scale document intelligence pipeline for converting complex U.S. legal PDFs into structured XML/JAXML outputs using Claude Sonnet-based LLM workflows, achieving zero content loss and reducing processing time from 3 hours to under 5 minutes for 20-document batches.
  • Developed a document intelligence pipeline for extracting and mapping footnotes to in-text references from complex PDF documents using GPT-5.2, Python, and LLM-based extraction workflows.
  • Developed an AI-powered document transformation pipeline that processes tracked changes in Word documents, applies rule and LLM-driven content modifications.
Tech:Claude SonnetGPT-5.2XML / JAXMLDocument AIPythonFastAPI

PRATHAM SOFTWARE

Sr Data Scientist
JUL 2024 – APR 2025
  • AI Ticket Automation for Logistics: Architected an AI-driven ticketing workflow using LangGraph, FastAPI, and Azure OpenAI to automate complex task categorization, summarization, and resolution. Engineered Retrieval-Augmented Generation (RAG) pipelines with Pinecone for vector search, structuring the agent routing, tool-calling, and fallback handling to successfully auto-resolve 90% of tickets and reduce manual operational effort by 80%.
  • LLM Text-to-SQL Pipeline: Led development using LangChain and OpenAI to translate natural language into complex SQL for multi-dialect engines. Designed semantic layers and integrated Pinecone for accurate schema retrieval, optimizing query latency by 75%.
  • AI Agent for NDA Conflict Resolution: Developed an agent with GPT-4 and Azure ML Studio to analyze NDAs and flag conflicts; integrated with Microsoft Teams for seamless contract review within enterprise workflows.
Tech:LangGraphLangChainAzure OpenAIPineconeFastAPIText-to-SQL

LANDMARK GROUP

Trainee Data Scientist
DEC 2022 – MAR 2024
  • Analyzed business requirements and provided data-driven solutions for stakeholders across various MARKETPLACE.
  • O2O Propensity Model (GCC Region): Developed a model achieving a 50% improvement over traditional methods, effectively targeting 90% of customers in the top 3 deciles, with an event rate of 1.5% or less across 6 countries.
  • Acquisition Model (Homecentre UAE): Targeted the top 5% of customers, achieving 80% coverage and a lift of over 2 in top deciles.
  • MLOps Pipeline (Centrepoint UAE): Designed and deployed a Stage 1 pipeline to generate customer propensity scores, enabling efficient forecasting for targeted campaigns.
Tech:PythonXGBoostDatabricksPySparkMLflow

EUNIMART PVT LTD

Machine Learning Engineer
NOV 2021 – SEP 2022
  • AI Service for SEO: Built a service for keyword and volume prediction to optimize omnichannel e-commerce platforms.
  • Personalized Onboarding Journey: Developed a user onboarding system with 95% statistical significance, improving user flow and reducing friction points.
  • Image Blur Removal: Enhanced and deployed an image deblurring solution for e-commerce platforms, reducing model size by 40%.
Tech:PythonBERTStatsmodelsAutoencodersXGBoostPyTorchLSTM

iNEURON INTELLIGENCE

Computer Vision Intern
MAY 2021 – AUG 2021
  • Developed a Computer Vision solution for visually impaired individuals, employing Multimodal learning (CV+NLP) for "Visual Question Answering."
  • Implemented the Hierarchical Co-attention paper, enhancing the model capabilities and contributing to advancements in attention mechanisms.
Tech:PythonTensorFlowFlaskCV+NLPData Labelling

EXPOSYS DATA LABS

Data Science Intern
JUL 2020 – AUG 2020
  • Implemented clustering algorithms on mall data to identify customer segments, optimizing targeted marketing strategies.
Tech:K-Means ClusteringPythonScikit-Learn

THE SMARTBRIDGE

Summer Intern
APR 2020 – MAY 2020
  • Implemented a Random Forest model for "Quality Prediction in the Mining Process."
  • Applied machine learning techniques to analyze and predict the quality of mining outputs, contributing to process optimization.
Tech:Random ForestStatistical ModelingPython
Academic Credential
M.TECH IN DATA SCIENCE
GITAM UNIVERSITY
JUN 2019 – MAY 2021
// 03.CONTACT_ME

Initiate Contact

Available for Applied Machine Learning, MLOps, LLM Systems & Research Engineering discussions.