Junior Data Scientist & Analyst
I build end-to-end ML pipelines that turn messy real-world data into decisions businesses can act on.
Final-year Computer Engineering student at Usman Institute of Technology, Karachi.
I completed a structured 90-day data science roadmap,not by watching tutorials, but by building real projects at every stage. From EDA and SQL to full ML pipelines with GridSearchCV and feature engineering.
I believe the best data scientists don't just write code, they tell stories with data and translate numbers into decisions that matter.
Built an end-to-end NLP system to detect fraudulent job postings, combining text analysis with structured data features across 18,000+ postings with a severe 4.84% fraud class imbalance. Deployed the trained pipeline as a live API using FastAPI and Docker, and built an interactive Power BI dashboard translating fraud patterns into business insights for HR and recruitment teams.
Segmented 4,338 e-commerce customers into behavioral groups using RFM (Recency, Frequency, Monetary) analysis on raw UK retail transaction data. Engineered customer-level features from 500K+ line-item transactions, addressed heavy right-skew in spend/frequency data with log transformation, and validated cluster count using both elbow method and silhouette score before settling on K=3.
Built a complete ML pipeline to predict which telecom customers will cancel their subscription. Handled class imbalance, compared two models, and tuned hyperparameters automatically.
Analyzed real transactional e-commerce data across 5 relational tables to answer core business questions: top customers, top products, monthly retention, and repeat-purchase value. Uncovered and fixed a critical data bug where the dataset's order-level customer ID (not a true per-person identifier) was silently breaking retention analysis, rebuilt the logic around the correct unique-customer identifier to reveal Olist's real, low repeat-purchase pattern.
Built a diagnostic classifier to predict heart disease from patient medical data, prioritizing recall over raw accuracy since a missed diagnosis is far costlier than a false alarm. Compared Random Forest against Decision Tree using learning curve analysis to catch overfitting before trusting either model's numbers, and used feature-level boxplot analysis to identify which clinical markers actually separated disease from no-disease cases.
Built baseline regression model with feature engineering, outlier detection, and residual diagnostics. Identified key price drivers through coefficient analysis.
Complete EDA on 1,000 student records. Applied statistical hypothesis testing to prove test prep impact with scientific rigor.
Open to junior data science roles, freelance projects, and collaborations.