Available for Opportunities
Karachi, Pakistan

Hadia Siddiqui

Junior Data Scientist & Analyst

I build end-to-end ML pipelines that turn messy real-world data into decisions businesses can act on.

Scroll to explore

The person behind
the pipelines

Final-year Computer Engineering student at Usman Institute of Technology, Karachi.

I completed a structured 90-day data science roadmap,not by watching tutorials, but by building real projects at every stage. From EDA and SQL to full ML pipelines with GridSearchCV and feature engineering.

I believe the best data scientists don't just write code, they tell stories with data and translate numbers into decisions that matter.

Currently: Open to junior data science and analyst roles locally and remotely.
NameHadia Siddiqui
LocationKarachi, Pakistan
UniversityUsman Institute of Technology
DegreeB.E. Computer Engineering
Graduating2026
FocusML · NLP · Data Analysis
StatusOpen to Work

My toolkit

🐍 Languages
Python SQL
📦 Libraries
PandasNumPy Scikit-learn Matplotlib SeabornSciPy NLP/TF-IDF
🤖 Machine Learning
Linear RegressionLogistic Regression Decision TreesRandom Forest GridSearchCVCross Validation L1/L2 Regularization Class Imbalance K-Mean Clustering
⚙️ Data Engineering
sklearn PipelinesColumnTransformer Feature EngineeringEDA Hypothesis TestingOutlier Detection
🗄️ Databases
MySQLSQLite JoinsSubqueries Window Functions
📊 BI & Visualization
Power BIDAX Power Query
🛠️ Tools
VS CodeJupyter GitGitHubKaggle

What I've built

Fake Job Posting Detection project thumbnail
01 / PROJECT · NEW

Fake Job Posting Detection (NLP + Deployment)

Python · scikit-learn · TF-IDF · FastAPI · Docker · Power BI

Built an end-to-end NLP system to detect fraudulent job postings, combining text analysis with structured data features across 18,000+ postings with a severe 4.84% fraud class imbalance. Deployed the trained pipeline as a live API using FastAPI and Docker, and built an interactive Power BI dashboard translating fraud patterns into business insights for HR and recruitment teams.

🎯 63.4% Recall, 0.375 F1-Score on severely imbalanced data
📊 18,000+ job postings analyzed, EMSCAD dataset
⚙️ TF-IDF + Chi-Square feature selection, tuned Linear SVM via RandomizedSearchCV
Power BI dashboard overview
Overview / KPIs
Industry and region fraud breakdown
Industry / Region Breakdown
Missing info vs fraud correlation
Missing Info vs Fraud
→ View on GitHub
RFM Customer Segmentation
02 / PROJECT

RFM Customer Segmentation

Python · K-Means Clustering · Pandas · scikit-learn · (Power BI — coming soon)

Segmented 4,338 e-commerce customers into behavioral groups using RFM (Recency, Frequency, Monetary) analysis on raw UK retail transaction data. Engineered customer-level features from 500K+ line-item transactions, addressed heavy right-skew in spend/frequency data with log transformation, and validated cluster count using both elbow method and silhouette score before settling on K=3.

🎯 Silhouette score 0.416 at K=3 — clean separation across 3 customer segments
📊 4,338 customers segmented from UCI Online Retail II dataset (500K+ raw transactions)
⚙️ RFM feature engineering · log transform for skew · StandardScaler · K-Means · cluster profiling
→ View on GitHub
Customer Churn Prediction project thumbnail
03 / PROJECT

Customer Churn Prediction

Python · scikit-learn · Pipeline · GridSearchCV

Built a complete ML pipeline to predict which telecom customers will cancel their subscription. Handled class imbalance, compared two models, and tuned hyperparameters automatically.

🎯 90% Recall, up from 47% baseline
📊 7,000+ customer records analyzed
⚙️ GridSearchCV across 108 parameter combinations
→ View on GitHub
E-Commerce Customer Analytics project thumbnail
04 / PROJECT

E-Commerce Customer Analytics

MySQL · SQL · (Power BI — coming soon)

Analyzed real transactional e-commerce data across 5 relational tables to answer core business questions: top customers, top products, monthly retention, and repeat-purchase value. Uncovered and fixed a critical data bug where the dataset's order-level customer ID (not a true per-person identifier) was silently breaking retention analysis, rebuilt the logic around the correct unique-customer identifier to reveal Olist's real, low repeat-purchase pattern.

🎯 Repeat customers = 3.1% of the base but 5.9% of revenue ~2x higher per-capita value than one-time buyers
📊 ~96,000 customers · ~99,000 orders analyzed via JOINs, self-joins, and multi-level subqueries
⚙️ MySQL · data cleaning (encoding, type mismatches, invalid dates) · self-joins · CTEs
→ View on GitHub
Heart Disease Prediction project thumbnail
05 / PROJECT

Heart Disease Prediction

Python · Random Forest · Decision Tree · sklearn Pipelines · Learning Curves

Built a diagnostic classifier to predict heart disease from patient medical data, prioritizing recall over raw accuracy since a missed diagnosis is far costlier than a false alarm. Compared Random Forest against Decision Tree using learning curve analysis to catch overfitting before trusting either model's numbers, and used feature-level boxplot analysis to identify which clinical markers actually separated disease from no-disease cases.

🎯 84% accuracy (Random Forest) vs 78% (Decision Tree) Random Forest chosen for its smaller train/validation gap, not just the higher number
📊 UCI Heart Disease dataset · 13 features after dropping high-missingness/non-predictive columns
⚙️ Leak-safe sklearn Pipeline (fit preprocessing on train only) · learning curves for overfitting diagnosis · recall prioritized as the clinical metric that matters
→ View on GitHub
House Price Prediction project thumbnail
06 / PROJECT

House Price Prediction

Python · Linear Regression · Feature Engineering

Built baseline regression model with feature engineering, outlier detection, and residual diagnostics. Identified key price drivers through coefficient analysis.

🎯 R² = 0.64 — 64% variance explained
📊 545 properties analyzed
⚙️ IQR outlier detection + interaction features
→ View on GitHub
Student Performance Analysis project thumbnail
07 / PROJECT

Student Performance Analysis

Python · Pandas · Seaborn · SciPy

Complete EDA on 1,000 student records. Applied statistical hypothesis testing to prove test prep impact with scientific rigor.

🎯 23% score improvement proven (p < 0.05)
📊 1,000 students across 8 features
⚙️ t-test · correlation analysis · IQR outliers
→ View on GitHub

Where I learned
to think in data

Bachelor of Electrical Engineering
Computer Engineering
Usman Institute of Technology
Karachi, Pakistan
2022 — 2026
Relevant Coursework
Data Structures Algorithms Database Systems Statistics Linear Algebra Probability Theory

Let's Work Together

Open to junior data science roles, freelance projects, and collaborations.