Role
Job Description
About the role
We are seeking a Data Scientist to join our hybrid team in Bengaluru. You will focus on building robust predictive models and extracting insights from complex datasets. This role requires a strong foundation in statistical analysis and machine learning, with a specific emphasis on end-to-end model development from data exploration to deployment monitoring. You will work closely with engineering and product teams to solve business problems using data-driven approaches.
Responsibilities
Perform Exploratory Data Analysis (EDA) to understand data distributions, identify anomalies, and define key features for modeling.
Design and implement feature engineering pipelines to transform raw data into high-quality inputs for machine learning algorithms.
Build and train classification models using Logistic Regression, Decision Trees, Random Forest, and XGBoost.
Write efficient Python code using Pandas, NumPy, and Scikit-learn to process large and messy datasets.
Develop SQL queries to extract and aggregate data from relational databases for analysis and model training.
Evaluate model performance using statistical metrics and validate results against business objectives.
Implement basic model monitoring scripts to track data drift and model degradation in production environments.
Required skills
Proficiency in Python with extensive experience in Pandas, NumPy, and Scikit-learn.
Strong command of SQL for data extraction, manipulation, and analysis.
Solid understanding of machine learning algorithms, specifically Logistic Regression, Decision Trees, Random Forest, and XGBoost.
Experience with statistical hypothesis testing and model evaluation metrics (AUC, F1, Precision, Recall).
Ability to handle large, unstructured, or messy datasets with minimal data cleaning overhead.
Basic knowledge of MLOps concepts and model monitoring practices.
Nice to have
Experience with cloud platforms like AWS or GCP for data storage and compute.
Familiarity with version control systems like Git for code management.
Exposure to containerization tools like Docker for model deployment.
Knowledge of A/B testing frameworks for experiment design.
What success looks like
You deliver accurate predictive models that improve key business KPIs by measurable percentages.
Your feature engineering pipelines reduce model training time and improve data quality.
You establish reliable monitoring systems that detect model drift before it impacts business outcomes.
You collaborate effectively with engineers to integrate models into production systems with minimal friction.
Skills
What you bring
Must have
