About the role
Data Scientist Job requirements Experience Range: With at least 4 years of hands-on experience in advanced data science, including statistical analysis and machine learning, and up to 6 years in similar roles Key Responsibilities: - Design and implement robust statistical models using advanced hypothesis testing, regression, and forecasting techniques to deliver actionable business insights - Develop and optimize machine learning algorithms for classification, prediction, and probabilistic graph models utilizing Python, PySpark, and R - Conduct comprehensive statistical analysis with SAS, SPSS, and R Studio to support data-driven decision-making - Build, train, and deploy scalable models using ML frameworks such as TensorFlow, PyTorch, Sci-Kit Learn, CNTK, Keras, and MXNet - Apply advanced time series forecasting methods, including exponential smoothing, ARIMA, and ARIMAX, to analyze trends and predict outcomes - Streamline model deployment and lifecycle management in production environments using KubeFlow and BentoML - Implement and validate data quality checks with Great Expectations and Evidently AI to ensure dataset integrity - Present complex data findings to stakeholders, translating insights into actionable recommendations that drive business outcomes Required Skills: - Advanced application of hypothesis testing methodologies, including T-Test and Z-Test - Expert-level regression analysis (linear and logistic) for predictive modeling - Proficient programming in Python and PySpark for data manipulation and model development - Extensive experience with statistical analysis using SAS and SPSS - Hands-on expertise in probabilistic graph models for complex data relationships - Mastery of time series forecasting techniques (exponential smoothing, ARIMA, ARIMAX) - Implementation of classification algorithms such as decision trees and support vector machines (SVM) - Deep familiarity with ML frameworks: TensorFlow, PyTorch, Sci-Kit Learn, CNTK, Keras, MXNet - Calculation and application of distance metrics (Hamming, Euclidean, Manhattan) - Skilled in R and R Studio for statistical analysis and visualization Preferred Skills: - Practical experience with Great Expectations and Evidently AI for advanced data validation - Proficiency in cloud-based model deployment tools such as KubeFlow and BentoML - Background in large-scale data processing and distributed computing environments - Expertise in feature engineering and model interpretability techniques - Familiarity with cloud-based data science platforms such as AWS SageMaker, Azure ML, or Google Cloud AI Platform Desired Qualifications: - Bachelor's degree in Computer Science, Statistics, Mathematics, Data Science, or a closely related discipline - Certification in Data Science or Machine Learning from a recognized institution, such as Microsoft Certified: Azure Data Scientist Associate or TensorFlow Developer Certificate