SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Socure is building identity trust infrastructure for the digital economy, verifying identities in real time and stopping fraud before it starts. The Big Data R&D team is responsible for developing the core identity graph and entity-resolution capabilities that power Socure's Deceased Monitoring and compliance products.
In this Data Scientist II role, you will contribute to the design and implementation of machine learning, data mining, statistical, and graph-based algorithms to analyze very large PII datasets. You'll work on entity-resolution and identity-matching algorithms that drive Socure's compliance solutions, building and maintaining data-processing pipelines using Spark/PySpark and AWS (EMR, S3). You'll support senior data scientists with feature engineering, data exploration, error analysis, and A/B test setup for new models and signals.
Key responsibilities include evaluating new third-party and internal data sources by profiling data quality, designing offline experiments, and summarizing impact on coverage and model performance. You'll implement and maintain SQL and Python/R code for data extraction, transformation, and validation, contributing to code reviews and testing. You'll also provide analytical support to compliance and regulatory product teams through ad hoc investigations, dashboards, and data deep dives.
You'll communicate findings clearly to peers and cross-functional partners (Product, Engineering, Client Analysis), focusing on key insights and trade-offs. The role requires working effectively in a fast-paced, cross-functional environment with ownership of well-scoped tasks.
Required qualifications: Master's degree with 2+ years of data science/analytics experience, or Ph.D. with 1+ years, or equivalent practical experience. Proficiency in Python or Scala, solid SQL skills for large datasets, hands-on Spark/PySpark experience, and familiarity with common ML libraries (scikit-learn, XGBoost, TensorFlow/PyTorch). UNIX environment and AWS ecosystem experience required. Graph techniques or graph databases (Neo4j, AWS Neptune, GraphFrames) are a strong plus. Bonus skills include Elasticsearch, DynamoDB, and Airflow. Strong problem-solving ability and ability to iterate quickly with feedback.