SlipstreamJobsFresh Startup & VC-Backed Jobs

Principal Software Architect - Data Platform

SecurityScorecard - New York, NY, United States - Hybrid - posted 2026-09-11

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 270,000 - 330,000 / annual

SecurityScorecard is the global leader in cybersecurity ratings, continuously rating over 12 million companies across 64 countries. The company's patented technology is used by 25,000+ organizations for self-monitoring, third-party risk management, board reporting, and cyber insurance underwriting. You will serve as Principal Software Architect for the data platform, an individual contributor role reporting to the Chief Architect. This is a high-impact position where you own the system design of a mission-critical data platform that ingests internet-scale measurement data, processes it through streaming, microbatch, and batch paths, stores it affordably at scale, and serves analytics in real time. The data itself is the product—not a byproduct—which raises the stakes on correctness. Published ratings are disputed by companies and used by underwriters for pricing decisions, making data quality, contracts, and lineage architecture problems rather than administrative ones. You will work alongside a Principal Architect focused on AI/agentic systems and a Principal Front-end Architect, partnering closely with engineering leadership, Product, and Data Science. Your role emphasizes influence through technical credibility and clear reasoning rather than authority. You will prototype to validate architectural decisions, set direction through Technical Design Reviews (TDRs) and standards, and lead the data domain while bringing distributed systems judgment to review designs across the wider platform. Key responsibilities include: owning end-to-end system design from ingestion through serving layer; defining service boundaries and data contracts between producers and consumers; designing the lakehouse architecture (table format, partitioning, schema evolution, compaction, metadata management); architecting the analytical serving layer for three conflicting workload classes (low-latency product queries, ad-hoc internal analytics, bulk external feeds); setting direction on data stack languages and frameworks; engineering data quality and observability into the platform (validation, quarantine paths, freshness/completeness SLOs, drift detection, lineage); designing for correctness and reproducibility in the ratings pipeline including backfills and historical restatements; writing TDRs and design docs that set architecture direction; reviewing TDRs from across engineering; and mentoring senior and staff engineers on data system design. The platform runs on Kafka for event streaming, Flink (Java) for stream processing, Spark for batch/microbatch, with some high-performance components in C++. Storage uses ClickHouse for analytics, PostgreSQL, and AWS infrastructure with Kubernetes, Terraform, Helm, and ArgoCD. The broader platform uses Node.js, TypeScript, React, and is actively expanding AI/ML infrastructure. REQUIREMENTS: - 10+ years of software or data engineering experience, including significant time architecting large-scale data platforms - Deep expertise in stream and batch processing at scale with Kafka, Flink, and Spark or close equivalents; clear judgment about which path a given workload belongs in - Strong Python and PySpark; solid Java for Flink stream processing; enough Scala to read and reason about existing Spark codebases - Hands-on production experience designing lakehouse storage: columnar formats (Parquet), open table formats (Iceberg), and associated partitioning, compaction, and schema evolution decisions - Experience architecting OLAP and analytical serving layers (ClickHouse, Druid, Pinot, BigQuery, Snowflake or similar) - Strong distributed systems fundamentals as applied to data: exactly-once vs. at-least-once semantics, ordering, backpressure, late and out-of-order data, pipeline failure modes - Track record of building data quality, contracts, and observability as engineered system properties (assertions, schema enforcement, lineage in code) - Experience owning large-scale data migrations, including preserving history and correctness through cutover - Track record of influence without authority: presenting technical direction to skeptical engineers and earning genuine buy-in; giving rigorous design review feedback on systems you didn't build - Strong technical writing and mentorship: TDRs, design docs, decision records that teams can act on without hand-holding; history of raising technical bar around you - Comfort operating as a senior individual contributor, driving outcomes through prototyping and technical credibility PREFERRED: - Experience building data layers for ML/LLM systems (feature stores, vector stores, retrieval pipelines) - Internet-scale scan, telemetry, observability data experience, or cybersecurity industry background - Track record of reducing platform costs at scale through storage tiering, query governance, or compute right-sizing

Similar roles