SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer III – Data Platform

ID.me - Mountain View, CA, United States - In-office - posted 2026-08-05

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ID.me is seeking a Site Reliability Engineer to join the Core Platform Engineering organization. The SRE team builds automation, observability, and operational foundations that ensure ID.me's services are reliable, scalable, and secure. As an SRE, you will play a pivotal role in building the platform and governance processes required to safely scale, deploy, and operate a high volume of machine-generated applications and features. You will design and implement automated guardrails that maintain high standards for resilience and security in an AI-accelerated development environment. Key responsibilities include: - Build and maintain automated reliability tooling, infrastructure as code, and observability systems that enhance uptime and service performance - Develop monitoring, logging, and alerting frameworks (e.g., Prometheus, Grafana, OpenTelemetry) to detect and remediate issues proactively - Implement automated architectural reviews and reliability guardrails for agent-developed applications to ensure machine-generated code meets long-term maintainability and performance standards - Partner with engineering teams to design and implement scalable, fault-tolerant systems that meet defined SLIs and SLOs - Automate repetitive operational tasks and develop self-healing and auto-remediation mechanisms to minimize human intervention - Participate in on-call rotations and lead incident response efforts, performing post-incident reviews and driving systemic improvements - Improve the deployment and release process using CI/CD pipelines and progressive delivery techniques - Champion observability, reliability, and operational readiness reviews as part of the development process - Collaborate with Security and Compliance teams to ensure production systems meet FedRAMP, NIST, and internal policy requirements - Contribute to documentation, runbooks, and internal tooling to enhance knowledge sharing and operational maturity Minimum qualifications: Bachelor's degree in Computer Science, Software Engineering, or related technical field; 3-5 years of experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering; 2+ years hands-on experience managing and scaling services in cloud environments (AWS, GCP, Azure); 1+ years proficiency in at least one modern programming language (Java, Go, Python, Ruby, JavaScript). This role is based in Mountain View, CA or McLean, VA and requires full-time in-office attendance, 5 days per week.

Similar roles