SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
SingleStore is seeking a Software Engineer to join the Observability Team, playing a critical role in designing and delivering core capabilities for SingleStore's observability platform. This is a hands-on engineering role with end-to-end ownership of features and projects at the intersection of distributed systems, cloud infrastructure, database technology, and AI-powered observability.
You will solve complex system-level problems, contribute meaningfully to technical direction, and partner with Product and customer-facing teams to ensure the platform meets the needs of enterprise customers. The observability team provides comprehensive customer observability over their workloads—enabling customers to understand usage patterns, identify bottlenecks, optimize performance, and leverage powerful alerting capabilities. The platform unifies traces, logs, and metrics through an open-source-first approach.
Key responsibilities include:
- Design and implement scalable observability features for traces, logs, and metrics spanning ingestion, processing, storage, and visualization
- Work across control plane and data plane components in a multi-cloud environment (AWS, GCP, Azure), ensuring reliable operation and data consistency at scale
- Build high-throughput data pipelines that process telemetry data using OpenTelemetry Collector and related open-source tooling
- Develop and maintain alerting capabilities with Alertmanager, enabling customers to define, tune, and manage alerts with routing, inhibition, and notification management
- Optimize time-series data storage and query performance using SingleStore DB, handling high-cardinality data and complex analytical queries
- Contribute to data visualization dashboards in Grafana, creating intuitive customer-facing experiences to explore telemetry data
- Collaborate closely with Product Management to translate customer and business requirements into robust technical solutions
- Investigate and resolve difficult issues in production and development environments, debugging data synchronization across distributed systems and cloud providers
- Participate in on-call rotations to ensure system reliability and respond to incidents promptly
SingleStore is a venture-backed, cloud-native database company headquartered in San Francisco with offices in Sunnyvale, Raleigh, Seattle, Boston, London, Lisbon, Bangalore, Dublin, and Kyiv. The company delivers a distributed SQL database that unifies transactions and analytics, empowering digital leaders to deliver exceptional, real-time data experiences.
REQUIREMENTS:
- 2+ years of professional software development experience building distributed systems or backend services
- Strong proficiency in Go (Golang); experience with Rust, Python, or C++ also valuable
- Deep understanding of distributed systems concepts: scalability, consistency, high availability, concurrency, and failure modes
- Familiarity with distributed systems managed via Kubernetes
- Demonstrated ability to design and build reliable, high-performance system software
- Experience working in environments where performance, scalability, and reliability are critical
- Familiarity with observability concepts: traces, logs, metrics, APM, and monitoring patterns
- Strong problem-solving and debugging skills with ability to root-cause complex production issues
- Excellent communication skills, both written and verbal, with ability to collaborate in multicultural, remote-first teams
- Code quality mindset: value simplicity, performance, maintainability, and thorough testing
PREFERRED QUALIFICATIONS:
- Experience with time-series data and understanding of metrics cardinality challenges
- Proficiency with SQL and experience working with relational or distributed databases
- Experience building cloud-native SaaS platforms with multi-tenant architecture
- Multi-cloud experience: working with AWS, GCP, Azure, or other cloud providers in production
- Kubernetes proficiency: operating, monitoring, or developing for Kubernetes clusters
- Open-source observability tools: hands-on experience with Grafana, Alertmanager, Loki, Tempo, OpenTelemetry Collector, OTLP protocol
- OpenTelemetry expertise: experience with or active contributions to OTel projects
- Time-series database experience: Prometheus TSDB, InfluxDB, Mimir, TimescaleDB, or SingleStore
- Experience with data pipeline technologies: Apache Kafka, Parquet, Arrow, or stream processing frameworks (Flink, etc.)
- Experience working with AI agents or LLM-powered applications, including agentic workflows for observability
- Experience in a SaaS or cloud-native company delivering managed services to customers