SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 139,725 - 209,588 / annual
OneTrust is seeking a Staff Software Engineer to join the Reporting and Data Platform team as a hands-on individual contributor. This role focuses on designing, building, operating, and improving distributed backend services and data-processing platforms that power reporting and data-driven experiences across the company's AI-Ready Governance Platform.
You will work across Java microservices and Python/PySpark data pipelines, with a strong emphasis on reliability, scalability, performance, and observability. You will own complex features and technical improvements from discovery through production rollout, making substantial hands-on contributions to backend services and data-processing pipelines.
Key responsibilities include:
**Technical Ownership & Delivery**: Own complex features from discovery through production, investigate ambiguous problems to identify root causes, review code and technical designs, and apply AI-assisted engineering tools (Devin, Claude, etc.) to accelerate delivery while maintaining production quality.
**Backend & Distributed Systems**: Design and implement production services using Java, Spring Boot, and Maven; develop event-driven functionality using Kafka; improve service performance, scalability, and fault tolerance; and diagnose issues across services, queues, databases, and dependencies.
**Data Engineering**: Build and maintain ingestion and transformation pipelines using Python, PySpark, Azure Databricks, and Delta Lake for batch and streaming workloads. Implement schema evolution, checkpoint management, deduplication, and data-quality controls. Optimize Spark joins, partitioning, and cluster utilization while protecting tenant boundaries.
**Observability & Operations**: Improve observability across services and pipelines using metrics, structured logs, traces, and business telemetry. Build actionable dashboards and monitors using Datadog and Grafana. Participate in on-call rotation and incident response, reducing alert noise and operational toil through engineering improvements.
Success in this role means requiring limited direction after understanding desired outcomes, owning work through design, implementation, testing, deployment, and ongoing operation, using production evidence to prioritize improvements, and making sound trade-offs among delivery speed, reliability, performance, security, cost, and maintainability.
OneTrust is embracing an office-first culture, encouraging three days per week in office for most roles, with flexibility depending on the specific position scope.
**Requirements**
Required Experience:
- Strong professional experience building and operating production software systems as a highly autonomous individual contributor
- Strong proficiency in Java and Spring Boot, with experience designing and operating distributed systems and microservices
- Production experience with asynchronous or event-driven systems, preferably Apache Kafka
- Strong experience with Python, PySpark, Apache Spark, and Delta Lake, plus production experience with Azure Databricks or comparable managed Spark platform
- Hands-on experience with test-driven development, automated testing strategies, and quality gates supporting fast, reliable delivery
- Strong understanding of metrics, logs, distributed tracing, dashboards, monitoring, and alerting, including hands-on experience with Datadog and Grafana
- Experience creating or responding to PagerDuty incidents, Datadog alerts, or equivalent production alerting workflows, and willingness to participate in on-call rotation
- Experience using AI engineering tools such as Devin, Claude, or similar systems to produce production-ready code, tests, documentation, and operational improvements
- Strong design-thinking skills and ability to reduce delivery-cycle time through clear architecture, smaller increments, reusable patterns, and pragmatic technical trade-offs
- Ability to independently diagnose complex performance and reliability problems and communicate implementation decisions and technical trade-offs clearly
Preferred Experience:
- Experience with both batch and streaming data pipelines and optimizing Spark or Databricks workloads for performance, reliability, and cost
- Experience with Databricks SQL, Databricks SDKs, Delta operations, and schema migrations
- Familiarity with Azure Blob Storage, Azure Identity, and Azure Key Vault
- Experience operating reporting, analytics, dashboard, or large-scale export systems
- Experience with Kubernetes, containers, CI/CD, and infrastructure as code
- Experience defining or applying service-level indicators, service-level objectives, and error budgets
- Understanding of data governance, encryption, audit-ability, and tenant isolation
- Experience modernizing established production systems incrementally