SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
FreedomPay is a fast-growing commerce platform company serving major enterprises across retail, hospitality, gaming, sports, entertainment, foodservice, education, healthcare, and financial services. The company operates a unified, globally-scaled technology stack and maintains world-class security standards, including PCI validation for Point-to-Point Encryption with EMV.
The Platform Operations team is seeking a Senior Platform Operations Engineer to drive automation and AI-driven strategies that maintain operational readiness as the company scales. This role bridges platform operations, engineering, and infrastructure teams to identify automation opportunities, implement AI-driven tooling for faster issue detection and resolution, and continuously improve monitoring and operational processes.
Key Responsibilities:
- Identify monitoring and performance optimization opportunities across the platform; implement improvements and continuously optimize system monitoring.
- Apply AI and machine learning to platform operations, including anomaly detection, alert correlation, and noise reduction to detect and resolve issues faster.
- Conduct root cause analysis for platform incidents; implement and maintain monitoring tools; develop proactive strategies to prevent recurrence.
- Serve as an escalation point for complex technical issues related to platform and services performance and stability; manage problems and coordinate resolution efforts to minimize business impact.
- Build and maintain AI-assisted tooling and automation that reduces manual operational work and accelerates incident response.
- Document procedure improvements and maintain detailed SOPs and alert response documentation to support operational consistency and knowledge retention.
- Continuously build technical expertise in observability and AI-driven operations; act as a subject matter expert and maintain strong stakeholder relationships across the company.
Requirements:
- BS degree in Computer Science or equivalent relevant experience.
- 4+ years of hands-on experience in observability, monitoring, or site reliability roles within highly available, high-throughput, or transaction processing environments.
- Demonstrated experience applying AI or machine learning tools to operational problems such as anomaly detection, alert correlation, or automated remediation.
- Strong incident management track record with a history of driving issues to resolution and minimizing business impact.
- Excellent communication and organizational skills with strong ownership and service orientation.
- Deep proficiency with an enterprise APM or observability platform (Dynatrace, Datadog, New Relic, or comparable), including dashboard design, alerting strategy, and performance analysis.
- Hands-on experience with AIOps and ML-driven observability, including anomaly detection, alert correlation, and intelligent noise reduction.
- Experience with log management and analysis tools such as Splunk or similar platforms for signal correlation and root cause analysis.
- Experience with AI-assisted development, agent workflows, and LLM-based tooling (Claude, Codex, or similar) applied to incident response, runbook automation, and knowledge tooling.
- Proficiency in scripting and automation, including PowerShell and/or Python, to build observability tooling, automate alert response, and drive remediation.
- Experience using Azure Logic Apps, Azure Functions, or Azure AI Foundry to build AI-driven automation workflows.
- Working knowledge of monitoring modern containerized workloads.
- Administration of third-party tools with Configuration-As-Code.
Preferred Qualifications:
- Experience with payment processing systems or financial services.
- Knowledge of PCI policies and best practices.
- Experience building AI agent workflows, LLM-based tooling, or automation pipelines.
- Experience mentoring and sharing knowledge with team members.