SlipstreamJobsFresh Startup & VC-Backed Jobs

DevOps & AI/ML Infrastructure Engineer

CreatorIQ - New York, NY, United States - Hybrid - posted 2026-09-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

CreatorIQ is seeking a DevOps & AI/ML Infrastructure Engineer to support and improve cloud and ML/AI infrastructure, automate deployments, and maintain CI/CD pipelines. The role ensures efficient, secure, and scalable development workflows while collaborating with Software Engineers, ML Engineers, Product Support, QA, and Security teams. Key Responsibilities: Cloud Infrastructure, Security & Reliability: Support and maintain scalable, highly available, and secure cloud infrastructure. Provision and manage cloud resources using Infrastructure as Code (Terraform, Terragrunt, CloudFormation). Implement cloud security best practices including IAM/role-based access controls, encryption, vulnerability management, and secure configurations. Support containerized environments and orchestration platforms. Apply DevSecOps principles and participate in disaster recovery planning and testing. CI/CD, Automation & Deployment: Maintain and optimize CI/CD pipelines using GitLab CI/CD and Jenkins, supporting application and ML model deployments. Improve deployment reliability and support zero-downtime deployment strategies. Automate configuration management, infrastructure provisioning, and routine operational processes. Troubleshoot deployment and pipeline issues and develop scripts to reduce manual work. AI & Agentic Infrastructure: Help design, deploy, operate, and secure infrastructure supporting AI and agentic products, including MCP, agents, integrations, internal tooling, and customer-facing use cases. Use AI-assisted engineering tools and coding copilots to improve DevOps productivity. Evaluate and adopt practical AI-enabled workflows for infrastructure management and troubleshooting. MLOps & ML Platform Infrastructure: Operate and scale ML platform infrastructure, including Databricks interactive clusters, jobs compute, ML pipelines, and Model Serving endpoints. Manage production model-serving infrastructure for high-throughput inference workloads. Maintain infrastructure-level monitoring for model drift, data quality, inference performance, and serving health. Partner with ML Engineering on reliable CI/CD and production deployment of ML models. Observability, Incident Response & Collaboration: Maintain monitoring, logging, metrics, and alerting solutions using Prometheus, Grafana, Coralogix, and CloudWatch. Support incident response and perform Root Cause Analysis for infrastructure issues. Improve system observability through effective log aggregation and metrics collection. Partner with engineering teams to improve deployment workflows and integrate automated testing into CI/CD pipelines. Collaborate with IT Security and respond to engineering and Product Support requests. Maintain accurate technical documentation and work effectively across international time zones. Requirements: - 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or similar infrastructure-focused role - 2+ years of hands-on experience with AWS services (EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar) - 2+ years of experience with containerized environments and orchestration platforms (Kubernetes, Amazon EKS) - Strong experience building and maintaining CI/CD pipelines using GitLab CI/CD or Jenkins - Hands-on experience with Infrastructure as Code (Terraform, Terragrunt, CloudFormation, or similar) - Strong Linux system administration and troubleshooting skills - Solid understanding of networking fundamentals (routing, load balancing, network security) - Scripting experience with Python, Bash, or similar languages - Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases - Experience supporting data, ML, or other compute-intensive production workloads Valuable additions: - Google Cloud experience, particularly multi-cloud environments - Familiarity with Helm and service mesh technologies (Istio, Linkerd, Traefik) - Experience with serverless and event-driven architectures (AWS Lambda, API Gateway, SQS) - Cloud and infrastructure security practices (vulnerability management, Nessus, Prowler, Trivy) - Knowledge of security standards, compliance requirements, and cloud security best practices - Experience with observability and monitoring platforms (Coralogix, Prometheus, Grafana) - FinOps experience (cloud cost monitoring and optimization) - Experience with API gateways or API management platforms (Kong, Apigee) - MLOps platforms and practices (Databricks, model serving, ML pipelines, model monitoring)

Similar roles