SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Software Engineer, Agent Platform

Commure - Rio de Janeiro, Brazil - In-office - posted 2026-09-15

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Commure is building an AI Operating System for healthcare, integrating ambient AI, intelligent agents, and autonomous RCM processing across 60+ EHRs. The company processes $25B+ in annual claims and supports 200M+ patient interactions for 500+ healthcare organizations. Recently raised $70M at a $7B valuation and named to Fortune Future 50 and 2026 AI Breakthrough Awards. You will own the platform powering Commure's fleet of OpenClaw agents—the infrastructure, runtime environments, and tools that enable autonomous agents to execute operational workflows, interact with internal systems, and turn organizational knowledge into action. This role sits at the intersection of platform engineering and applied AI. Key responsibilities include: - Own infrastructure and runtime for the OpenClaw agent fleet - Build, configure, deploy, and troubleshoot agents on Linux VMs - Develop Python services, automation, operational tooling, and agent integrations - Design agent harnesses supporting tool use, memory, context management, permissions, scheduling, retries, and human escalation - Create standardized platform for engineers to develop, test, deploy, and update agents safely - Deploy and operate workloads across Linux VMs, Kubernetes, and cloud infrastructure (Google Cloud or AWS) - Use Bash and Linux tooling to investigate system behavior and resolve production issues - Build reliable systems for agent scheduling, background execution, queues, and concurrency management - Develop integrations enabling agents to interact with internal services, databases, APIs, and operational tools - Establish secure approaches to secrets, credentials, identity, and access control for autonomous agents - Build evaluation frameworks measuring agent reliability, task completion, tool accuracy, latency, and cost - Implement production observability across agent traces, prompts, tool calls, infrastructure, and resource consumption - Create dashboards, alerting, and incident-response processes for agent and platform health - Diagnose failures across application code, model behavior, third-party tools, networks, containers, VMs, and cloud infrastructure - Improve platform efficiency and resilience as agent usage grows - Work closely with product engineers and internal teams to convert agent prototypes into production systems - Establish best practices for developing and operating autonomous agents Within first several months, you will develop deep understanding of existing agents and infrastructure, make the fleet easier to deploy and maintain, establish reliability metrics, reduce prototype-to-production time, improve safeguards around credentials and permissions, and create a platform enabling teams to build increasingly capable agents. REQUIREMENTS: - Professional experience building and operating production software, infrastructure, or internal platforms - Strong proficiency in Python (primary language for agent development, platform tooling, integrations, and automation) - Excellent familiarity with Linux environments and hands-on experience developing, deploying, and debugging on VMs - Strong Bash and command-line skills (shell scripting, process management, networking diagnostics, filesystem operations, package management, log investigation) - Experience managing software across fleets of VMs (configuration, deployment, monitoring, access management, failure recovery) - Strong hands-on experience with Google Cloud Platform or Amazon Web Services (compute, networking, storage, IAM, observability) - Experience deploying and managing Linux VMs and containerized workloads in cloud environments - Hands-on experience with Kubernetes, Docker, and containerized production workloads - Experience building infrastructure, deployment systems, internal developer platforms, or distributed application runtimes - Experience with CI/CD, infrastructure as code, configuration management, secrets management, and automated deployments - Hands-on experience building with large language models, autonomous agents, or agent harnesses - Understanding of agent concepts (tool calling, memory, context management, planning, retries, evaluations, permissions, human-in-the-loop execution) - Experience monitoring and debugging distributed systems using logs, metrics, traces, dashboards, and alerts - Strong understanding of production reliability (failure recovery, idempotency, rate limiting, concurrency management, incident response) - Ability to balance fast experimentation with security and reliability required for production systems - Strong ownership and communication skills, comfort working across product, infrastructure, security, and operational teams - Builder mindset and excitement about creating infrastructure for rapidly evolving agent software category NICE TO HAVE: - Built, extended, deployed, or operated an OpenClaw agent - Contributed to OpenClaw or another open-source agent framework - Experience operating a fleet of autonomous or long-running AI agents - Experience building multi-tenant agent platforms or internal developer platforms - Experience with agent evaluation, tracing, replay, simulation, or regression testing - Experience with browser agents, computer-use agents, or agents interacting with external systems - Familiarity with model gateways, inference providers, prompt versioning, token accounting, and LLM cost optimization - Experience securing workloads accessing sensitive data or high-impact production tools - Familiarity with Swift and experience with native Apple client applications - Experience in healthcare, revenue cycle management, or regulated industries

Similar roles