SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 313,000 - 369,000 / annual
Horizon3.ai is a fast-growing cybersecurity company building autonomous pentesting and defensive systems. The NodeZero platform delivers production-safe autonomous pentests and security assessments across enterprise environments, used by ITOps/SecOps teams, consultants, and managed security providers.
You will build AI agents that reason from proven attack paths to specific control changes that remediate them—EDR policies, firewall rules, conditional access, detection content, cloud IAM, GPO—and autonomously verify the fix by re-running the attack. This is a closed-loop defense system: find, fix, verify, at enterprise scale.
The core challenge is that your agents run inside customer production environments and modify live security controls. A bad change creates an outage or a new vulnerability. The reasoning problem and the safety problem are identical. You will own both, with real outcome signals on short loops and hundreds of thousands of prior tests to learn from.
Key responsibilities:
- Build reasoning systems that map attack paths and exploitation telemetry to specific, ranked control changes, balancing effectiveness against operational blast radius.
- Transform pentest data into training and evaluation datasets; extract signal from why attacks succeed in some environments and fail in others.
- Design and run counterfactual experiments in test environments to validate whether changes break attack chains, measure operational cost, and assess generalization.
- Design the reasoning layer over heterogeneous control planes so agents work across vendor APIs with different policy models without hardcoded playbooks.
- Design the safety architecture for autonomous change: dry-run, simulation, blast-radius classification, approval gates, staged rollout, rollback, and audit trails.
- Collaborate with attack and detection engineers to define target behavior and diagnose failure modes: ineffective remediations, over-broad changes, business-breaking edits, recommendations that don't hold on re-test.
- Own problems end-to-end in a 0→1 environment where requirements are ambiguous, systems move fast, and reliability is critical.
Requirements:
- Strong ML engineering experience building, evaluating, and deploying production AI systems with hands-on deep learning, transformer models, and PyTorch.
- Hands-on experience with post-training large language models (supervised fine-tuning, distillation, preference optimization, RL) OR designing agentic systems with tool use, planning, and long-horizon execution that work outside demos.
- Track record building evaluation systems for open-ended tasks with no clean labels, where success is judged by outcome.
- Experience reasoning over structured, heterogeneous, messy real-world data (configurations, graphs, logs, policy documents) rather than clean benchmarks.
- Strong software engineering fundamentals and production-quality Python code shipping and maintenance (not just scripts/POCs).
- Experience with data pipelines, distributed systems, and cloud infrastructure, preferably AWS.
- Ability to work across model behavior, APIs, and infrastructure; collaborate with attack engineers, detection engineers, product, and infrastructure teams.
- Ability to independently research unfamiliar systems and rapidly become the team expert.
- Strong written and verbal communication with clear technical documentation.
- Master's in Computer Science, Machine Learning, or related field, or equivalent practical experience, plus 4+ years professional engineering experience.
- Serious interest in how attackers and defenders operate; willingness to learn the domain in depth (pentesting/SOC background not required but domain understanding is essential).
Preferred:
- Background in detection engineering, purple teaming, security engineering, offensive security, or incident response.
- Hands-on familiarity with security control planes and APIs: EDR (CrowdStrike, SentinelOne, Defender), firewalls/segmentation (Palo Alto, Fortinet), identity/conditional access (Entra ID, Okta), SIEM/detection (Splunk, Sentinel), cloud IAM, WAF, MDM, GPO.
- Experience with causal/counterfactual inference, graph reasoning, planning, search over large state spaces, Neo4j, attack-path analysis.
- Experience building automation with write actions in production systems, including safety and change-management machinery.
- Experience with adversarial robustness or prompt injection, especially where agents consume untrusted environment input.
- Experience integrating ML into production multi-tenant SaaS or customer-controlled/air-gapped environments.