SlipstreamJobsFresh Startup & VC-Backed Jobs

Lead Endpoint Researcher

PointFive - Tel Aviv, Israel - In-office - posted 2026-09-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

PointFive is an AI Efficiency OS that helps engineering and FinOps teams manage AI spend and cloud costs. The company recently closed a $60M Series B led by Accel and was founded by the team behind IntSights (acquired by Rapid7). TokenShift is PointFive's newest product—an endpoint runtime for AI coding agents that provides organizations control, visibility, and efficiency across agent execution and model consumption. As AI agents move from chat interfaces into developer machines, they need access to shells, filesystems, credentials, and local compute while enterprises increasingly run models locally across heterogeneous endpoint environments. In this role, you will research how AI coding agents execute and isolate workloads, design secure sandboxing mechanisms, and build infrastructure for deploying and operating local models across diverse endpoint environments. You'll work across macOS, Windows, Linux, WSL, and developer containers to enable agents with powerful capabilities while preventing unrestricted machine access. Key responsibilities include: - Researching AI coding agent runtimes by reverse-engineering tools like Claude Code, Cursor, Copilot CLI, and emerging agents to understand how they execute commands, access files, and interact with the operating system - Designing agent sandboxing architectures that safely constrain filesystem access, networking, process execution, credentials, and sensitive capabilities - Exploring and implementing OS-native isolation mechanisms including containers, namespaces, sandbox profiles, virtualization, and permission models - Building endpoint harnesses that intercept, control, observe, and modify agent execution flows - Designing infrastructure for deploying and running local LLMs across heterogeneous endpoint fleets, accounting for OS, CPU architecture, GPU availability, and hardware acceleration - Evaluating local inference runtimes and model formats such as llama.cpp, MLX, ONNX Runtime, Ollama, and vLLM-compatible runtimes - Developing mechanisms for model discovery, installation, versioning, updates, health monitoring, and lifecycle management across developer machines - Researching hardware-aware model selection and routing for different endpoint configurations - Building benchmarks measuring model latency, throughput, memory consumption, startup time, and developer experience on real hardware - Investigating security boundaries between local models, cloud-hosted models, agents, MCP servers, shells, and enterprise resources - Collaborating with core CLI and platform teams to move endpoint research from prototype to production - Tracking the rapidly evolving AI agent, sandboxing, and local inference ecosystems to identify technologies PointFive should support Tech stack: Go, TypeScript, Cloudflare, Snowflake, local inference runtimes, and OS-native endpoint technologies across macOS, Windows, and Linux. REQUIREMENTS Must-have: - Deep understanding of operating system fundamentals: processes, permissions, filesystems, IPC, networking, signals, and process trees - Strong understanding of sandboxing and workload isolation concepts - Hands-on systems experience across multiple operating systems, ideally including Linux, macOS, and Windows - Experience building or debugging software that runs directly on developer endpoints, workstations, or servers - Comfort operating close to the metal—binary behavior, system calls, OS APIs, environment variables, config files, process injection/interception, and runtime behavior - Investigative, reverse-engineering mindset; you enjoy taking apart systems you didn't build and figuring out how they actually work - Ability to move comfortably between research prototypes and production-quality systems code Nice-to-have: - Experience running or deploying local LLMs using frameworks such as llama.cpp, MLX, Ollama, ONNX Runtime, or similar inference runtimes - Understanding of model quantization, KV-cache behavior, GPU memory constraints, model loading, and inference performance - Experience with NVIDIA CUDA, Apple Metal, DirectML, ROCm, or other hardware acceleration stacks - Experience with Linux namespaces, seccomp, cgroups, AppArmor, SELinux, eBPF, or container runtimes - Familiarity with macOS sandboxing, Endpoint Security Framework, launchd, XPC, virtualization, or Apple Silicon - Familiarity with Windows security primitives, Job Objects, AppContainers, Windows Sandbox, Hyper-V, WSL, ETW, or Windows process APIs - Hands-on familiarity with AI coding agents such as Cursor, Claude Code, Copilot CLI, and Codex - Experience with endpoint management or fleet deployment technologies such as MDM, Intune, Jamf, SCCM, or enterprise software distribution systems - Experience with containers, microVMs, or lightweight virtualization technologies - TypeScript experience

Similar roles