SlipstreamJobsFresh Startup & VC-Backed Jobs

Agent Architect

Pencil - Remote - Remote - posted 2026-08-06

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Pencil is building an agentic OS for marketing—a platform that moves beyond simple generative AI integration toward sophisticated multi-agent architectures designed for global brands and small businesses alike. As Agent Architect, you will design the core "brain" of Pencil's creative engine, responsible for architecting agent logic, tool-calling structures, and evaluation loops that power both the SaaS platform and bespoke client solutions. Your work spans two high-impact pillars: (1) Core Systems—designing and scaling foundational agents that power the Pencil platform to serve 1,000+ brands with unique voices simultaneously; and (2) Client Solutions—architecting custom workflows for world-class brands that require specific brand DNA and complex creative logic. Key responsibilities include designing and implementing multi-agent workflows with task decomposition, state management, and tool integration (RAG, APIs); moving beyond vibe-based testing to implement rigorous evaluation frameworks (LLM-as-a-judge, promptfoo, DSPy) to measure and improve agent performance at scale; developing reusable agents and prompt libraries; partnering with AI Engineering and Product teams to determine whether behaviors should be handled via prompting, RAG, or fine-tuning; and acting as technical lead for complex client deployments, translating creative briefs into deterministic AI workflows. You bring 3+ years of direct GenAI experience with deep technical understanding of LLMs (GPT-4, Claude, Gemini) and multimodal models (Stable Diffusion, Midjourney, video generation). You think in systems—understanding latent space, context windows, and how one agent's output becomes another's input. You are comfortable with Python, JSON structures, and API documentation, with experience in orchestration frameworks (LangChain, CrewAI, AutoGen) being a major plus. You are obsessed with evaluation and believe that if you can't measure a prompt's performance, you shouldn't ship it. You can bridge creative and technical worlds, translating creative direction ("make it feel more punchy") into concrete technical adjustments like temperature tuning and few-shot prompting strategies. Success is measured by system reliability (reducing failure rates in complex agent workflows), architectural efficiency (minimizing token usage and latency while increasing output quality), brand alignment (automated evaluation of outputs against brand guidelines), and scale (successful rollout of core agents across the Pencil user base).

Similar roles