SlipstreamJobsFresh Startup & VC-Backed Jobs

Agent Architect

Pencil - Remote - Remote - posted 2026-08-06

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Pencil is building an agentic OS for marketing—a platform that moves beyond simple generative AI integration to create professional, brand-safe, and scalable multi-agent systems. The company serves both global brands and small businesses with complex creative workflows. As Agent Architect, you will design the core "brain" of Pencil's platform, responsible for architecting agent logic, tool-calling structures, and evaluation loops that power creative output across text, image, and video. This role bridges creative intent and machine execution, ensuring agents are robust, predictable, and capable of high-fidelity results. Your work spans two pillars: (1) Core Systems—designing and scaling foundational agents powering the Pencil SaaS platform; (2) Client Solutions—architecting custom workflows for world-class brands requiring specific brand DNA and complex creative logic. Key responsibilities include designing multi-agent workflows with task decomposition, state management, and tool-use (RAG, API integration); implementing rigorous evaluation frameworks (LLM-as-a-judge, promptfoo, DSPy) to measure agent performance at scale; developing reusable agents and prompt libraries to serve 1,000+ brands with unique voices; partnering with AI Engineering and Product teams to determine whether behaviors should be handled via prompting, RAG, or fine-tuning; and acting as technical lead for complex client deployments, translating creative briefs into deterministic AI workflows. You bring 3+ years of direct GenAI experience with deep technical understanding of LLMs (GPT-4, Claude, Gemini) and multimodal models (Stable Diffusion, Midjourney, video generation). You think in systems—understanding latent space, context windows, and how agent outputs feed into subsequent inputs. You're comfortable with Python, JSON, and API documentation, with experience in orchestration frameworks (LangChain, CrewAI, AutoGen). You're obsessed with evaluation and benchmarking—if you can't measure performance, you don't ship. You bridge creative and technical worlds, translating creative direction into concrete prompting strategies and parameter adjustments. Success is measured by reducing failure rates in complex agent workflows, minimizing token usage and latency while improving output quality, achieving high brand alignment scores using automated evaluation, and successfully rolling out core agents across the entire Pencil user base.

Similar roles