SlipstreamJobsFresh Startup & VC-Backed Jobs

Agent Architect

Pencil - Remote - Remote - posted 2026-08-06

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Pencil is building an agentic OS for marketing—a platform that moves beyond simple text-in/text-out AI interfaces toward sophisticated multi-agent architectures designed for global brands and small businesses alike. The company is focused on making generative AI professional, brand-safe, and scalable for creative work. As Agent Architect, you will design and engineer the "brain" of Pencil's platform. This is a systems-level role bridging creative intent and machine execution. You'll work across two high-impact pillars: (1) Core Systems—designing foundational agents powering the Pencil SaaS platform, and (2) Client Solutions—architecting custom workflows for world-class brands requiring specific brand DNA and complex creative logic. Key responsibilities include: - Designing and implementing multi-agent workflows with task decomposition, state management, and tool-use (RAG, API integration, etc.) - Implementing rigorous evaluation frameworks (LLM-as-a-judge, promptfoo, DSPy) to measure and improve agent performance at scale - Developing reusable agents and prompt libraries to serve 1,000+ brands with unique voices simultaneously - Partnering with AI Engineering and Product teams to determine whether behaviors should be handled via prompting, RAG, or fine-tuning - Acting as technical lead for complex client deployments, translating creative briefs into deterministic AI workflows You bring 3+ years of direct GenAI experience with deep technical understanding of LLMs (GPT-4, Claude, Gemini) and multimodal models (Stable Diffusion, Midjourney, video generation). You think in systems—understanding latent space, context windows, and how one agent's output feeds another's input. You're comfortable with Python, JSON, and API documentation, with experience in orchestration frameworks (LangChain, CrewAI, AutoGen) being a major plus. Critically, you're obsessed with evaluation and measurement—you believe if you can't measure a prompt's performance, you shouldn't ship it. You can translate creative direction ("make it feel more punchy") into technical adjustments (temperature tuning, few-shot prompting). You see hallucinations as logic puzzles to solve and are excited by the challenge of making AI follow complex brand guidelines with high fidelity. Success is measured by system reliability (reducing agent workflow failure rates), architectural efficiency (minimizing token usage/latency while improving quality), brand alignment (automated evaluation against brand guidelines), and scale (core agents adopted across the user base).

Similar roles