SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff AI Engineer - AI Labs

dLocal - Madrid, Madrid, Spain - Hybrid - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

dLocal is the financial infrastructure powering global commerce in emerging markets, operating in 60+ countries. The AI Lab is a specialized team tasked with validating emerging AI and automation technologies and de-risking their adoption across the organization. You will join as a Staff AI Engineer—a senior individual contributor role with significant technical influence but no direct people management responsibility. You own technology scouting, prototyping, and evaluation for dLocal's AI adoption strategy. Key Responsibilities: Technology Scouting & Evaluation: Run instrumented spikes and benchmarking on new models, tools, and frameworks (LLMs, agentic systems, vector databases, orchestration frameworks, copilots, assistants). Compare vendor and open-source options across quality, cost, latency, security, and integration complexity. Deliver concise decision memos with clear recommendations (adopt, watch, or avoid). Evaluation Harnesses & Sandboxes: Design and maintain evaluation environments with datasets, prompts, scenarios, and telemetry to test models under realistic constraints. Build automation and tooling to measure quality, robustness, latency, and cost, including regression tracking over time. Ensure every evaluated technology has benchmark coverage and documented risk/limitations views. Prototyping & Technical Validation: Build working prototypes to understand how technologies behave under realistic conditions, not just in vendor demos. Explore architecture, integration patterns, operational constraints, security boundaries, and failure modes. Determine what must be true for a proof of concept to become viable production capability. Recommendations, Readiness & Hand-offs: Translate technical findings into clear decision memos for technical and non-technical stakeholders. For validated technologies, produce readiness guidance covering recommended patterns, guardrails, known limitations, operational considerations, and integration requirements. Coordinate hand-offs to engineering teams responsible for productionization and support transitions when deep context is required. Governance, Risk & Standards: Work with Security, Legal, Compliance, and other AI teams to document risk assessments, mitigations, and governance recommendations. Maintain checklists, decision templates, and lightweight standards reusable across evaluations. Incorporate learnings from third-party AI tooling (external copilots, AWS AI suite) into adoption guidelines. Collaboration, Mentoring & Community: Partner with other AI teams and domain teams to ensure clear boundaries and smooth collaboration. Participate in hiring as a technical evaluator and culture champion. Mentor engineers on evaluation methods, benchmarking, and experimental design. Share knowledge through internal write-ups, tech talks, and occasional external meetups and conferences. Requirements: • 8+ years of software engineering experience, including significant experience operating at senior or Staff-level scope • Deep hands-on experience building and evaluating systems based on LLMs and modern AI tooling • Strong software engineering fundamentals and ability to rapidly build high-quality experimental systems • Experience building agentic or multi-step AI systems involving tool use, orchestration, state, retrieval, or external integrations • Strong knowledge of cloud infrastructure, preferably AWS, and ability to run experimental workloads securely and cost-consciously • Experience with observability, telemetry, testing, and benchmarking of complex systems • Ability to reason about system architecture, reliability, scalability, asynchronous workflows, and distributed components • Track record of designing experiments or benchmarks that influenced meaningful technical decisions • Track record designing and running benchmarks that compare AI models and tools under real constraints • Experience constructing evaluation datasets: task selection, labeling, holdout discipline, and maintaining useful sets as models improve • Working knowledge of LLM-as-judge methods and their failure modes, alongside human evaluation, inter-annotator agreement, and judgment on when each is appropriate • Ability to reason about statistical significance on small samples and state confidence honestly rather than over-reading results • Familiarity with regression tracking, telemetry, and versioning to keep results reproducible over time • Ability to turn ambiguous "we should try this new thing" ideas into well-scoped evaluation plans with clear hypotheses and metrics • Comfortable making trade-off calls across quality, latency, cost, and vendor lock-in, and documenting them clearly • Experience writing short, opinionated decision memos that help others move fast • Ability to explain technical results to non-specialists in concrete, concise terms • Experience working with platform, product, and operations teams to align evaluations with real use cases • Ability to influence without authority, aligning teams around shared standards and guardrails • Curious and biased toward experimentation, combined with disciplined measurement and risk awareness • Comfortable in a small, high-leverage team without embedded PMs; able to structure own work and keep stakeholders informed • Builder attitude: preference for reusable tools, templates, and playbooks over one-off work

Similar roles