SlipstreamJobsFresh Startup & VC-Backed Jobs

Data Engineer

Graft - San Francisco, CA, United States - Hybrid

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 180,000 - 220,000 / annual

You.com is building the AI Search Infrastructure that powers modern AI systems, creating a trusted knowledge layer for agents, applications, and enterprises to retrieve real-time, accurate, and citation-backed information. The platform combines proprietary vertical indexes with LLM-optimized retrieval systems. As a Data Engineer, you will be a hands-on contributor building and scaling the modern data platform. You'll work closely with Finance, Engineering, Product, and Analytics teams to develop reliable, high-performance data pipelines and systems supporting both batch and real-time data processing. Your work will enable data activation, ensuring high-quality data flows into the warehouse and outward to business tools like Salesforce, while powering next-generation AI-driven applications including agent-based systems. Key responsibilities include: - Building and maintaining scalable data pipelines (batch and streaming) using Databricks, Spark, Kafka, and AWS services - Developing pipelines from source systems (Salesforce, billing, product events, API logs) into clean analytics layers - Designing and optimizing ETL/ELT workflows using DBT, PySpark, and SQL - Partnering with Finance on revenue accounting, COGS, margin reporting, and forecasting models - Enabling marketing and growth data use cases such as segmentation, campaign targeting, and lifecycle analytics - Developing reverse ETL pipelines to sync data from the warehouse to tools like Salesforce, HubSpot, and Braze - Creating curated datasets to support analytics, reporting, and go-to-market initiatives - Building dashboards and reporting layers for marketing and business performance tracking - Supporting AI/ML and agent-based applications by preparing datasets for MCP integrations and AI-driven applications - Monitoring pipeline performance, troubleshooting issues, and ensuring data reliability and quality - Implementing data quality checks, validations, and alerting mechanisms - Collaborating with cross-functional teams to define data contracts and ensure consistency Requirements: - 6+ years of experience in data engineering or related field - Strong hands-on experience with Databricks, AWS (S3, Glue, Athena, EMR, etc.), and Kafka - Proficiency in Python (PySpark) and SQL for large-scale data processing - Experience building and maintaining ETL/ELT pipelines (DBT/Airflow or similar preferred) - Experience with data ingestion tools such as Fivetran or similar - Familiarity with reverse ETL/data activation workflows and syncing data to tools like Salesforce, HubSpot, Braze - Exposure to or experience with AI/ML data pipelines, including RAG architectures, vector databases, or embeddings workflows - Familiarity with agent-based systems, MCP integrations, or LLM-powered applications (strong plus) - Experience working with Finance and building finance-specific metrics and pipelines (strong plus) - Understanding of data modeling and working with large-scale datasets (batch and streaming) - Experience creating dashboards and supporting reporting workflows using BI tools - Strong problem-solving skills and ability to debug production data issues - Strong communication skills and ability to work collaboratively across teams

Similar roles