Talent Apply
Log in
All jobs
RF

Senior Infrastructure Engineer, Applied AI

Recruiting From Scratch (RFS Group)
San Francisco, CA
On-siteUSD 250,000 - 350,000 / year

About this role

Senior Infrastructure Engineer, Applied AI

Full time

San Francisco, CA | Onsite

6 - 10 Years of Experience

Who is Recruiting from Scratch:

Recruiting from Scratch is a premier talent firm that focuses on placing the best product managers, software, and hardware talent at innovative companies. Our team is 100% remote and we work with teams across the United States to help them hire.

Senior Infrastructure Engineer, Applied AI

Location

San Francisco, CA

Fully onsite, 5 days per week, Monday–Friday.

Company Stage of Funding

High-Growth, Venture-Backed AI Infrastructure Company

Office Type

On-site — 5 days per week

Salary

$250,000 – $350,000 Base

OTE: $300,000 – $450,000

Equity

Competitive Equity

Visa

Open to visa transfers, including OPT and H-1B transfers.

Experience

  • 6–10 years of experience building and maintaining platform infrastructure, distributed systems, or large-scale production systems, with continued full-stack engineering capabilities.

  • Employment Type

  • Full-time

  • Hiring Count

  • Growth Hiring — 6–8 hires

  • Company Description

  • This is a high-growth AI infrastructure company building the data generation, human-in-the-loop systems, and evaluation infrastructure that powers leading AI teams and advanced AI applications. Its platform helps organizations generate, manage, and evaluate large-scale datasets used to develop and improve AI systems.

  • The company has experienced rapid growth and is expanding its infrastructure to support increasing data volumes, demanding workloads, and multiple product verticals. Its platform supports complex, data-intensive workflows across areas such as healthcare, finance, and other enterprise applications.

  • As the company scales, it needs a dedicated senior infrastructure engineer to own the foundational platform that enables engineering teams to build and ship products reliably. This person will help establish the architecture, systems, and engineering practices required to support high-throughput workloads, distributed processing, and continued company growth.

  • The engineering team is collaborative and fast-moving. You will work closely with vertical-focused engineers, mentor junior team members, and build shared infrastructure that multiple teams depend on. The role combines deep infrastructure and distributed systems expertise with enough full-stack capability to understand and contribute across the broader product.

  • This is a high-impact opportunity for an engineer who enjoys solving complex infrastructure problems, taking ownership of critical systems, and building scalable foundations in an early-stage, high-growth environment.

  • What You Will Do

    1. Design & Build Core Platform Infrastructure
  • Design, build, and maintain the shared infrastructure powering data generation platforms, human-in-the-loop systems, and AI evaluation pipelines.

  • Take ownership of foundational systems used across multiple product verticals.

  • Architect scalable, reliable services that support growing workloads and increasing data volumes.

  • Build platform capabilities that enable other engineering teams to develop and deploy products efficiently.

  • Identify infrastructure bottlenecks and design solutions that improve system performance, scalability, and maintainability.

  • Establish clear architectural patterns and reusable infrastructure components.

  • Make pragmatic decisions about system design, performance, cost, and long-term reliability.

    1. Build & Operate Distributed Systems at Scale
  • Architect and develop distributed systems capable of processing large datasets and high-throughput workloads.

  • Design services and workflows that remain reliable as traffic, data volumes, and system complexity increase.

  • Work with Kubernetes and cloud environments, including AWS and GCP.

  • Build and maintain asynchronous processing systems and event-driven workflows.

  • Use technologies such as Kafka, Redis, and Elasticsearch where appropriate to support data processing, messaging, caching, and search.

  • Diagnose performance issues, reliability problems, and complex production failures.

  • Improve system resilience, fault tolerance, and operational efficiency.

  • Ensure infrastructure can support the company's growing product and customer requirements.

    1. Establish Reliability, Observability & Engineering Standards
  • Establish engineering best practices for system design, deployment, monitoring, and reliability.

  • Build observability into core systems through appropriate logging, metrics, tracing, and alerting.

  • Define and improve reliability standards and operational practices for critical infrastructure.

  • Identify failure modes and develop strategies for graceful degradation and recovery.

  • Improve deployment workflows and infrastructure management practices.

  • Partner with engineering teams to ensure platform services are dependable and easy to use.

  • Help create consistent approaches to testing, incident response, documentation, and production readiness.

  • Balance rapid delivery with the reliability required for mission-critical, high-throughput systems.

    1. Mentor Engineers & Drive Technical Execution
  • Mentor junior engineers and help raise the team's technical standards.

  • Collaborate with engineers across product verticals to understand infrastructure requirements.

  • Translate team needs into reusable platform capabilities and reliable services.

  • Lead technical discussions around architecture, scalability, performance, and reliability.

  • Provide guidance on code quality, system design, deployment practices, and operational ownership.

  • Remain hands-on with implementation while helping other engineers make effective technical decisions.

  • Take initiative in identifying infrastructure gaps and driving solutions through production.

  • Help build a collaborative, high-performance engineering culture as the company grows.

  • Ideal Candidate Background

  • Experience Requirements

  • 6–10 years of professional software engineering experience.

  • Strong experience building and maintaining production distributed systems or platform infrastructure.

  • Experience designing and operating systems in cloud environments, ideally AWS or GCP.

  • Experience building infrastructure that supports large-scale data processing or high-throughput workloads.

  • Experience working at a venture-backed startup, ideally as an early engineer or founding engineer at a company that raised meaningful institutional funding.

  • Alternatively, experience at a company specializing in data infrastructure or large-scale data management.

  • Experience taking ownership of foundational systems used by multiple engineering teams.

  • Ability to mentor engineers while remaining hands-on as an individual contributor.

  • Full-stack engineering capabilities, with a clear strength in backend, infrastructure, and platform engineering.

  • Comfort operating in a fast-moving environment where infrastructure needs evolve alongside the product.

  • A bachelor's degree in Computer Science or a closely related technical field is preferred, with a strong academic background or equivalent demonstrated technical expertise.

  • Technical Requirements

  • Strong experience designing, building, and operating distributed systems.

  • Deep knowledge of platform infrastructure and production system architecture.

  • Strong Kubernetes expertise.

  • Experience with cloud infrastructure, particularly AWS or GCP.

  • Understanding of high-throughput data processing and scalable service design.

  • Experience with infrastructure deployment, production operations, and system troubleshooting.

  • Strong understanding of observability, monitoring, logging, tracing, and reliability engineering.

  • Experience designing fault-tolerant services and resilient distributed workflows.

  • Strong system design and architectural decision-making skills.

  • Familiarity with infrastructure automation, deployment practices, and production incident management.

  • Ability to reason about performance, availability, scalability, and operational complexity.

  • AI, Data & Infrastructure Requirements

  • Experience building infrastructure for data-intensive or AI-related products is strongly preferred.

  • Understanding of the infrastructure required to support data generation, data processing, and evaluation pipelines.

  • Experience supporting human-in-the-loop workflows, data annotation, or data operations is valuable.

  • Experience working with high-volume datasets and distributed processing workloads.

  • Familiarity with Kafka or other event-streaming and messaging technologies.

  • Experience with Redis, Elasticsearch, or comparable caching and search systems.

  • Experience building platform services that support multiple product teams or business verticals.

  • Ability to build reliable infrastructure that allows other engineers to move quickly without compromising system stability.

  • Understanding of production observability, system reliability, and operational readiness.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →