Talent Apply
Log in
All jobs
M

Distributed Systems Engineer

Moveworks
Remote
Remote

About this role

About the Role You're building the runtime infrastructure that powers Moveworks' AI agents — the systems that orchestrate, execute, and deliver agent responses to millions of enterprise users in real time. This is a distributed systems engineering role at the heart of the agentic AI wave, not an ML role, focused on correctness, observability, and low latency. You'll own the end-to-end runtime that enables agents to plan, execute workflows, call tools, wait for human input, and resume. What You'll Do

  • Build an agent orchestration engine: a state machine that coordinates planning, execution, and user interaction across multiple LLM calls and tool invocations
  • Manage distributed sessions: lease-based ownership using DynamoDB conditional writes, heartbeat protocols, and crash recovery via checkpointing
  • Create event-driven pipelines: SQS FIFO queues for ordered delivery, Kafka consumers for event processing, and real-time streaming via gRPC and Socket
  • Implement IO-structured concurrency: Python asyncio TaskGroups running multiple concurrent tasks per session with fail-fast semantics and graceful cancellation
  • Develop observability infrastructure: OpenTelemetry instrumentation and distributed trace context propagation across async boundaries
  • Design caching/state layers: Redis and DynamoDB KV stores with per-org/per-bot scoping, batch read optimization, and hot-reload configuration What We're Looking For
  • Deep experience in at least 3 of the following: Distributed systems (consistency models, idempotency, exactly-once delivery, distributed locking/leasing); Concurrent/async programming (Python asyncio, Go goroutines, structured concurrency, cancellation handling); Event-driven architectures (SQS, Kafka, pub/sub, backpressure, delivery guarantees); Database systems for infrastructure (DynamoDB, Redis); Observability (OpenTelemetry, distributed tracing, Prometheus); gRPC/protobuf (streaming RPCs, service interface design, error handling)
  • 2+ years building production backend/infrastructure systems
  • Strong in Python or Go (ideally both)
  • Experience designing and operating systems that handle real traffic at scale
  • Comfort with ambiguity — these are novel problems without textbook solutions Nice to Have
  • Experience with end-to-end workflow orchestration and multi-step ML/artificial intelligence tooling
  • Prior work on real-time, low-latency distributed systems at scale Compensation & Benefits
  • Not disclosed

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →