About this role
About the Role You're building the runtime infrastructure that powers Moveworks' AI agents — the systems that orchestrate, execute, and deliver agent responses to millions of enterprise users in real time. This is a distributed systems engineering role at the heart of the agentic AI wave, not an ML role, focused on correctness, observability, and low latency. You'll own the end-to-end runtime that enables agents to plan, execute workflows, call tools, wait for human input, and resume. What You'll Do
- Build an agent orchestration engine: a state machine that coordinates planning, execution, and user interaction across multiple LLM calls and tool invocations
- Manage distributed sessions: lease-based ownership using DynamoDB conditional writes, heartbeat protocols, and crash recovery via checkpointing
- Create event-driven pipelines: SQS FIFO queues for ordered delivery, Kafka consumers for event processing, and real-time streaming via gRPC and Socket
- Implement IO-structured concurrency: Python asyncio TaskGroups running multiple concurrent tasks per session with fail-fast semantics and graceful cancellation
- Develop observability infrastructure: OpenTelemetry instrumentation and distributed trace context propagation across async boundaries
- Design caching/state layers: Redis and DynamoDB KV stores with per-org/per-bot scoping, batch read optimization, and hot-reload configuration What We're Looking For
- Deep experience in at least 3 of the following: Distributed systems (consistency models, idempotency, exactly-once delivery, distributed locking/leasing); Concurrent/async programming (Python asyncio, Go goroutines, structured concurrency, cancellation handling); Event-driven architectures (SQS, Kafka, pub/sub, backpressure, delivery guarantees); Database systems for infrastructure (DynamoDB, Redis); Observability (OpenTelemetry, distributed tracing, Prometheus); gRPC/protobuf (streaming RPCs, service interface design, error handling)
- 2+ years building production backend/infrastructure systems
- Strong in Python or Go (ideally both)
- Experience designing and operating systems that handle real traffic at scale
- Comfort with ambiguity — these are novel problems without textbook solutions Nice to Have
- Experience with end-to-end workflow orchestration and multi-step ML/artificial intelligence tooling
- Prior work on real-time, low-latency distributed systems at scale Compensation & Benefits
- Not disclosed