
Infrastructure engineer role focused on building and operating high-concurrency distributed systems that support simulation environments. The position involves maintaining production infrastructure at scale and optimizing performance for thousands of concurrent sessions.
Key Responsibilities:
- Operate distributed infrastructure supporting high-concurrency simulation environments
- Maintain and scale Firecracker microVM orchestration and bare metal compute clusters
- Build and improve high-throughput distributed storage and queuing systems
- Design and operate telemetry, logging, and observability systems
- Improve reliability, cold starts, uptime, and cost efficiency across the infrastructure stack
- Partner with research engineers to integrate training pipelines with infrastructure
- Handle on-call responsibilities for critical customer workloads
- Drive cost optimization and performance improvements
3 to 8 years of experience in distributed systems or infrastructure engineering
Experience maintaining and scaling high-stakes production infrastructure
Experience with storage-efficient snapshotting, copy-on-write, or forkable state systems
Experience scaling infrastructure at a later-stage startup
2+ years of big tech infrastructure experience
Large-scale distributed systems experience including queuing and distributed storage
Bare metal compute and Linux kernel-level experience
Proficiency in Python with Rust familiarity
Experience with VM or container orchestration
Experience building telemetry or observability pipelines at scale