Golang/Kubernetes Engineer
Golang/Kubernetes Engineer
> ROLE OVERVIEW
We are seeking a Senior Engineer to lead development of the ArangoDB operator, responsible for ensuring robust lifecycle management, scaling, back-up, and replication capabilities for ArangoDB running on Kubernetes. You will shape the architecture of core operator components, design and implement new features, influence best practices for deploying stateful systems on Kubernetes, and ensure that ArangoDB continues to meet enterprise-grade expectations for performance, consistency, and stability. This is an opportunity to own a critical part of our ecosystem and directly impact how modern AI-driven applications run in production. Arango delivers a unified, natively multimodel contextual data platform that powers AI agents, assistants, and applications with the unified, current, and trusted business context needed to reason, decide, and act at scale.
> CORE RESPONSIBILITIES
- Lead the design and development of the ArangoDB Kubernetes Operator.
- Own core operator architecture and implement features for lifecycle management, scaling, backup, and replication.
- Ensure high availability, performance, and stability of ArangoDB on Kubernetes.
- Define best practices for running stateful systems on Kubernetes.
- Collaborate with engineering teams to align operator capabilities with enterprise and cloud requirements.
- Troubleshoot complex issues across Kubernetes and distributed systems.
> HARD REQUIREMENTS & SPECS
- 4+ years of programming and Cloud experience
- Deep expertise in Kubernetes, CRDs and operators (deployments, custom resources like ArangoDeployment, ArangoBackup, ArangoLocalStorage, replication, etc.)
- Strong experience in Go (the operator is written in Go) plus solid understanding of concurrency, storage, and distributed system concerns.
- Strong experience in Cloud solutions, especially AWS. AWS GovCloud is an additional point
- Proven background in designing and managing production-grade distributed databases or stateful systems, including storage management and data replication. (e.g., persistent volumes, snapshot/backup, failover, cluster scaling)
Is the AI extraction inaccurate? Report an issue