Senior Devops Engineer – Postgresql & Kafka On Kubernetes
Mirantis
Remote
Additionalinformation
Additional Information
What does Mirantis offer you?
- Work with an established Silicon Valley leader in the cloud infrastructure industry;
- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
We are a Leader for Container Management in G2 (#2 after AWS)!
Companydescription
Company Description
Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/
Jobdescription
Job Description
We are looking for a Senior Data Platform Engineer to run the PostgreSQL and Apache Kafka platform behind k0rdent-ai — our multi-tenant control plane for enterprise GPU infrastructure. Every cluster provisioned, every GPU-hour consumed, and every tenant action is recorded on this platform, so it must be reliable, secure, and recoverable. This is a hands-on DevOps / SRE role for stateful systems. You will deploy and operate PostgreSQL (CloudNativePG) and Kafka (Strimzi) on Kubernetes using operators, Helm, and GitOps — across cloud and bare-metal clusters, in a global control plane and multiple regions. You will own automation, upgrades, observability, backups, and disaster recovery, and give product teams self-service access to databases and topics through code. Main Responsibilities
- Run PostgreSQL and Kafka on Kubernetes: Deploy, upgrade, scale, and operate CloudNativePG and Strimzi clusters with Helm charts and operator custom resources, across cloud and bare-metal Kubernetes.
- GitOps & Infrastructure as Code: Manage everything declaratively with Argo CD (or Flux), Helm, and Terraform; build CI/CD pipelines for database and Kafka changes, schema migrations, and operator upgrades.
- High Availability & Disaster Recovery: Design and operate high availability within each region and cross-region disaster recovery — PostgreSQL replica clusters, backups and point-in-time recovery, Kafka MirrorMaker 2 — with defined RPO/RTO targets and regular failover drills.
- Self-Service Platform: Give product teams self-service provisioning of databases, users, topics, and ACLs through code, with secure defaults and guardrails.
- Observability & Operations: Build monitoring and alerting (Prometheus, Grafana, OpenTelemetry) for replication lag, backups, consumer lag, and capacity; participate in on-call and lead incident reviews.
- Security & Multi-Tenancy: Implement TLS/mTLS, secrets management (Vault/OpenBao), network policies, Kafka ACLs, and per-tenant isolation; integrate database and Kafka access with our identity platform (Keycloak / OIDC).
- Data Movement & CDC: Operate the outbox and change-data-capture pipelines (Debezium, Kafka Connect) that keep databases and event streams consistent between the global control plane and regions.
- Automation & Tooling: Write automation and small services in Go or Python to remove manual work, and partner with application teams on schema standards and performance troubleshooting.
Qualifications
Must Have
- Experience: 8+ years in DevOps, SRE, platform, or infrastructure engineering, including 3+ years running stateful systems (databases or message streaming) in production.
- Kubernetes: Strong hands-on Kubernetes operations — StatefulSets, persistent volumes and storage classes, pod disruption budgets, scheduling and affinity, network policies, and cluster upgrades — on cloud and/or bare-metal clusters.
- Operators & Helm: Production experience running PostgreSQL and/or Kafka through Kubernetes operators (CloudNativePG, Zalando, or Crunchy PGO; Strimzi, Confluent for Kubernetes, or similar), and authoring and maintaining Helm charts. Depth with one operator matters more than the specific product.
- GitOps & IaC: Day-to-day use of Argo CD or Flux, Terraform, and CI/CD pipelines (GitHub Actions, GitLab CI, or similar) in a declarative, review-driven workflow.
- PostgreSQL Operations: Practical PostgreSQL administration — replication and failover, backup and point-in-time recovery, connection pooling (PgBouncer), version upgrades, and performance troubleshooting.
- Kafka Operations: Running Kafka in production — brokers, topics and partitions, replication, consumer groups, monitoring, and capacity planning.
- Reliability & DR: Designing and testing high availability and disaster recovery for stateful systems, with defined RPO/RTO.
- Automation: Strong scripting and programming in Go (preferred) or Python, plus Bash.
Nice to Have
- Cross-region replication (CloudNativePG replica clusters, Kafka MirrorMaker 2) and global/regional multi-site architectures.
- Debezium, Kafka Connect, and transactional outbox / CDC patterns.
- Identity integration for data platforms (Keycloak, OIDC, LDAP/Active Directory).
- Observability stacks (Prometheus, Grafana, OpenTelemetry, VictoriaMetrics).
- Secrets and security tooling (Vault/OpenBao, cert-manager, mTLS).
- Bare-metal infrastructure, Ceph or object storage, or GPU/AI infrastructure.
- Contributions to CloudNativePG, Strimzi, or related open-source projects.
- Experience in SOC 2 or ISO 27001 environments.
Education and Experience Bachelor's degree in Computer Science or a related field, or equivalent practical experience.
Similar jobs
Senior Software Engineer, Security Factory: Code Security
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve ...
Senior Growth Analyst
What We DoWe’re on a mission to empower animal healthcare professionals with opportunities to earn more and achieve greater flexibility i...
Senior Brand & Integrated Marketing Manager
What We DoWe’re on a mission to empower animal healthcare professionals with opportunities to earn more and achieve greater flexibility i...
Product Manager
About the RoleWe're hiring a Senior Product Manager to own a core part of our supply chain orchestration product. You'll work with custom...
Senior Revenue Operations Manager
GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve ...
Director, Biometrics Strategy And Business Development (Uk)
eClinical Solutions helps life sciences organizations around the world accelerate clinical development initiatives with expert data servi...