The role
You'll own reliability end to end for our infrastructure platforms, at least one of which carries a 99.99% availability target: define how we measure it, build the automation and guardrails that protect it, and raise the bar of the DevOps engineers working alongside you.
This is a hands-on role with real ownership. You'll make architectural calls and be accountable for them, lead when incidents happen, and mentor others rather than just executing tickets. You'll work closely with the Head of Engineering and the Tech Lead, who own the product roadmap and the application platform; you own how reliably it all runs.
What you will do
Own reliability. Define SLIs, SLOs and error budgets with product and engineering teams, and use them to drive priorities and alerting.
Design for high availability. Architect multi-AZ Kubernetes platforms with clear RTO/RPO targets; design, document and regularly test disaster-recovery procedures, including the stateful layer (databases, storage, tested restores).
Plan capacity. Load-test critical paths, forecast growth, right-size clusters and workloads, and know where systems will break before they do.
Own change and delivery. Run GitOps-based deployments; define how changes are reviewed, rolled out and rolled back, and make every production change traceable to its cause.
Build the platform as code. Provision and evolve cloud infrastructure with Terraform or AWS CDK, with security (IAM, network segmentation, secrets, RBAC) built in from the start.
Make systems observable. Design metrics, logging, tracing and alerting so that on-call pages on user-facing symptoms, not noise.
Lead incident response. Participate in on-call, lead incidents when they happen, run blameless post-mortems and drive the follow-ups to completion.
Enable and mentor. Coach development teams on containerization, delivery patterns and operational readiness; mentor DevOps engineers and set the technical standards for the team.
What we need from you
Production ownership. Several years operating business-critical systems in production, including on-call and incident leadership. You can point to systems you kept reliable and explain what you changed to get there.
Kubernetes depth. Hands-on with core objects (Deployments, StatefulSets, Services, Ingress, ConfigMaps, Secrets, HPA) and cluster-level concerns: scheduling, resource management, upgrades, RBAC, persistent storage and troubleshooting on managed (EKS/GKE/AKS) or self-managed clusters.
GitOps, CI/CD and change management. Strong GitLab CI/CD (.gitlab-ci.yml, runner optimization, artifacts, secret handling); experience running GitOps deployments with ArgoCD or Flux; a track record of safe rollout and rollback practices.
Infrastructure as Code. Proven experience with Terraform or AWS CDK: modules/constructs, state management, provisioning at scale.
Platform security. Cloud IAM, network segmentation, secrets management and Kubernetes RBAC as everyday practice, plus container image security and minimal multi-stage Dockerfiles.
Observability and SLOs. Experience designing SLIs/SLOs and building alerting around them, using tools such as Prometheus, Grafana, OpenTelemetry or Datadog.
HA, DR and data. Designed and operated multi-AZ architectures; defined and actually tested disaster recovery, including backups and restores, database replication and upgrades for stateful services.
Capacity and performance. Experience with load testing, capacity forecasting and right-sizing, and diagnosing performance bottlenecks in production.
Fundamentals. Solid Linux, networking (DNS, TLS, load balancing) and AWS knowledge; proficiency in Python or Bash for automation.
Leadership. Experience mentoring engineers and driving technical decisions across teams.
Nice to have
Multi-region or active-active architectures
Service mesh (Istio, Linkerd or Cilium) for mTLS, traffic management and resilience patterns
Progressive delivery (canary / blue-green) with Argo Rollouts or Flagger
Supply-chain security and policy-as-code (e.g. SBOMs, image signing, Kyverno/OPA)
Ansible or other configuration-management tooling
A second cloud, ideally GCP
Chaos engineering or resilience testing
Cost awareness and FinOps practices
Start-up / scale-up experience
How we work
Small team, high trust β you'll talk to the CEO as easily as to the person sitting next to you
Open to experimenting with new technologies: if you can show it solves a real problem, you'll get the room to try it
Relaxed, low-ego environment β we care about the work being good, not about hierarchy or hours at a desk
Room to shape how we build: our engineering practices are still being written, and you'll have a real say in them
About Opplane
Opplane specializes in providing advanced data-focused solutions for financial services, telecommunication, and reg-tech to accelerate their digital transformation journey. Opplane leadership team is comprised of Silicon Valley serial entrepreneurs and experienced executives. Its expertise comes from years of specific industry experience at some of the worldβs top companies, such as PayPal, Xerox Parc, Amazon, Wells Fargo, SoFi in the areas of product management, data technology, data governance, data privacy, security, machine learning, and risk management.
Why Opplane
π Global & Multicultural β Diverse perspectives, global collaboration (US, Portugal, India and Singapore offices)
β‘ Startup Energy β Fast-moving, impact-driven environment
πͺ Ownership Mindset β Engineers own what they build
π€ Collaborative & Friendly β Open, curious, and supportive culture
Most organizations deploying AI to engineering cannot say what they got for it and respond by buying more of it or by arguing. You will build the evidence instead and then use it to decide where the next capability goes.
The measurement program is the first project. The scope is applied AI across the delivery lifecycle, with real internal users, a platform team to build on, and leadership that will act on what you find.