Role Overview

The SRE (Infrastructure) Lead Engineer owns the reliability, operability, and continuous improvement of the platform that powers our retail robotics products. This is a hands-on technical leadership role for someone who can operate production systems directly, improve infrastructure through code, and guide a small Infra/SRE team toward stronger reliability, security, automation, and cost discipline.

Our current platform is primarily Azure-based, but this role does not require Azure-only experience. We value strong hands-on cloud infrastructure experience on any major cloud platform, with the ability to learn the specifics of Azure, AKS, and our tooling quickly.

You will work closely with Backend, Frontend, Robotics, Security, Product, and Operations teams to keep our cloud and Kubernetes environments dependable for live store operations, robot and smart shelf workflows, telemetry processing, internal tooling, and customer-facing services.

Current Platform State

We operate a production platform supporting:

  • Retail robotics SaaS services and internal operations tools.
  • Kubernetes-based application workloads across development, staging, infrastructure, and production environments.
  • Robot, smart shelf, and store operations systems that depend on reliable cloud-to-edge communication.
  • GitOps-style Kubernetes manifests and environment overlays.
  • Infrastructure as Code using Azure Bicep, Terraform, and Terragrunt.
  • CI/CD automation with GitHub Actions, including cloud authentication through OIDC.
  • Observability through Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, and alerting workflows.
  • You will inherit systems that are already running and help evolve them into a more standardized, automated, and resilient platform.

Tech Stack

  • Cloud: Microsoft Azure today; AWS or GCP experience is also welcome if paired with strong cloud fundamentals
  • Compute & Runtime: AKS, Kubernetes, Docker, App Service, Functions
  • Infrastructure as Code: Azure Bicep, Terraform, Terragrunt
  • GitOps & Manifests: Argo CD, Kustomize, Helm, Kubernetes YAML
  • Networking & Ingress: VNet, subnets, NSG, load balancers, Traefik, ingress-nginx, cert-manager, TLS
  • Identity & Security: Microsoft Entra ID, Azure RBAC, workload identity, managed identities, Key Vault, GitHub Actions OIDC
  • Data & Storage: Azure Cosmos DB, PostgreSQL, Redis / Redis Enterprise, Blob Storage, Storage Accounts
  • Observability: Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, Prometheus-compatible metrics, scheduled query alerts
  • Languages & Scripting: Bash, Python, TypeScript, C# or similar production scripting/application languages
  • Development Tools: Git/GitHub, GitHub Actions, Azure CLI, kubectl, Helm, Argo CD CLI

Why This Role Matters

  • Production reliability: our store operations, robot workflows, and customer systems depend on stable infrastructure.
  • Hands-on platform evolution: this role improves running systems through code, automation, standards, and direct operation.
  • Cloud-edge complexity: our platform connects cloud services, Kubernetes workloads, store systems, robots, smart shelves, and telemetry flows.
  • High business impact: reliability, deployment quality, and cost control directly affect operational efficiency and customer trust.
  • Team leadership: you will shape the practices, rituals, and technical judgment of a small Infra/SRE team.
  • Standardization opportunity: we are actively improving IaC ownership, environment separation, observability, incident response, and deployment guardrails.

Key Responsibilities

Reliability & Operations

  • Own reliability practices for production cloud and Kubernetes systems, including SLOs, SLIs, error budgets, alert quality, and operational readiness.
  • Lead incident response for infrastructure-related outages, act as Incident Commander when needed, and drive blameless post-incident reviews.
  • Build and maintain runbooks, dashboards, alerts, and operational tooling that help engineers diagnose and recover systems quickly.
  • Improve backup, restore, disaster recovery, and business continuity practices for critical compute, data, storage, and configuration systems.
  • Identify recurring operational pain and remove it through automation, better design, or clearer ownership.
  • Infrastructure & Platform Engineering

    • Design, operate, and improve Kubernetes environments, including AKS clusters, workload scheduling, autoscaling, ingress, TLS, RBAC, and cluster upgrades.
    • Maintain and evolve Infrastructure as Code using Bicep, Terraform, and Terragrunt across development, staging, infrastructure, and production environments.
    • Own GitOps-style application deployment patterns using Argo CD, Kustomize, Helm, and environment overlays.
    • Improve CI/CD pipelines with safe promotion flows, validation, drift detection, security checks, and rollback strategies.
    • Manage cloud networking, identity, secrets, certificates, and access controls in partnership with Security Engineering.
    • Support application teams with platform guidance for resource requests, scaling, dependency management, release safety, and production readiness.
    • Observability, Security, and Cost

      • Standardize logs, metrics, traces, dashboards, and alerts across services and infrastructure.
      • Operate and improve observability tooling such as Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, and scheduled query alerts.
      • Partner with Security Engineering on RBAC, workload identity, managed identities, Key Vault usage, secret rotation, cloud access boundaries, and auditability.
      • Drive cloud cost visibility and optimization through tagging, right-sizing, autoscaling, log retention controls, budget review, and FinOps practices.
      • Keep reliability, security, and cost decisions practical for a growing robotics business.
      • Team Leadership

        • Lead and mentor 3-6 Infra/SRE/Platform engineers while staying deeply hands-on.
        • Set technical direction for infrastructure reliability, deployment automation, observability, and cloud operations.
        • Run design reviews, code reviews, operational reviews, and regular improvement planning.
        • Build healthy on-call practices, escalation paths, incident roles, and postmortem follow-through.
        • Communicate priorities, risks, and tradeoffs clearly to engineering leadership and cross-functional partners.
        • Cross-Team Collaboration

          • Work with Backend and Frontend teams to make services easier to deploy, observe, scale, and operate.
          • Work with Robotics and Operations teams to understand how infrastructure behavior affects live robot and store workflows.
          • Work with Product and Engineering Management to translate non-functional requirements into roadmap items and engineering guardrails.
          • Help teams adopt standard platform patterns without blocking delivery.

Qualifications

Must Have

  • 5+ years of professional infrastructure, platform, SRE, DevOps, backend infrastructure, or cloud engineering experience.
  • 2+ years in a senior, lead, or technical leadership role with responsibility for production reliability or infrastructure direction.
  • Strong hands-on experience operating production systems on at least one major cloud platform such as Azure, AWS, or GCP.
  • Production Kubernetes experience, including deployments, services, ingress, autoscaling, resource management, RBAC, troubleshooting, and upgrades.
  • Hands-on Infrastructure as Code experience with Terraform, Bicep, CloudFormation, Pulumi, CDK, or similar tools.
  • Experience designing or operating CI/CD pipelines and deployment workflows for production services.
  • Strong Linux, networking, DNS, TLS, and cloud identity fundamentals.
  • Practical observability experience with metrics, logs, traces, dashboards, alerting, and incident response.
  • Experience leading incidents, writing postmortems, and driving corrective actions to completion.
  • Ability to write scripts or small tools in Bash, Python, TypeScript, Go, C#, or a comparable language.
  • Clear communication skills and the ability to work across engineering, operations, security, and product teams.
  • Nice to Have

    • Azure production experience, especially AKS, Azure Monitor, Application Insights, Log Analytics, Key Vault, Entra ID, managed identities, and Azure RBAC.
    • Experience with Argo CD, Kustomize, Helm, cert-manager, Traefik, ingress-nginx, Grafana, Loki, Alloy, or Prometheus-style monitoring.
    • Experience with GitHub Actions OIDC, workload identity, or federated cloud authentication patterns.
    • Experience operating data services such as PostgreSQL, Redis, Cosmos DB, MongoDB-compatible databases, or cloud storage systems.
    • Experience with drift detection, policy-as-code, cloud guardrails, or compliance-oriented infrastructure workflows.
    • Background in robotics, IoT, retail operations, edge computing, telemetry systems, or other cyber-physical production environments.
    • Experience building or improving an on-call program for a growing engineering organization.
    • Japanese language skills are helpful but not required.

Success Profile

  • Hands-on operator: comfortable debugging real incidents, reading manifests, reviewing Terraform/Bicep, and using cloud/Kubernetes CLIs directly.
  • Reliability-minded: thinks in SLOs, failure modes, blast radius, recovery paths, and operational feedback loops.
  • Pragmatic leader: balances engineering quality with business urgency and team capacity.
  • Automation-oriented: turns repeated manual work into durable tools, workflows, or platform patterns.
  • Security-aware: treats access, secrets, identity, network boundaries, and auditability as part of daily infrastructure work.
  • Cost-conscious: understands that cloud architecture must be reliable and financially sustainable.
  • Collaborative teacher: raises the operational maturity of surrounding teams through guidance, reviews, and shared standards.

Vision & Growth

  • Build a stronger SRE and platform engineering practice for retail robotics.
  • Standardize infrastructure ownership across cloud resources, Kubernetes clusters, manifests, CI/CD, observability, and incident response.
  • Help scale the platform from today’s operational needs toward enterprise-grade reliability across more stores, robots, smart shelves, and customer environments.
  • Shape the long-term infrastructure roadmap while remaining close to the systems that keep the business running.

SRE (Infrastructure) Lead Engineer

Apply On Company Site
Back to search page