Appealing Points:

  • Play a key technical role in a global SRE/DevOps operations team, driving operational excellence for a large-scale, mission-critical connected vehicle insights SaaS platform used across automotive, insurance, and emerging autonomous-driving use cases.
  • Lead high-impact initiatives such as root cause analysis, post-incident reviews, and service reliability improvements, while mentoring and coaching junior engineers and DevOps team members.
  • Gain exposure to a modern, cross-functional environment — collaborating with global stakeholders and offshore delivery teams while working with technologies like Kubernetes, CI/CD, Infrastructure as Code, and observability tooling.

Annual Salary: ¥10,000,000 and above

Job Description:

The role sits within a SaaS operations team supporting a connected vehicle insights offering, delivered as a dedicated SaaS environment for customers across multiple industries, including automotive and insurance. The platform enables retrieval, management, and analysis of large-scale vehicle sensor data, supporting near-real-time analytics for use cases such as autonomous driving and advanced insurance services. The role is technical in nature, focused on maintaining cloud-based SaaS services, automating security patching and deployment processes, resolving operational issues, and collaborating cross-functionally to ensure service stability — contributing to operational excellence through modern engineering practices such as SRE.

Job Responsibilities:

  • Execute day-to-day SaaS operational tasks as part of a global operations team spanning multiple geographies, ensuring 24x7 customer support.
  • Communicate directly with customers (typically via their consulting service teams) to resolve issues and handle customization requirements (e.g., custom plug-in deployments).
  • Manage SaaS inventory and apply security fixes, including kernel and non-kernel OS-level patches.
  • Provide on-call support, including pager duty alert handling as both first and second responder.
  • Collaborate with Product Management, Engineering, Architecture, and business stakeholders to drive operational readiness, service reliability, and customer success.
  • Lead operational reviews, root cause analysis (RCA), service improvement initiatives, and post-incident management.
  • Mentor and coach junior engineers, SREs, and DevOps team members to build a culture of ownership and accountability.

Required Qualifications:

  • 15+ years of experience in Site Reliability Engineering (SRE), DevOps, Cloud Operations, and Production Support for large-scale enterprise SaaS platforms.
  • Excellent verbal and written communication skills, with the ability to collaborate across engineering, product management, operations, and customer teams.
  • Strong analytical, troubleshooting, and problem-solving skills with a focus on operational excellence and continuous improvement.
  • Extensive experience managing mission-critical production environments, ensuring high availability, performance, reliability, and operational stability.
  • Proven experience in Incident Management, Major Incident Response, Problem Management, and 24x7 On-Call Operations using tools such as PagerDuty.
  • Strong expertise in Linux administration, Java/JVM performance analysis, networking, security, and cloud infrastructure, including cloud virtual servers (VSI) and Bare Metal environments.
  • Deep understanding of DevOps practices, Infrastructure as Code (IaC), automation, CI/CD pipelines, and operational tooling using Python, Bash, Ansible, Elastic Stack, and InfluxDB.
  • Experience implementing observability and monitoring solutions, including metrics, logging, tracing, alerting, and capacity planning.
  • Strong understanding of IoT-based products and platforms, preferably connected vehicle insights or similar connected ecosystem solutions.
  • Ability to communicate and coordinate effectively with Japanese customers, global stakeholders, and offshore delivery teams.
  • Experience defining and managing Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and reliability engineering best practices.
  • Strong leadership skills with the ability to manage technical escalations, drive cross-functional collaboration, and lead operational transformation initiatives.

Preferred Qualifications:

  • Strong knowledge of Maximo, WebSphere Liberty, DB2, Hadoop, and other enterprise application platforms.
  • Hands-on experience troubleshooting complex issues across application, middleware, database, cloud, and infrastructure layers.
  • Experience implementing and driving SRE practices, reliability frameworks, and operational excellence programs.
  • Familiarity with Kubernetes, container platforms, cloud-native technologies, and modern observability solutions.
  • Experience leading cloud migration, operational modernization, and automation initiatives within enterprise environments.
  • Exposure to OpenTelemetry, APM platforms, observability engineering, and enterprise monitoring frameworks.
  • Experience working with global customers, particularly in the Japan region, and with multicultural operational teams.
  • Experience driving end-to-end cloud migration projects involving complex integrations.
  • Strong architectural understanding of cloud-native platforms and distributed systems, with proven experience partnering with architects and leading DevOps teams.

Language Skills: Fluent level Japanese (JLPT N2 and above) and Business level English

Company Description:

We're a trusted partner in Digital Engineering and Enterprise Modernization, leveraging deep technical expertise to help clients stay ahead. Our solutions empower clients to outpace competition. Partnering with industry leaders worldwide, including top US companies and banks, we drive innovation across Healthcare and Life Sciences; Banking, Financial Services, and Insurance; Software and Hi-Tech; and Emerging Verticals.

. Skillset Required: Site Reliability Engineering (SRE), DevOps, Cloud Operations, Production Support, SaaS platforms, verbal communication, written communication, cross-functional collaboration, analytical skills, troubleshooting, problem-solving, operational excellence, continuous improvement, Incident Management, Major Incident Response, Problem Management, 24x7 On-Call Operations, PagerDuty, Linux administration, Java/JVM performance analysis, networking, security, cloud infrastructure, cloud virtual servers (VSI), Bare Metal environments, Infrastructure as Code (IaC), automation, CI/CD pipelines, operational tooling, Python, Bash, Ansible, Elastic Stack, InfluxDB, observability, monitoring solutions, metrics, logging, tracing, alerting, capacity planning, IoT-based products, connected vehicle insights, Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, reliability engineering, technical escalations, cross-functional collaboration, operational transformation, Maximo, WebSphere Liberty, DB2, Hadoop, enterprise application platforms, troubleshooting application issues, middleware troubleshooting, database troubleshooting, cloud troubleshooting, infrastructure troubleshooting, SRE practices, reliability frameworks, operational excellence, Kubernetes, container platforms, cloud-native technologies, OpenTelemetry, APM platforms, observability engineering, enterprise monitoring frameworks, cloud migration, operational modernization, automation, cloud-native platforms, distributed systems, leadership skills
似たような求人
ATOMS ( Tokyo ) 38 minutes ago
GMV ( Soto ) 16 hours ago

Fidel Consulting KKからの続きを読む
Fidel Consulting KK 6 days ago
Fidel Consulting KK 18 hours ago
Fidel Consulting KK 2 hours ago

SaaS Operations Lead

企業サイトでの申請
Back to search page