Appealing Points

  • Full autonomy to build, from scratch, the foundational infrastructure supporting AI-agent and LLM-based decision-making products for enterprise clients
  • Direct, organization-wide impact on developer productivity by comprehensively improving CI/CD, container infrastructure, IaC, and observability
  • Opportunity to build a highly marketable career by combining Platform Engineering and Embedded SRE experience in a rapidly expanding AI and data science field

Annual Salary: ¥10,000,000 and above

Job Description

The essential mission of this position is to deliver value by working closely with the development team — establishing a foundation that allows the team to confidently deliver products, while maintaining responsibility for system reliability. Beyond simply building and operating infrastructure, the goal is to quantitatively visualize "reliability" through SLI/SLO design and observability improvements, creating a system where the development team can autonomously guarantee quality.

Job Responsibilities

Reliability Engineering

  • Integrate into each development team to support the design and implementation of SLI/SLO
  • Establish an observability platform integrating metrics, logs, and traces to quantitatively visualize system status
  • Enhance the organization's ability to learn from failures through incident response processes and a post-mortem culture
  • Actively participate in reliability-related design reviews, such as capacity planning and failure design

Platform Construction and Operation

  • Design and maintain CI/CD pipelines
  • Promote the construction and automation of container infrastructure such as Kubernetes and ECS
  • Standardize and codify infrastructure using IaC tools such as Terraform and AWS CDK
  • Improve Developer Experience (DX) through a developer portal and internal tools
  • Design and implement an evaluation platform for continuously improving the quality of AI agent and LLM applications
  • Develop common components, evaluation platforms, development/operation platforms, and internal knowledge to accelerate AI agent development

Collaboration with the Application Development Team

  • Work closely with the development team to continuously improve infrastructure while gathering on-site needs
  • Contribute to improving the development team's autonomy through documentation and knowledge sharing

Required Qualifications

  • Practical experience in web application development and operation (regardless of programming language)
  • Practical experience designing, building, and operating production infrastructure in cloud environments (AWS / GCP / Azure, etc.) — on-premises-only experience is not eligible
  • Experience primarily in infrastructure with some involvement in application development (ideally a 70% infrastructure / 30% development ratio; development-focused candidates with infrastructure build-phase experience are also acceptable)
  • Practical experience in container orchestration technology, along with designing, operating, and building observability infrastructure for SLI/SLO
  • Experience and challenges in developing and operating AI, in professional and/or personal settings
  • Experience leading development teams and collaborating with roles such as product managers, customer success specialists, and consultants
  • Experience leading development teams through requirements definition with product managers, and through technical decision-making and team management as a tech lead
  • Practical experience in incident response and incident management, including developing incident response processes and conducting post-mortems
  • Communication experience during inquiries and incident response, including customer support/escalation with customer success and consultants, and explaining incidents to stakeholders (management, customers, etc.)

Preferred Qualifications

  • Platform Engineering or Internal Developer Platform (IDP) development experience, including designing and building internal platforms to improve Developer Experience (DX)
  • AI/ML workload design and operation experience, including model deployment and monitoring, GPU infrastructure management, and ML pipeline development

Language Required - Fluent level Japanese and Business level English

About Company

Founded in Tokyo in 2012, Company is a data science firm that redefines the structure of organizational decision-making.

For over a decade, we have partnered with Japan's leading enterprise brands to decode market mechanisms, quantify competitive dynamics, and build the scientific foundation for strategic advantage. That work — rigorous, applied, and deeply embedded in client operations — distinguishes us from firms that deliver analysis and stop there.

We define our offering as "AIDE": AI Decision Engine. It connects insight to strategy to execution, embedding a continuously evolving decision-making structure into the organization itself. The goal is not better data. It is leadership that acts with conviction, and decisions that compound over time.

. Skillset Required: Platform Engineering, Embedded SRE, CI/CD, container infrastructure, Infrastructure as Code, observability, SLI, SLO design, incident response, post-mortem culture, reliability-related design reviews, capacity planning, failure design, Kubernetes, ECS, Terraform, AWS CDK, Developer Experience (DX), developer portal, internal tools, AI agent applications evaluation, infrastructure automation, cloud environments, web application development, production infrastructure design and operation, container orchestration technology, observability infrastructure, AI development and operation, team leadership, collaboration with product managers, customer success, consultants, technical decision-making, incident management, customer support, escalation, communication, Internal Developer Platform development, AI/ML workload design, model deployment, GPU infrastructure management, ML pipeline development
Similar jobs

More from Fidel Consulting KK
Fidel Consulting KK 7 days ago
Fidel Consulting KK 18 hours ago
Fidel Consulting KK 2 hours ago

Senior Software Engineer

Apply On Company Site
Back to search page