Production Engineer

Hackajob
Hackajob

Product

India

Posted on Aug 5, 2026
hackajob is collaborating with HSBC to connect them with exceptional professionals for this role.

In This Role You Will

  • Provide 24/7 on-call production support for cloud-native AI platform services, including model inference service, data pipeline, AI middleware and backend system troubleshooting
  • Monitor production environment health metrics, system logs, error alerts and performance anomalies; rapidly diagnose, triage and resolve production incidents and outages.
  • Handle critical service issues during UK and US working hours, conduct root cause analysis (RCA), prepare incident reports, and implement permanent fixes to avoid recurrence.
  • Take the DevOPS lead to ensure reliability, performance, and compliance through monitoring, load/performance testing, optimisation, and incident response to maintain low latency, high throughput, and stable operations (including onsite/offsite production support).
  • Manage CI/CD pipeline configuration, optimization and daily operation for AI platform projects, ensure efficient and reliable automated build, test and deployment.
  • Optimize deployment workflow to reduce release failure rate and improve delivery efficiency for cross-border teams.
  • Establish and iterate standardized engineering operation specifications, production operation manuals, and on-call response SOPs for the AI platform.
  • Promote engineering best practices including log standardization, alert optimization, monitoring coverage enhancement, and fault tolerance improvement.

To be successful in this role you should meet the following requirements

  • Bachelor’s degree in Computer Science, Information Technology, Software Engineering or related technical field.
  • 5+ years of professional working experience in Production Support, DevOps or cloud platform operation; AI platform operation experience is a strong plus.
  • Based in overseas and able to accept flexible working hours to cover UK/US timezone support and on-call shifts.
  • Proficient in Linux system operation, troubleshooting, shell scripting and server maintenance.Experience with cloud platform operation (public cloud preferred) and AI system operation & maintenance is highly preferred.
  • Solid experience with CI/CD tools (Jenkins/GitLab CI/Azure DevOps) and Git version control management.
  • Familiar with Kubernetes container orchestration, Docker, microservice operation and maintenance.

You’ll achieve more when you join HSBC.