Hong Kong, Hong Kong SAR, Hong Kong

Site Reliability Engineer – SRE / Data & Platform

 Job Description:

We are looking for a Site Reliability Engineer with strong production reliability and observability experience, together with hands-on exposure to data pipelines, streaming or data platforms.

You will work in a fast-paced technology environment supporting high-performance production systems, with a focus on reliability, monitoring, automation and system performance.

Key Responsibilities

  • Own and improve the reliability, availability and performance of production systems.
  • Build and enhance observability, monitoring, dashboards and alerting for critical services.
  • Troubleshoot production issues, perform root-cause analysis and drive reliability improvements.
  • Support Kubernetes/cloud-based infrastructure and production workloads.
  • Develop and maintain CI/CD and GitOps workflows.
  • Work with data pipelines, streaming workloads or data platform infrastructure to support reliable and scalable data processing.
  • Monitor data flow, system performance and potential issues such as latency, failures, capacity constraints and data loss.
  • Develop automation and operational tools using Python or other programming/scripting languages.
  • Participate in production support and on-call activities.

Requirements

  • 3–7 years of experience in SRE, DevOps, Platform Engineering, Production Engineering or a related role.
  • Strong hands-on experience with production systems, monitoring/observability and incident troubleshooting.
  • Must have practical exposure to data-related infrastructure, such as:
    • Data pipelines / data ingestion
    • Streaming platforms
    • ETL / ELT
    • Kafka / Kinesis or similar technologies
    • Spark / Flink or similar processing frameworks
    • Data platforms / data lakes / OLAP databases
  • Experience with Kubernetes and cloud environments such as AWS, GCP or Azure.
  • Experience with CI/CD, GitOps or infrastructure automation.
  • Proficiency in Python or another programming/scripting language.
  • Good understanding of system performance, reliability and troubleshooting.
  • Strong communication and problem-solving skills.
  • Hong Kong-based candidates are highly preferred.
  Required Skills:

Performance Production Support Data Cloud Support Data Processing ETL Pipelines Spark Kafka Gcp Analysis CI/CD Azure Reliability Availability DevOps Infrastructure Automation Programming AWS Kubernetes Databases Troubleshooting Engineering Python Communication