Job Openings
Senior AI Platform Operations Engineer
About the job Senior AI Platform Operations Engineer
Key Responsibilities:
- Monitor availability, detect outages, and optimize performance of Azure AI cloud platform
- Ensure resilience and efficiency of RE:AI (OpenShift-powered) platform
- Lead incident response, root cause analysis, and disaster recovery planning
- Manage cybersecurity operations: IAM, SIEM/SOAR, vulnerability management, access control
- Support audits, compliance, and regulatory alignment
- Collaborate with developer teams to embed monitoring, automation, and security into AI/ML workflows
- Drive operational excellence through automation and observability
Key Requirements:
- Degree in Computer Science/Engineering
- 4–6 years in cloud operations/administration
- Strong expertise in Azure monitoring (Monitor, Log Analytics, App Insights)
- Hands-on with OpenShift observability and performance tuning
- Experience in incident management, SRE practices, disaster recovery
- Cloud security operations (IAM, SIEM/SOAR, firewalls, EDR)
- Proficiency in IaC (Terraform, Bicep, ARM) and scripting (PowerShell, Python)
- Familiarity with AI/ML infra (AKS, GPU VMs, pipelines, model hosting)
- Knowledge of compliance frameworks (ISO 27001, CIS, NIST)
- Excellent problem-solving and leadership in high-pressure scenarios