About the job Lead Infrastructure & DevOps
1. Infrastructure & OS Management
- Manage and maintain Linux-based systems (Rocky Linux, CentOS, RHEL) in production and staging environments.
- Perform kernel patching, security updates, and system hardening following industry best practices.
- Troubleshoot and debug complex OS-level issues (performance, memory, I/O, kernel panics).
- Manage and support Windows Server environments where applicable.
2. Security & Compliance
- Perform regular security scanning and vulnerability assessments using tools like Nessus, Qualys, or OpenVAS.
- Conduct code security scans (SAST/DAST) and integrate them into CI/CD pipelines.
- Configure and maintain LDAP authentication for centralized access management.
- Ensure all systems comply with internal security policies and external regulatory standards.
- Implement and manage WAF (Web Application Firewall) configurations to protect against OWASP Top 10 threats.
3. Automation & Configuration Management
- Write and maintain Ansible playbooks for automated provisioning, patching, and configuration management.
- Manage GitHub Enterprise repositories, branching strategies, and CI/CD workflows (GitHub Actions preferred).
- Develop automation scripts using Python, Shell, or Go to reduce toil and improve operational efficiency.
4. Application & Web Server Management
- Configure and maintain Apache and Nginx web servers, including performance tuning and reverse proxy setup.
- Troubleshoot web server logs, SSL/TLS issues, and load balancing configurations.
5. Container Orchestration & Cloud-Native Technologies
- Deploy, manage, and scale applications on Kubernetes clusters (EKS, AKS, or on-prem).
- Build and maintain Docker images with secure base layers and minimal attack surfaces.
- Monitor cluster health, resource usage, and implement auto-scaling policies.
6. Monitoring & Incident Management
- Deploy and maintain monitoring tools (Prometheus, Grafana, Datadog, or ELK stack).
- Define SLIs, SLOs, and error budgets; drive blameless post-mortems.
- Participate in on-call rotations and lead incident response efforts.
What Qualifies You
- 8+ years in SRE, DevOps, or Systems Engineering roles.
- Experience with Terraform or other IaC tools.
- Familiarity with service mesh (Istio/Linkerd) and CNI plugins (Calico/Cilium) or similar.
- Understanding of zero-trust networking and SSO integrations.
- Certifications: RHCE, CKA, or CISSP are a plus.
OS: Expert in Rocky Linux, CentOS, RHEL (6+ years); Windows Server experience
Security: WAF, vulnerability scanning, LDAP, OS hardening, kernel patching
Automation: Ansible (playbooks, roles, AWX/Tower), Python/Shell scripting
CI/CD/SCM: GitHub Enterprise administration, GitHub Actions, branch protection rules
Containers: Docker (build/security scanning), Kubernetes (deployment, networking, storage)
Web Servers: Apache, Nginx (configuration, SSL, reverse proxy)
Monitoring: Prometheus, Grafana, ELK, or commercial tools (Datadog/New Relic)
Code Security: SAST/DAST tools (SonarQube, Snyk, Checkmarx)
Soft Skills & Cultural Fit
- Strong communication and documentation skills.
- Ability to mentor junior engineers and conduct technical interviews.
- Collaborative mindset with DevOps/DevSecOps philosophy.
- Calm under pressure during incidents; drive for continuous improvement.
XTIUM is an equal opportunity employer.