About the job GPU Infrastructure NOC Engineer
Pay: $75,000.00 - $140,000.00 per year
Why This Is a Great Opportunity
- Monitor and support cutting-edge GPU and AI infrastructure powering demanding enterprise workloads.
- Go beyond traditional NOC monitoring by building automation, improving tooling, and making operations smarter.
- Get hands-on exposure to GPU, HPC, networking, data center infrastructure, and AI-enabled operations.
- Play a direct role in protecting customer uptime and meeting demanding SLA commitments.
- Help create and continuously improve runbooks, SOPs, monitoring processes, and incident-response workflows.
- Work alongside technical teams, OEMs, data center operators, and network providers to solve real infrastructure problems.
- Join a fast-growing environment where your ideas can directly improve how the NOC operates.
- Receive bonus and equity opportunities in addition to competitive base compensation.
Location: Remote nationwide. This is a fully remote role supporting a 24/7 global operations environment through rotating shifts, including nights and weekends.
Note: Must have 2+ years of NOC, network operations, or infrastructure monitoring experience, hands-on GPU or HPC infrastructure experience, and Python or Bash scripting experience. Candidates must be willing to work rotating 24/7 shifts, including nights and weekends. These are hard requirements.
About Us
We are building next-generation AI infrastructure designed to give enterprise customers fast, flexible access to high-performance GPU compute. Our operations team plays a critical role in keeping these environments reliable, responsive, and continuously improving. Confidential Employer.
Job Description
- Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments.
- Triage, troubleshoot, and resolve incidents while maintaining SLA requirements.
- Identify infrastructure issues early and take proactive action before they become customer-impacting incidents.
- Escalate appropriate issues to OEMs, data center operators, network providers, or other Tier 3 partners.
- Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work.
- Use Python, Bash, or similar scripting tools to automate monitoring, triage, reporting, and operational workflows.
- Explore and implement AI-enabled workflows that improve NOC speed, accuracy, and efficiency.
- Create, maintain, and continuously improve technical runbooks and standard operating procedures.
- Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements.
- Track SLA and incident metrics and identify opportunities to improve reliability and response times.
- Communicate clearly and proactively with customers and internal stakeholders during incidents.
- Coordinate with data center operators, OEMs, network providers, and other third parties to resolve customer-impacting issues.
- Support a 24/7 rotating operations schedule, including nights, weekends, and other assigned shifts.
Qualifications
- 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment.
- Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments.
- Hands-on Python or Bash scripting experience.
- Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms.
- Strong troubleshooting and incident-response skills.
- Experience working with network, compute, storage, or data center infrastructure.
- Ability to understand technical issues quickly and communicate effectively during incidents.
- Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement.
- Must be comfortable working rotating 24/7 shifts, including nights and weekends.
- Ability to work independently in a remote environment while collaborating effectively with global technical teams.
- Additional languages beyond English are a plus.
Why You Will Love Working Here
- Work directly with advanced GPU and HPC infrastructure rather than generic IT systems.
- Build technical depth across AI compute, networking, data centers, monitoring, and automation.
- Have a voice in how the NOC operates and help improve processes rather than simply following them.
- Work with modern monitoring, automation, and AI-enabled operations tools.
- Gain exposure to complex enterprise infrastructure and high-availability environments.
- Remote nationwide flexibility with multiple shift options.
- Medical, dental, and vision insurance.
- 401(k).
- Paid maternity and paternity leave.
- Bonus and equity opportunities.
JPC-1916
Benefits:
- Dental insurance
- Paid time off
- Retirement plan
- Vision insurance