Job Openings GPU Infrastructure NOC Engineer

About the job GPU Infrastructure NOC Engineer

Pay: $75,000.00 - $140,000.00 per year

Why This Is a Great Opportunity

  • Monitor and support cutting-edge GPU and AI infrastructure powering demanding enterprise workloads.
  • Go beyond traditional NOC monitoring by building automation, improving tooling, and making operations smarter.
  • Get hands-on exposure to GPU, HPC, networking, data center infrastructure, and AI-enabled operations.
  • Play a direct role in protecting customer uptime and meeting demanding SLA commitments.
  • Help create and continuously improve runbooks, SOPs, monitoring processes, and incident-response workflows.
  • Work alongside technical teams, OEMs, data center operators, and network providers to solve real infrastructure problems.
  • Join a fast-growing environment where your ideas can directly improve how the NOC operates.
  • Receive bonus and equity opportunities in addition to competitive base compensation.

Location: Remote nationwide. This is a fully remote role supporting a 24/7 global operations environment through rotating shifts, including nights and weekends.

Note: Must have 2+ years of NOC, network operations, or infrastructure monitoring experience, hands-on GPU or HPC infrastructure experience, and Python or Bash scripting experience. Candidates must be willing to work rotating 24/7 shifts, including nights and weekends. These are hard requirements.

About Us

We are building next-generation AI infrastructure designed to give enterprise customers fast, flexible access to high-performance GPU compute. Our operations team plays a critical role in keeping these environments reliable, responsive, and continuously improving. Confidential Employer.

Job Description

  • Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments.
  • Triage, troubleshoot, and resolve incidents while maintaining SLA requirements.
  • Identify infrastructure issues early and take proactive action before they become customer-impacting incidents.
  • Escalate appropriate issues to OEMs, data center operators, network providers, or other Tier 3 partners.
  • Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work.
  • Use Python, Bash, or similar scripting tools to automate monitoring, triage, reporting, and operational workflows.
  • Explore and implement AI-enabled workflows that improve NOC speed, accuracy, and efficiency.
  • Create, maintain, and continuously improve technical runbooks and standard operating procedures.
  • Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements.
  • Track SLA and incident metrics and identify opportunities to improve reliability and response times.
  • Communicate clearly and proactively with customers and internal stakeholders during incidents.
  • Coordinate with data center operators, OEMs, network providers, and other third parties to resolve customer-impacting issues.
  • Support a 24/7 rotating operations schedule, including nights, weekends, and other assigned shifts.

Qualifications

  • 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment.
  • Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments.
  • Hands-on Python or Bash scripting experience.
  • Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms.
  • Strong troubleshooting and incident-response skills.
  • Experience working with network, compute, storage, or data center infrastructure.
  • Ability to understand technical issues quickly and communicate effectively during incidents.
  • Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement.
  • Must be comfortable working rotating 24/7 shifts, including nights and weekends.
  • Ability to work independently in a remote environment while collaborating effectively with global technical teams.
  • Additional languages beyond English are a plus.

Why You Will Love Working Here

  • Work directly with advanced GPU and HPC infrastructure rather than generic IT systems.
  • Build technical depth across AI compute, networking, data centers, monitoring, and automation.
  • Have a voice in how the NOC operates and help improve processes rather than simply following them.
  • Work with modern monitoring, automation, and AI-enabled operations tools.
  • Gain exposure to complex enterprise infrastructure and high-availability environments.
  • Remote nationwide flexibility with multiple shift options.
  • Medical, dental, and vision insurance.
  • 401(k).
  • Paid maternity and paternity leave.
  • Bonus and equity opportunities.

JPC-1916

Benefits:

  • Dental insurance
  • Paid time off
  • Retirement plan
  • Vision insurance