Job Openings JR-193050 Infrastructure Platform Engineer

About the job JR-193050 Infrastructure Platform Engineer

Role Overview:

We are looking for Infrastructure Platform Engineers. The ideal candidate will own and manage everything below the application layer, including compute infrastructure, container orchestration, RAG pipelines, CI/CD, observability, load testing, and security.

The candidate should have strong experience building, deploying, and operating production-grade AI and LLM platforms. They must be capable of managing multiple infrastructure initiatives simultaneously while supporting different product teams, ensuring scalability, reliability, security, and operational excellence across the platform.

Location: Bangalore/Pune (Onsite / Offshore)

We are looking for two resources:

  • 1 Onsite resource.
  • 1 Offshore resource (the offshore candidate can be an Infrastructure Platform Engineer with backend engineering skills).

Responsibilities

  • Design, build, and manage scalable infrastructure for AI/LLM platforms.
  • Own Kubernetes, container orchestration, cloud infrastructure, and CI/CD pipelines.
  • Deploy and maintain RAG pipeline infrastructure and GPU-based AI workloads.
  • Implement Infrastructure as Code (Terraform) and GitOps practices.
  • Ensure platform reliability through monitoring, observability, incident response, and on-call support.
  • Optimize platform performance, scalability, security, and infrastructure costs.
  • Drive InfoSec compliance, security reviews, and infrastructure governance.
  • Collaborate with engineering teams to deliver reliable, production-ready infrastructure across multiple products.

Experience Requirements

  • 5–8 years of experience in Platform Engineering, Site Reliability Engineering (SRE), or Infrastructure Engineering roles.
  • Strong infrastructure background; candidates with primarily full-stack experience and limited infrastructure ownership will not be considered.
  • Hands-on experience deploying and managing infrastructure for at least one AI or LLM product in production.
  • Proven ownership of RAG pipeline infrastructure rather than only experimentation with LLM technologies.
  • Experience managing competing infrastructure priorities across two or more product teams simultaneously.
  • Ability to self-organize and execute across parallel initiatives without dedicated project management support.
  • Experience serving as the primary technical owner during InfoSec or security reviews for customer-facing products.
  • Ability to prepare technical evidence, answer security-related questions, and independently drive reviews to completion.
  • Hands-on experience supporting production services with SLAs, incident management, on-call rotations, and post-mortem analysis.

Technical Skills

Container Orchestration & Runtime

  • Kubernetes / EKS: Production cluster operations, workload management, and RBAC administration. (Must-have)
  • Docker: Container image creation, multi-stage builds, and lifecycle management. (Must-have)
  • Helm: Chart authoring, templating, and release management. (Must-have)

Cloud Platforms

Required: Azure

Preferred: AWS

  • Azure Key Vault: Secrets management, certificates, and managed identity binding. (Must-have)
  • Azure Monitor and Log Analytics Workspace. (Must-have)
  • Azure DevOps Pipelines: YAML pipelines, service connections, and approval workflows. (Must-have)
  • Azure Kubernetes Service (AKS). (Nice to have)
  • Azure API Management: Gateway configuration, OAuth 2.0, and rate limiting. (Nice to have)

Infrastructure as Code

  • Terraform: Modules, state management, and remote backends. (Must-have)
  • GitOps tools such as ArgoCD or Flux CD, including the app-of-apps pattern. (Must-have)
  • Pulumi (Python or TypeScript). (Nice to have)

AI / ML Infrastructure

  • Vector database operations, including provisioning, indexing, and tuning using technologies such as Milvus, Weaviate, pgvector, or FAISS. (Must-have)
  • GPU compute provisioning and management in data center environments. (Must-have)
  • RAG pipeline infrastructure, including ingestion, chunking, and embedding pipelines. (Must-have)
  • LLM inference infrastructure, including serving, scaling, and latency optimization using tools such as NIM and Triton. (Must-have)
  • NIM deployment and lifecycle management. (Nice to have)

CI/CD and Deployment

  • GitHub Actions: Workflow design, reusable actions, and secrets management. (Must-have)
  • Azure DevOps Pipelines: YAML pipelines and approval processes. (Must-have)
  • Container registry management, including tagging, scanning, and promotion using ECR or ACR. (Must-have)

Observability and Monitoring

  • Datadog: Metrics, APM, logging, dashboards, and SLO management. (Must-have)
  • Elasticsearch / OpenSearch: Indexing, querying, and ILM policies. (Must-have)
  • OpenTelemetry: Instrumentation, collectors, and exporters. (Must-have)
  • Distributed tracing using Lightstep or equivalent solutions. (Must-have)
  • Azure Monitor and Log Analytics integration. (Must-have)
  • Incident response and on-call management using PagerDuty or equivalent tools. (Must-have)

Performance, Security & FinOps

  • Load testing using tools such as k6, Locust, or JMeter, including scenario design and results analysis. (Must-have)
  • Security posture management, including secrets hygiene, network isolation, and RBAC auditing. (Must-have)
  • InfoSec review readiness, controls documentation, and evidence packaging. (Must-have)
  • Python scripting for infrastructure automation and ingestion pipelines. (Nice to have)