Senior Engineer – AI Platform

 Job Description:

Role Overview We are looking for a hands-on Senior Engineer to design, build, and operate an enterprise AI platform for a digital bank. Our goal is to provide a fully governed, safe, and efficient AI environment that staff actually want to use. You will take ownership across three core areas: an enterprise-grade inference gateway, identity and distribution systems, and self-hosted GPU inference clusters.

Core Responsibilities

Inference Serving & On-Premise GPU Operations

  • Size hardware capacity and collaborate with infrastructure teams to handle hypervisors, rack/power setups, and GPU passthrough across bare-metal or data center environments.
  • Deploy, operate, and fine-tune model serving frameworks using techniques like continuous batching, quantization, tensor parallelism, and KV cache optimization.
  • Implement tenancy admission controls, quota management, queuing, and failover mechanisms to cloud inference when local capacity is saturated.
  • Track and publish key performance indicators—such as p95 time-to-first-token, tokens per second, and cost per million tokens—to evaluate fixed OpEx break-even points against cloud alternatives.

Governed Gateway Architecture & Identity Integration

  • Construct and maintain an inference gateway that handles rate limiting, per-user key management, model allowlists, and strict zero data retention policies.
  • Connect governance systems directly to corporate single sign-on (OIDC / OAuth2) and automate onboarding, offboarding, and role claims via SCIM.
  • Ensure total financial transparency by attributing real-time API spend, token usage, and operational costs down to individual users and departments.

Tooling Distribution & User Enablement

  • Package, distribute, and manage agent configurations and execution harnesses across a diverse enterprise machine fleet.
  • Maintain a secure, CI-driven registry for versioned skills, knowledge-base bundles, and agent modules featuring seamless rollback capabilities.
  • Onboard, train, and support non-technical teams (such as risk, finance, and operations) to maximize productivity using approved AI tools.

What We Are Looking For

Essential Qualifications & Skills

  • Production Platform Engineering: 5+ years of experience managing critical production infrastructure, with direct involvement in incident response and on-call rotations.
  • Core Technical Stack: Proficiency in Python (primary) and TypeScript, alongside hands-on experience with Kubernetes, Linux, containers, Infrastructure as Code, and CI/CD pipelines.
  • Enterprise Identity Systems: Strong grasp of JWT claims, OIDC, OAuth2, SCIM protocols, and enterprise directory integrations.
  • Gateway & Multi-Tenancy Architecture: Background in building or maintaining API/LLM gateways, rate limits, quota systems, usage metering, and multi-tenant isolation.
  • On-Premise Infrastructure Comfort: Experience navigating capacity planning, hypervisors, hardware passthrough, and enterprise change management processes.
  • FinOps & Attribution Mindset: Ability to evaluate platforms through unit economics, capacity utilization, and cost attribution.
  • Technical Enablement: Skill in simplifying complex AI tools and training non-technical colleagues to use them effectively.

Preferred / Bonus Qualifications

  • Experience with Model Context Protocol (MCP), sovereign hosting constraints, model evaluation, or highly regulated financial services environments.
  • Hands-on expertise tuning high-throughput GPU model serving stacks such as vLLM, SGLang, Triton, or TGI.

Platform Guiding Principles

  • Usability First: Approved pathways must offer a better user experience than unapproved workarounds to prevent control bypasses.
  • Attribution Before Scale: Total visibility into spend and usage per user is required before expanding system scale.
  • Allocated Capacity: GPU resources are scarce and access is governed through explicit capacity allocation policies.
  • Pragmatic Ownership: Open-source or licensed solutions are used for standard components, while internal focus is dedicated to security and governance boundaries.
  Required Skills:

JWT Environment Hardware Data Center Incident Response Performance CI/CD pipelines Data Cloud API Support Access AI Usability Transparency Financial Services Pipelines Operations Ownership User Experience Onboarding CI/CD Components Architecture Change Management Optimization Infrastructure Economics Integration Kubernetes TypeScript Security Linux Finance Planning Design Engineering Training Python Management