Job Openings JR-193246 AI Application Engineer

About the job JR-193246 AI Application Engineer

Project Description:

Design, build and operate production-grade LLM applications powering intelligent enterprise products.

Location: Bangalore or Pune

Responsibilities include:

  • Design and implement end-to-end LLM application pipelines using modern orchestration frameworks.
  • Own prompt engineering, model quality, RAG accuracy and AI safety across multiple AI applications.
  • Build and optimize multi-step LLM workflows including intent classification, query rewriting, retrieval, reranking and response synthesis.
  • Develop scalable Retrieval-Augmented Generation (RAG) systems with high retrieval precision and low latency.
  • Implement production-grade AI guardrails, jailbreak protection, prompt injection mitigation and hallucination reduction strategies.
  • Design automated evaluation pipelines for prompts, retrieval quality and model performance.
  • Optimize latency, throughput and cost across complex LLM call chains.
  • Integrate NVIDIA AI technologies including NIM, NeMo, NeMo Guardrails and NVIDIA Riva.
  • Build streaming AI APIs and conversational experiences supporting multi-turn interactions.
  • Collaborate with platform, backend and product engineering teams to deliver enterprise AI solutions.
  • Monitor production AI systems, identify quality degradation and continuously improve model performance.

Requirements:

  • 4+ years of software engineering experience.
  • At least 2 years building production LLM-powered applications.
  • Expert-level Python development.
  • Strong experience with LLM prompt engineering and prompt optimization.
  • Experience designing production RAG architectures.
  • Experience with LangChain, LlamaIndex or custom orchestration frameworks.
  • Strong understanding of embeddings, vector databases and retrieval optimization.
  • Experience implementing AI safety, guardrails and prompt injection protection.
  • Experience designing automated AI evaluation pipelines (RAGAS, TruLens or equivalent).
  • Experience optimizing latency and quality across multi-stage LLM pipelines.
  • Experience developing streaming AI APIs.
  • Strong understanding of REST APIs and asynchronous Python development.
  • Experience deploying AI applications using NVIDIA NIM and NeMo technologies.
  • Excellent debugging, analytical and communication skills.
  • Strong English communication skills.