Job Openings AI Data Engineer (Document Intelligence & GenAI)

About the job AI Data Engineer (Document Intelligence & GenAI)

Key Responsibilities 
  • Design, develop, and maintain scalable document ingestion and content extraction pipelines capable of processing structured and unstructured data, including PDFs, Office documents, images, audio, and video.
  • Build and optimize AI-powered document processing solutions using OCR, Vision Language Models (VLMs), Large Language Models (LLMs), and multimodal AI technologies for intelligent content extraction.
  • Develop and enhance embedding models, semantic search capabilities, and vector databases to improve knowledge retrieval and enterprise-specific AI performance.
  • Implement knowledge ingestion, indexing, vectorization, and Retrieval-Augmented Generation (RAG) pipelines to support enterprise AI assistants and agentic AI applications.
  • Build batch and streaming data pipelines using modern data engineering frameworks and integrate data from APIs, databases, cloud platforms, and enterprise systems.
  • Collaborate with AI engineers, platform engineers, and data teams to integrate AI pipelines with cloud platforms such as Azure, Microsoft Fabric, Databricks, and Delta Lake.
  • Apply data governance, security, version control, and CI/CD best practices while ensuring reliable, scalable, and maintainable AI data pipelines.
  • Continuously evaluate and implement emerging AI, document intelligence, and data engineering technologies to improve platform performance and automation.

Key Requirements 

  • Bachelor's Degree in Computer Science, Data Engineering, Artificial Intelligence, Engineering, or a related discipline.
  • 1–3 years of experience in Data Engineering, AI Data Platforms, Document Intelligence, or related fields. Fresh graduates with relevant projects are encouraged to apply.
  • Strong programming skills in Python, SQL, and experience with PySpark or Apache Spark for large-scale data processing.
  • Hands-on experience with OCR, Vision Language Models (VLMs), Large Language Models (LLMs), prompt engineering, Text-to-Speech technologies, and AI model fine-tuning techniques such as PEFT and LoRA.
  • Experience building document processing pipelines, semantic search solutions, vector databases, and Retrieval-Augmented Generation (RAG) architectures.
  • Familiarity with cloud platforms and modern data technologies such as Microsoft Azure, Microsoft Fabric, Databricks, Delta Lake, Kafka, LangChain, LlamaIndex, or similar frameworks.
  • Understanding of CI/CD, version control, data governance, metadata management, and secure data access principles.
  • Strong analytical, problem-solving, communication, and documentation skills with the ability to work independently and collaboratively in Agile environments.