Job Openings
AI Data Engineer (Document Intelligence & GenAI)
About the job AI Data Engineer (Document Intelligence & GenAI)
Key Responsibilities
- Design, develop, and maintain scalable document ingestion and content extraction pipelines capable of processing structured and unstructured data, including PDFs, Office documents, images, audio, and video.
- Build and optimize AI-powered document processing solutions using OCR, Vision Language Models (VLMs), Large Language Models (LLMs), and multimodal AI technologies for intelligent content extraction.
- Develop and enhance embedding models, semantic search capabilities, and vector databases to improve knowledge retrieval and enterprise-specific AI performance.
- Implement knowledge ingestion, indexing, vectorization, and Retrieval-Augmented Generation (RAG) pipelines to support enterprise AI assistants and agentic AI applications.
- Build batch and streaming data pipelines using modern data engineering frameworks and integrate data from APIs, databases, cloud platforms, and enterprise systems.
- Collaborate with AI engineers, platform engineers, and data teams to integrate AI pipelines with cloud platforms such as Azure, Microsoft Fabric, Databricks, and Delta Lake.
- Apply data governance, security, version control, and CI/CD best practices while ensuring reliable, scalable, and maintainable AI data pipelines.
- Continuously evaluate and implement emerging AI, document intelligence, and data engineering technologies to improve platform performance and automation.
Key Requirements
- Bachelor's Degree in Computer Science, Data Engineering, Artificial Intelligence, Engineering, or a related discipline.
- 1–3 years of experience in Data Engineering, AI Data Platforms, Document Intelligence, or related fields. Fresh graduates with relevant projects are encouraged to apply.
- Strong programming skills in Python, SQL, and experience with PySpark or Apache Spark for large-scale data processing.
- Hands-on experience with OCR, Vision Language Models (VLMs), Large Language Models (LLMs), prompt engineering, Text-to-Speech technologies, and AI model fine-tuning techniques such as PEFT and LoRA.
- Experience building document processing pipelines, semantic search solutions, vector databases, and Retrieval-Augmented Generation (RAG) architectures.
- Familiarity with cloud platforms and modern data technologies such as Microsoft Azure, Microsoft Fabric, Databricks, Delta Lake, Kafka, LangChain, LlamaIndex, or similar frameworks.
- Understanding of CI/CD, version control, data governance, metadata management, and secure data access principles.
- Strong analytical, problem-solving, communication, and documentation skills with the ability to work independently and collaboratively in Agile environments.