Lead RAG & Data Infrastructure Engineer
Build the secure retrieval and ingestion layer that lets the platform search large private document sets in seconds.
Mission
Large language models need accurate, secure, real-time context. Your job is to build the pipelines that feed them. As a Lead RAG & Data Infrastructure Engineer at LexScale iQ, you will design storage and retrieval layers for large private document sets. You will be directly responsible for data sovereignty and strict data handling within ring-fenced client networks.
Responsibilities
Architect Vector Pipelines: Design and optimize high-throughput data ingestion pipelines that clean, chunk, embed, and index massive corpuses of unstructured enterprise data.
Optimize Retrieval Performance: Implement and test advanced RAG strategies (hierarchical node parsing, hybrid keyword/vector search, re-ranking models) to drive down context latency below 500ms.
Enforce Data Sovereignty: Deploy secure vector databases inside protected cloud environments or client-managed networks, maintaining isolation from public web access.
Manage Embedding Latency: Monitor and optimize embedding compute costs and model drift, ensuring data indices remain dynamically synced with the clients’ live production databases.
Requirements
Vector Database Mastery: Definitive, hands-on experience deploying and scaling vector databases in production environments handling millions of high-dimensional vectors.
Advanced Chunking Strategy: You understand that basic token-splitting doesn’t work for complex corporate data. You must possess deep knowledge of semantic chunking, parent-child retrieval, and metadata filtering.
Infrastructure Pragmatism: Strong command of container tooling and enterprise cloud architecture. You care deeply about data privacy laws, compliance boundaries, and encrypted storage pipelines.
Department
Data Architecture
Location
Remote
Compensation
$160,000 – $210,000 USD + Production Hand-off Bonuses
