Data Engineer
- Brampton, ON
- Sur place
- Publié 27 août 2026
- 1 poste
Ouvre un site externe
- Type d’emploi
- Contrat
- Niveau d’expérience
- Expérimenté · 5+ ans
- Postuler avant le
- 26 sept. 2026
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Mid-Senior level
- Mode de candidature
- La candidature directe est offerte
Résumé du poste
Architect and build ETL pipelines for structured and unstructured data to support AI agents. Manage vector database infrastructure and implement real-time event-driven architectures for embedding updates.
Détails du poste
You will architect ETL pipelines, manage vector databases, enforce governance, and build real-time data flows that continuously update embeddings and indexes. Your work ensures AI agents operate with fresh, trustworthy information. What You Will Do • Data Pipelines: Build ETL flows for structured/unstructured data, ensuring normalization, deduplication, and semantic consistency. • Vector Infrastructure: Manage pgvector, Azure AI Search, Redis vector indexing, and hybrid search layers. • Data Governance: Implement zero-trust access, privacy controls, and compliance within AI context pipelines. • Real-time Processing: Build event-driven architectures that continuously refresh embeddings and indexes. Required Qualifications • Deep experience with distributed data systems, SQL, and orchestration tools. • Experience tuning high-throughput database infrastructure. • Knowledge of Google’s GECX is a plus. • Familiarity with chunking strategies and embedding models. Skillset Requirements • ETL & Data Modeling: Designing pipelines for structured/unstructured data, normalization, deduplication, and semantic consistency. • Vector Databases: pgvector, Redis, Azure AI Search, hybrid search, and index optimization. • Distributed Data Systems: Kafka, Spark, Flink, or similar event-driven architectures. • Data Governance: Zero-trust access, privacy controls, compliance, and auditability. • Real-time Embedding Updates: Event-driven refresh pipelines for RAG and agent memory systems. • Chunking & Embeddings: Semantic chunking, metadata tagging, and embedding model selection. • Search Infrastructure: BM25, hybrid search, inverted indexes, and ranking algorithms. • Performance Tuning: High-throughput read/write optimization. • Data Quality & Lineage: Validation, schema enforcement, and lineage tracking (e.g., Great Expectations, OpenLineage).
Ce que vous ferez
Architect and build ETL pipelines for structured and unstructured data to support AI agents. Manage vector database infrastructure and implement real-time event-driven architectures for embedding updates.
Exigences
Requires deep experience with distributed data systems, SQL, and orchestration tools. Candidates should be familiar with chunking strategies, embedding models, and high-throughput database tuning.
Compétences indiquées
- Redis · Souhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- ETL
- Data Modeling
- Pgvector
- Redis
- Azure AI Search
- Kafka
- Spark
- Flink
- Data Governance
- RAG
- Semantic Chunking
- Embedding Models
- BM25
- Hybrid Search
- Performance Tuning
- Data Lineage
Domaines d’emploi
- Data & Analytics
- Technology
- Software
- Engineering
- Consulting
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Government of Ontario
Infirmier autorisé
CommanditéEmployeur directCandidature simplifiée- Sur place
- Lindsay, ON
- Publié 17 sept. 2026
Bédard Ressources Humaines
ITAD Services Representative #1265
CommanditéEmployeur directCandidature simplifiée- Sur place
- Mississauga, ON
- Publié 16 sept. 2026
BC Public Schools
Manager, Financial Planning and Analysis
CommanditéEmployeur directCandidature simplifiée- Sur place
- Victoria, BC
- Publié 18 sept. 2026