Agentic AI / Data Engineer - DC GPU
Offre en anglaisBuild and operate production agentic AI systems and data foundations on AMD Instinct GPU infrastructure. Develop memory layers, telemetry pipelines, and evaluation frameworks to ensure reliable and safe agent behavior.
- Sur place
- Oro-Medonte, ON
- Publié 27 août 2026
- Postuler avant le 26 sept. 2026
- 1 poste
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Forgeahead Solutions Corporation
Technical Lead and Senior Software Engineer
- Sur place
The Hilary and Galen Weston Foundation
Program Manager, Neuroscience
- Sur place
Government of Ontario
Chef des services de soins de santé
- Sur place
Résumé du poste
ADVANCE YOUR CAREER. ADVANCE THE WORLD. At AMD, we believe technology has the power to solve the world’s most important challenges. From advancing healthcare and scientific discovery to powering AI and the technologies people rely on every day, innovation at AMD is shaping the future. Whether you’re designing next-gen processors, enabling AI breakthroughs, or bringing leading edge products to market, every role at AMD contributes to something bigger — technology that moves the world forward. Join us and, together, we’ll advance your career. The Team AMD's Data Center GPU organization is transforming the industry with our AI based Graphic Processors. Our primary objective is to design exceptional products that drive the evolution of computing experiences, serving as the cornerstone for enterprise Data Centers, (AI) Artificial Intelligence, HPC and Embedded systems. If this resonates with you, come and joining our Data Center GPU organization where we are building amazing AI powered products with amazing people. The Role AMD's Applied AI team works with the world's most demanding AI operators — frontier labs, NeoCloud providers, and AI-native companies — to make AMD Instinct GPU infrastructure the easiest place to build and run AI. As an Agentic Data Engineer, you will build the data and agent systems that sit at the heart of this mission: production agentic AI applications running on AMD clusters, the data pipelines and memory/context databases that give those agents durable knowledge, and the skills frameworks and evaluation infrastructure that make agent behavior reliable, measurable, and safe. Your work spans two surfaces. Externally, you build agentic systems and their data foundations on customer AMD deployments — the reference implementations customers adopt when they move from inference to agents. Internally, you build the Applied AI team's own intelligence layer: engagement memory databases, fleet and telemetry data pipelines, and agent-executable skills libraries that encode deployment knowledge so every customer engagement makes the next one faster. This is a production engineering role. The systems you build run live, get depended on, and are held to production standards for quality, provenance, and security. The Person You are equal parts data engineer and applied AI engineer. You think about agents as data systems: what context they retrieve, what memory they accumulate, what tools they invoke, and how you would prove they behave correctly. You have shipped pipelines that other teams depend on and LLM applications that real users hit, and you know the difference between a demo agent and one that survives production. You hold strong opinions about context engineering, memory store design, and evaluation — and you can defend them with data. Key Responsibilities Build production agentic AI systems on AMD Instinct GPU infrastructure: agent orchestration, tool/function calling (including MCP-based integrations), skills frameworks, and streaming inference integration against ROCm-based serving stacks (vLLM, SGLang) Design and operate the memory and context data layer for agentic applications: vector, graph, and relational stores, embedding pipelines, retrieval and context-engineering strategies, and the freshness, provenance, and access-control policies that govern them Build the Applied AI team's engagement memory and fleet data infrastructure: pipelines that ingest deployment telemetry, incident histories, and field knowledge into structured, queryable, agent-consumable form Develop and maintain the skills library: reusable, versioned, agent-executable encodings of deployment and operational expertise, with the testing and review gates required before agents or engineers rely on them Build evaluation infrastructure for agentic systems: regression suites, LLM-as-judge pipelines, behavioral test harnesses, and production quality monitoring Harden agentic systems against real-world failure modes, including prompt injection through retrieved context and memory stores, data poisoning, and tool-misuse paths Create the reference architectures and open artifacts that make AMD the credible platform for agentic workloads, contributing upstream to the open-source agent, serving, and data ecosystem Partner with customer-facing engineers on live engagements: your systems deploy into customer environments, and you support their production behavior Preferred Experience 5+ years of software engineering with significant production data engineering: pipelines, storage systems, and data quality at scale (level flexible for exceptional candidates) Hands-on experience building LLM-powered and agentic applications in production: agent frameworks and orchestration, RAG and context engineering, tool calling, and multi-step workflows Depth in at least one memory/context storage paradigm — vector databases, graph databases, or hybrid retrieval architectures — and informed opinions about when each is wrong Experience designing evaluation frameworks for non-deterministic systems Strong Python; working fluency with modern data stack tooling (orchestration, streaming, warehouse/lakehouse) and containerized deployment on Kubernetes Familiarity with GPU inference serving (vLLM, SGLang, or comparable) and the performance characteristics of LLM workloads; ROCm/AMD Instinct experience a strong plus Security-conscious engineering instincts, particularly around untrusted content flowing into model context Open-source contribution history in the AI/ML or data infrastructure ecosystem is a plus Preferred Academic Credentials Bachelor's or Master's degree in Computer Science, Computer Engineering, Data Engineering, or equivalent practical experience This role is not eligible for visa sponsorship. Benefits offered are described: AMD benefits at a glance. AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law. We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process. AMD may use Artificial Intelligence to help screen, assess or select applicants for this position. AMD’s “Responsible AI Policy” is available here. This posting is for an existing vacancy.
Ce que vous ferez
Build and operate production agentic AI systems and data foundations on AMD Instinct GPU infrastructure. Develop memory layers, telemetry pipelines, and evaluation frameworks to ensure reliable and safe agent behavior.
Exigences
Requires 5+ years of software engineering experience with a focus on production data pipelines and LLM-powered applications. Proficiency in Python, modern data stack tooling, and experience with GPU inference serving is expected.
Compétences indiquées
- KubernetesSouhaitée
- PythonSouhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- Agentic AI
- Data Engineering
- Python
- LLM Orchestration
- RAG
- Vector Databases
- Graph Databases
- Kubernetes
- vLLM
- SGLang
- ROCm
- Context Engineering
- Evaluation Frameworks
- Streaming Pipelines
- GPU Inference Serving
- Prompt Injection Mitigation
Domaines d’emploi
- Software
- Data & Analytics
- Engineering
- Technology
- Manufacturing
Renseignements supplémentaires
- Formation minimale
- Baccalauréat
- Expérience minimale
- 5+ ans
- Postuler avant le
- 26 sept. 2026
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine
- Niveau d’expérience
- Mid-Senior level