Data Development Intern
Offre en anglaisThe intern will design, build, and maintain automated data pipelines for ETL, modeling, and quality control across demographic data products. They will also assist in migrating legacy workflows to modern architectures like Snowflake while utilizing AI coding tools to accelerate development.
- Hybride
- Toronto, ON
- Publié 10 août 2026
- 1 poste
D’autres postes auxquels postuler directement
Des possibilités semblables publiées par des employeurs qui recrutent sur Jobs.ca, sans formulaire externe.
Forgeahead Solutions Corporation
Technical Lead and Senior Software Engineer
- Sur place
Desjardins
Analyste d'affaires système(BSA) Guidewire
- Hybride
BC Public Schools
Manager, Physical and Psychological Safety
- Hybride
Résumé du poste
Role Objective: The Data Development Intern plays a central role in building, maintaining, and modernizing the data infrastructure that powers Environics Analytics' core demographic and behavioural data products. The Data Development team works across a wide range of data sources, including Statistics Canada, IRCC, CRA, and third-party survey data, applying rigorous ETL, quality control, and modeling pipelines to produce market-ready outputs from national to small-area geographies. You'll design and implement automated data pipelines in SQL and Python, support the migration of legacy workflows to modern architecture (including the team's move to Snowflake), and contribute to quality control systems that ensure the accuracy and consistency of our data products across vintages. This is a hands-on role with real ownership of production code, and strong performers will be well positioned for a full-time Data Engineer role on the team. What You'll Do: Design, build, and maintain automated data pipelines for ETL, modelling, and quality control across demographic data products. Help migrate and refactor legacy workflows into SQL (T-SQL) and Python, improving scalability, maintainability, and version control. Develop stored procedures, temp table-based workflows, and batch scripts to support large-scale data transformation. Build automated QC checks and validation logic to catch anomalies and inter-vintage inconsistencies early in the pipeline. Collaborate with data developers, Research Associates, and Technical Leads to translate data product methodology into reliable, repeatable code. Present design approaches before building, validate results after, and participate in code reviews. Use Azure DevOps and Git for version control and work item tracking; maintain documentation on SharePoint. Investigate and prototype new tools, libraries, or pipeline architectures, including Snowflake-native approaches, that improve team efficiency or product quality. Use AI coding tools (e.g., GitHub Copilot) as a core part of daily development to accelerate scripting, refactoring, and code review. Apply AI-assisted approaches to documentation and QC, such as generating test cases, drafting validation logic, or summarizing pipeline behaviour. Critically evaluate AI-generated code and output, verifying correctness and understanding the underlying SQL/Python well enough to own what ships. What You'll Learn: Practical, production experience in data engineering, automation, and AI-assisted development. Exposure to large-scale demographic, financial, and behavioural datasets. Insight into the full product development lifecycle at a leading data and analytics firm. Modern cloud data warehousing (Snowflake) alongside traditional SQL Server workflows. Agile development, version control, and code review practices. Best practices in quality control and data integrity at scale. Qualifications: Education Enrolled in or recently completed a graduate program (Master's) in Computer Science, Data Science, Statistics, Geography, Engineering, or a related quantitative field. Undergraduate candidates with strong relevant experience will also be considered. Experience Prior experience (coursework, research, co-op, or work) in data engineering, data analysis, or software development. Comfort working with large-scale structured datasets (millions of rows across related tables). Technical Skills Strong SQL, including window functions and set-based transformation logic; T-SQL experience is a plus. Proficiency in Python for data processing and automation, including pandas. Experience building or contributing to multi-step ETL pipelines. Comfort working in VS Code, Jupyter Notebook, and/or SQL Server Management Studio. Experience with Git and a willingness to learn Azure DevOps. Comfort using AI coding tools (e.g., GitHub Copilot) as part of your regular workflow. Bonus Skills Familiarity with Snowflake or other cloud data warehousing. Familiarity with ETL processes and APIs. Exposure to geospatial data or Canadian census geographies (e.g., DA, CT, CSD, CMA). Familiarity with dashboards, data visualization, or statistical concepts (imputation, aggregation, index construction). Exposure to workflow orchestration tools (e.g., Airflow) or distributed computing (e.g., Dask). Personal Attributes Strong problem-solving skills and eagerness to learn; comfortable identifying root causes and proposing systematic fixes. Detail-oriented, with good documentation and communication habits. Collaborative and open to feedback; comfortable working in a multidisciplinary team of researchers and data professionals. Able to clearly communicate technical findings to both technical and non-technical stakeholders. About Environics Analytics Environics Analytics (EA) is a marketing services company that specializes in geodemographic-based segmentation, site evaluation modelling, and custom analytics. EA is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. If you require any accommodation to participate in the hiring process, please note the request in your application. We welcome people of all abilities.
Ce que vous ferez
The intern will design, build, and maintain automated data pipelines for ETL, modeling, and quality control across demographic data products. They will also assist in migrating legacy workflows to modern architectures like Snowflake while utilizing AI coding tools to accelerate development.
Exigences
Candidates should be enrolled in or have recently completed a graduate program in a quantitative field, though strong undergraduate candidates are considered. Proficiency in SQL and Python, along with experience in data engineering or analysis, is required.
Compétences indiquées
- SQLSouhaitée
- Analyse de donnéesSouhaitée
- GitSouhaitée
- PythonSouhaitée
Autres compétences pertinentes
Relevées dans la description du poste. Confirmez les exigences importantes ci-dessus.
- SQL
- Python
- Data Engineering
- ETL
- Snowflake
- T-SQL
- Pandas
- Git
- Azure DevOps
- Data Modeling
- Quality Control
- Data Analysis
- Software Development
- GitHub Copilot
- Geospatial Data
- Statistical Concepts
- Dask (Software)
- VS Code
- Pipelines
- Workflow Management
- Apache Airflow
- Git (Version Control System)
- AI-Generated Code
- Snowflake (Data Warehouse)
- Willingness To Learn
- Research
- Application Programming Interface (API)
- Agile Methodology
- Artificial Intelligence
- Automation
- Dashboard
- Version Control
- Code Review
- Communication
- Computer Science
- Data Processing
- Data Infrastructure
- Data Integrity
- Extract Transform Load (ETL)
- Data Transformation
- Data Visualization
- Data Warehousing
- Demography
- Distributed Computing
- Marketing
- Scalability
- Problem Solving
- Python (Programming Language)
- SQL Server Management Studio
- New Product Development
Domaines d’emploi
- Data & Analytics
- Technology
- Software
- Engineering
- Science & Research
- Development Intern
- Python Developer
- Software Developers
Renseignements supplémentaires
- Formation minimale
- Baccalauréat
- Expérience minimale
- 0+ ans
- Langue de l’offre
- anglais
- Heures de travail
- 40 heures par semaine