وصف الوظيفة
Role Overview
L1 Data Engineer at DeepSource Technologies, Remote. You will design, build, and maintain the data architecture and infrastructure that supports the organization's data strategy, working hands-on to develop, test, and deploy reliable data solutions with scalable, efficient pipelines aligned to business requirements.
Role Purpose
To deliver scalable, cloud-native data solutions within the Microsoft Azure and Databricks ecosystem. You will take ownership of data pipelines and infrastructure, collaborating with senior engineers to ensure data quality, consistency, and alignment with organizational standards while continuously improving data platform capabilities.
Key Responsibilities
Data Pipeline & ETL Development
- Design, develop, and maintain scalable data pipelines and ETL/ELT workflows to support business intelligence and analytics use cases.
- Build and optimize data ingestion processes using Azure Data Factory and Databricks, ensuring data quality and consistency across all layers of the data platform.
- Transform and process large datasets using PySpark and Python, applying best practices for performance and maintainability.
Data Querying & Optimization
- Write and optimize complex SQL queries to support analytical reporting and data validation requirements.
Collaboration & Architecture
- Collaborate with data architects and senior engineers to implement and maintain data models aligned with organizational standards.
Monitoring & Troubleshooting
- Monitor, troubleshoot, and resolve pipeline failures and data quality issues, applying root-cause analysis to prevent recurrence.
Documentation & Continuous Improvement
- Contribute to documentation of data pipelines, data dictionaries, and engineering standards.
- Support the team in exploring and evaluating new tools and approaches to continuously improve the data infrastructure.
Qualifications & Experience
- 3+ years of professional experience in a Data Engineering or closely related role.
- Hands-on experience with Databricks, including notebook development, clusters, and job orchestration.
- Experience building and managing data pipelines with Azure Data Factory.
- Working knowledge of Azure Synapse Analytics, particularly Spark pool integration.
- Familiarity with data engineering principles including incremental loading, data lake architecture, and Delta Lake.
- Understanding of data governance and security concepts within a cloud data platform.
Skills & Competencies
- Strong proficiency in Python for data processing, transformation, and automation tasks.
- Hands-on experience with Pandas for data manipulation and PySpark for distributed data processing.
- Solid SQL skills, including query writing, optimization, and performance tuning.
Additional Information
Certification Requirement
- Candidates are expected to hold or be actively working toward the Databricks Certified Data Engineer Associate certification.
Certification Domains
- Databricks Lakehouse Platform architecture and capabilities.
- ETL and ELT workflows using Spark SQL and PySpark.
- Incremental data processing and structured streaming.
- Production pipeline development and orchestration.
- Data governance and security within the Databricks environment.
Nice to Have
- Experience with SQL Server migration projects, including schema conversion and data movement.
- Exposure to Terraform for Azure infrastructure provisioning and management.
- Familiarity with CI/CD practices applied to data engineering workflows.
- Experience with Delta Sharing or Lakehouse Federation concepts.