Azure Data Engineer is needed to perform the following duties:
· Design, develop, and maintain scalable data-processing applications using PySpark and Spark SQL on Azure Databricks.
· Elaboration: Develop distributed data-processing applications to cleanse, transform, standardize, and integrate large volumes of structured and semi-structured enterprise data. Build reusable processing components and scalable transformation logic to support downstream reporting, analytics, and operational applications.
· Design and implement data-ingestion processes for data originating from multiple source systems and formats. Elaboration: Develop ingestion processes using Azure Data Factory, Azure Data Lake Storage, Azure Databricks, and related services to acquire structured and semi-structured data from enterprise source systems and prepare it for downstream processing and analytics.
· Build, maintain, and optimize ETL/ELT pipelines using Azure Databricks, PySpark, Spark SQL, Azure Data Factory, ADLS, Delta Lake, and related technologies.
· Elaboration: Develop pipelines that apply business rules, filtering, joins, aggregations, standardization, and other transformation logic before making curated datasets available to downstream systems and analytical platforms.
· Analyze source and target data structures and implement data mappings, transformation rules, and integration logic.
· Elaboration: Review data structures, field definitions, business rules, relationships, and downstream requirements to map source information into standardized target structures and integrated enterprise datasets
· Implement data-validation, reconciliation, error-handling, and data-quality controls throughout ingestion and transformation processes.
· Elaboration: Establish validation rules to identify missing, invalid, duplicate, inconsistent, or unexpected data and perform reconciliation between source and target datasets to ensure accurate and reliable processing.
· Design, schedule, and maintain automated data workflows and pipeline-orchestration processes.
· Elaboration: Configure workflow dependencies, execution sequences, scheduling, monitoring, retries, and failure-handling mechanisms using Azure Data Factory, Databricks Workflows, or comparable orchestration technologies.
· Monitor, troubleshoot, and optimize distributed data-processing applications and cloud-based data pipelines.
· Elaboration: Analyze execution behavior, resource utilization, processing bottlenecks, failed tasks, inefficient transformations, and dependencies to improve performance, scalability, reliability, and operational efficiency.
· Support modernization, migration, and integration of legacy data-processing solutions into cloud-based Azure platforms.
· Elaboration: Analyze existing data-processing workflows and redesign or migrate them into scalable Azure-based solutions while maintaining required business logic, dependencies, and downstream interfaces.
· Collaborate with business analysts, data analysts, data scientists, application teams, and other stakeholders to understand data requirements and provide data for reporting, analytics, and operational decision-making.
· Elaboration: Translate business requirements into technical data-processing and transformation requirements and ensure that curated datasets accurately represent source-system data and downstream analytical needs.
· Apply data-management standards, client-specific business rules, security requirements, and technical documentation practices when implementing enterprise data solutions. Elaboration: Maintain documentation of data mappings, transformation rules, pipeline architecture, workflow dependencies, and technical procedures while applying applicable organizational, governance, and data-management standards.
Bachelor's Degree is required in Computer Science or Data Science or Industrial Engineering
© 2021 Intellilink Technologies. all rights reserved.
Developed by Intellilink Technologies.