Описание: EPAM builds enterprise software products, open source solutions, accelerators, and data products for ML modeling teams.
Задачи
- Support a team of data engineers building pipelines used by MLOps and ML Engineers on ML modeling teams;
- Develop, optimize, and maintain data transformation pipelines in Databricks using PySpark;
- Work with data stored in ADLS Gen2 and SAP HANA Data Lake, primarily in Delta/Parquet format;
- Implement and maintain data quality checks, including schema validation, deduplication, enrichment, and tagging;
- Communicate with stakeholders to understand business processes and model input data;
- Tune performance for large-scale datasets to ensure efficient processing.
Требования
- 3+ Years of experience in data engineering with proficiency in Python, PySpark, and Databricks;
- Hands-on experience with Delta Lake and Azure data lake technologies;
- Familiarity with software version control tools such as GitHub and Git;
- Experience with CI/CD frameworks such as GitHub Actions;
- Knowledge of data lake technologies and large-scale dataset performance tuning;
- Proficiency in English at B2+ level;
- Nice to have: familiarity with at least one other programming or scripting language, such as Java, SQL, or Scala.
Условия
Remote work in Kazakhstan.