Job description
Role Overview
GSSTech Group is seeking a Data Engineer - ETL/PySpark to design, build, and support data pipelines and data marts within a banking environment. This is a hands-on role based on-site requiring full ownership of the SDLC lifecycle from build through production deployment and post-production support.
Role Purpose
To design, develop, and maintain enterprise-grade ETL pipelines and data warehousing solutions that support banking and financial reporting needs, while ensuring data pipeline reliability, scalability, and adherence to banking data governance and compliance standards.
Key Responsibilities
ETL Pipeline Development
- Design, develop, and maintain ETL pipelines and data marts using PySpark and Python
- Build and maintain data warehousing solutions supporting banking and financial reporting needs
- Work across structured, semi-structured, and unstructured data sources
- Write clean, maintainable, and production-grade Python code following software engineering best practices
Performance & Optimization
- Perform data analysis and debugging using Oracle SQL and PySpark
- Debug and optimize PySpark jobs for performance and reliability
- Ensure data pipeline reliability, scalability, and adherence to banking data governance and compliance standards
SDLC & Deployment
- Own end-to-end SDLC activities: build, UAT support, UAT bug fixes, production deployment, and post-production support
- Participate in CI/CD pipeline processes, including testing and validation of data pipelines
Collaboration & Governance
- Collaborate with cross-functional teams including QA, DBAs, and business analysts through the release cycle
Qualifications & Experience
- 5+ years of commercial experience in a data-driven engineering role
- Hands-on experience building data marts and ETL pipelines
- Proven experience across the full SDLC: build, UAT, bug fixing, deployment, and post-production support
- Prior experience with banking clients or strong banking domain knowledge
- Strong data warehousing fundamentals
Skills & Competencies
- Expert-level PySpark and Python for ETL scripting
- Strong command of Oracle SQL for data analysis and debugging
- Strong understanding of software engineering concepts and best practices for production pipelines
- Experience with Spark, Hadoop, MapReduce, and Hive
- Proficiency with Pandas data libraries
- Experience with SQL and NoSQL DBMS
- Hands-on use of Jupyter notebooks
- Knowledge of CI/CD practices and data testing and validation methodologies
Additional Information
- Nice to have: Cloud experience (AWS/Azure/GCP)
- Nice to have: Airflow or other orchestration tools
- Nice to have: Experience with regulatory and compliance reporting in banking