وصف الوظيفة
Role Overview
Principal Core Infrastructure Engineer (AI/ML Forward Deployed Infrastructure Engineer) at Oracle. This is an IC4 level position responsible for designing, implementing, and maintaining infrastructure that supports customer AI and machine learning initiatives.
Role Purpose
Design, implement, and maintain AI/ML infrastructure that supports customer deployments at scale. Work hands-on with customer technical teams and internal cloud services to ensure AI/ML solutions are deployed efficiently, securely, and reliably while optimizing for performance and cost-effectiveness.
Key Responsibilities
Design & Architecture
- Lead the design and architecture of AI/ML infrastructure, considering performance, scalability, and security.
- Stay updated with the latest AI/ML technologies and trends, driving innovation within the team.
Implementation & Operations
- Implement and maintain AI/ML infrastructure, ensuring efficient and reliable operations.
- Monitor and optimize AI/ML infrastructure performance, identifying and resolving bottlenecks.
- Troubleshoot issues for proof of concept (POC) and production deployments.
Security & Compliance
- Ensure the secure deployment of AI/ML models and data, adhering to industry best practices.
Stakeholder & Customer Engagement
- Collaborate with customer technical teams to understand their AI/ML requirements and provide tailored solutions.
- Work closely with cloud services team to integrate AI/ML solutions into the cloud platform.
- Engage with customers and stakeholders to gather feedback and ensure their satisfaction with AI/ML offerings.
- Understand the challenges in working with large customers.
- Leverage understanding of business leaders, stakeholders, and customers to ensure proposed solutions meet their needs.
Documentation & Communication
- Document and communicate infrastructure designs, ensuring clear and concise documentation.
- Convey technical concepts to non-technical stakeholders.
Leadership & Mentorship
- Provide technical guidance and mentorship to junior engineers, fostering a culture of knowledge sharing.
- Coach and mentor junior team members, fostering continuous learning and knowledge sharing within and across teams.
- Contribute to the talent development pipeline by participating in candidate interviews, assessing candidates, and providing hiring recommendations.
Strategic & Planning
- Demonstrate the ability to think strategically about business, products, and technical challenges.
- Manage and coordinate moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for moderately sized projects or initiatives.
- Efficiently delegate, monitor, and prioritize work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.
- Identify and address moderately complex issues by analyzing a wide range of data and information to identify solutions in accordance with standard practices.
- Proactively escalate unresolved or critical issues with thorough assessment and suggest potential solutions.
- Review, contribute to, and document problem-solving strategies.
Continuous Improvement
- Develop ideas, recommend updates, and collaborate on the implementation of process improvements to increase the efficiency and effectiveness of processes, protocols, and workflows across teams.
- Evaluate the impact of improvements on key stakeholders.
- Solicit feedback from others on ideas for alternative approaches and methods for continued improvement.
Collaboration
- Collaborate across the organization to align on expectations and achieve shared objectives.
- Support inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.
Professional Development
- Pursue learning opportunities to expand knowledge and skills in new areas and stay abreast of the latest industry trends and best practices.
- Proactively seek and leverage ongoing feedback and training to improve skills.
Qualifications & Experience
- Experience in scripting and automation using tools such as Ansible, Terraform, Python, and/or Kubernetes.
- Experience with containerization technologies such as Docker and Kubernetes.
- Experience with orchestration tools such as Slurm and PBS for managing distributed systems.
- Strong Linux skills with hands-on experience in Oracle Linux, RHEL, CentOS, Ubuntu, and Debian distributions, including system administration, package management, shell scripting, and performance optimization.
- Proven ability to troubleshoot complex issues and drive resolution in a fast-paced environment.
Skills & Competencies
- Strong proficiency in at least one programming language such as Python, Rust, Go, Java, or Scala.
- Solid understanding of networking concepts, security principles, and best practices.
- Excellent problem-solving skills.
- Strong communication and collaboration skills with the ability to work effectively in cross-functional teams.
- Strong documentation skills with experience documenting infrastructure designs, configurations, procedures, and troubleshooting steps to facilitate knowledge sharing, ensure maintainability, and enhance team collaboration.
- Understanding of machine learning frameworks and libraries such as TensorFlow, PyTorch, or scikit-learn and their deployment in production environments.
- Familiarity with DevOps practices and tools for continuous integration, deployment, and monitoring such as Jenkins, GitLab CI/CD, and Prometheus.
- Strong experience with High-Performance Computing and GPU systems.
Additional Information
- Proven experience designing, implementing, and managing infrastructure for AI/ML or HPC workloads is preferred.