Role overview
About this role
At IBM Software, we transform client challenges into solutions. Building the world’s leading AI-powered, cloud-native products that shape the future of business and society. Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You’ll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation. With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM’s product and technology landscape. Here, you’ll have the tools and opportunities to advance your career while creating software that changes the world. As a Site Reliability Engineer, you will play a crucial role in supporting, maintaining, and operationally improving the cloud infrastructure. Working closely with various teams, your focus will be on ensuring the health and reliability of production and test systems. Your proactive approach will be essential in responding promptly to issues and alerts, contributing to the development of new capabilities, and collaborating with other SRE teams and program managers to deliver mission-critical services to the market. Key Duties: • 24x7 System Monitoring: Monitor the health of production and test systems around the clock, ensuring continuous reliability. • Rapid Issue Response: Respond promptly to production issues and alerts, providing swift resolution and maintaining system availability. • Capability Development: Support the development of new and existing capabilities for compute, storage, and network services. • Collaborative Partnership: Partner with other SRE teams and program managers, contributing to the seamless delivery of mission-critical services to the market. • Automation Execution: Execute changes in the production environment through automation, ensuring efficiency and minimizing downtime. • Cross-Functional Troubleshooting: Collaborate with engineering teams to provide initial assessments and possible workarounds for production issues. Troubleshoot and resolve production issues effectively. • Integration Planning: Work with support and development teams to identify and resolve issues. Discuss and plan integration tasks to enhance overall system performance. • System Monitoring and Troubleshooting: Strong skills in system monitoring, issue response, and troubleshooting for optimal system performance. • Automation Proficiency: Proficiency in automation for production environment changes, streamlining processes for efficiency. • Collaborative Mindset: Collaborative mindset with the ability to partner seamlessly with cross-functional teams for shared success. • Effective Communication Skills: Excellent communication skills, essential for effective integration planning and swift issue resolution.