Role overview
About this role
At IBM Software, we transform client challenges into solutions. Building the world’s leading AI-powered, cloud-native products that shape the future of business and society. Our legacy of innovation creates endless opportunities for IBMers to learn, grow, and make an impact on a global scale. Working in Software means joining a team fueled by curiosity and collaboration. You’ll work with diverse technologies, partners, and industries to design, develop, and deliver solutions that power digital transformation. With a culture that values innovation, growth, and continuous learning, IBM Software places you at the heart of IBM’s product and technology landscape. Here, you’ll have the tools and opportunities to advance your career while creating software that changes the world. As a Site Reliability Engineer, you will work in an agile, collaborative environment to build, deploy, configure, and maintain systems for the IBM client business. In this role, you will lead the problem resolution process for our clients, from analysis and troubleshooting, to deploying the latest software updates & fixes. Your primary responsibilities include: • 24x7 Observability: Be part of a worldwide team that monitors the health of production systems and services around the clock, ensuring continuous reliability and optimal customer experience. • Cross-Functional Troubleshooting: Collaborate with engineering teams to provide initial assessments and possible workarounds for production issues. Troubleshoot and resolve production issues effectively. • Deployment and Configuration: Leverage Continuous Delivery (CI/CD) tools to deploy services and configuration changes at enterprise scale. • Maintenance and Support: Tasks related to applying security patches and upgrades, and collaborating with Product support for issue resolution. Location Flexibility: By applying to this requisition, you acknowledge and agree to be considered for any of the listed locations associated with this position and are willing to work at the location where you are ultimately assigned. System Monitoring and Troubleshooting: knowledge in monitoring/observability, issue response, and troubleshooting for optimal system performance. Automation: knowledge in automation for production environment changes, streamlining processes for etticiency, and reducing toil. Linux: Knowledge of Linux operating systems. Operation and Support Experience: Understanding in handling day-to-day operations, alert management, incident support, migration tasks, and break-fix support. Scripting: knowledge or experience of Python, go or bash. Familiar with cloud providers like IBM Cloud, AWS, Azure or GCP. Kubernetes/OpenShift: knowledge or experience of Kubernetes/OpenShift environments. Automation/Scripting: knowledge or experience of Ansible, Python, Terraform, and CI/CD tools such as Jenkins, IBM Continuous Delivery, ArgoCD. Monitoring/Observability: knowledge or experience cratting alerts and dashboards using tools such as Instana, New Relic, Grafana/Prometheus. DBA: Interest or experience configuring and maintaining SQL, NoSQL, and data streaming technologies (e.g. PostgreSQL, CouchDB, Redis, Katka, Spark, etc.). United States Infrastructure & Technology Hybrid Internship Multiple Cities (0147) International Business Machines Corporation