Site Reliability Engineer

Full Time, onsite
Deloitte
Wilayah Persekutuan Kuala Lumpur, Malaysia

Salary undisclosed

Apply on

Linkedn

Original

Simplified

Are you ready to unleash your potential?

At Deloitte, our purpose is to make an impact that matters for our clients, our people, and the communities we serve.

We believe we have a responsibility to be a force for good, and WorldImpact is our portfolio of initiatives focused on making a tangible impact on society’s biggest challenges and creating a better future. We strive to advise clients on how to deliver purpose-led growth and embed more equitable, inclusive as well as sustainable business practices.

Hence, we seek talented individuals driven to excel and innovate, working together to achieve our shared goals.

We are committed to creating positive work experiences that foster a culture of respect and inclusion, where diverse perspectives are celebrated, and everyone is recognised for their contributions.

Ready to unleash your potential with us? Join the winning team now!

Work you’ll do:

As a Site Reliability Engineer (SRE), you will play a key role in maintaining the reliability and performance of critical services. Your expertise will help bridge the gap between development and operations, ensuring robust, scalable, and responsive infrastructure. This role emphasizes strong system architecture and design principles, focusing on key SRE practices such as Service Level Objectives (SLOs), Service Level Indicators (SLIs), and the reduction of operational toil. You will collaborate closely with diverse teams to drive reliability improvements and foster a culture of continuous learning and accountability.

You will,

Design and implement resilient system architectures that support high availability and scalability.
Develop automation tools and scripts to enhance operational efficiency and reduce manual effort.
Define, track, and analyze SLOs and SLIs to ensure reliability and performance meet business needs.
Conduct thorough post-mortem analyses following incidents, driving continuous improvement through root cause identification and solution implementation.
Collaborate with development and operations teams to establish best practices in system reliability and incident management.
Troubleshoot and resolve issues related to database performance, network connectivity, and deployment failures, including diagnosing problems at the underlying platform level (e.g., Kubernetes, virtual machines).
Ensure that issues are resolved within the stipulated Service Level Agreements (SLAs), maintaining high standards of service delivery.
Identify and troubleshoot performance bottlenecks across systems, providing actionable recommendations for enhancements.
Maintain detailed documentation of processes and incident responses to support knowledge sharing and compliance.

Your role as a leader:

At Deloitte, we believe in the importance of empowering our people to be leaders at all levels. We connect our purpose and shared values to identify issues as well as to make an impact that matters to our clients, people and the communities. Additionally, Senior Consultants / Assistant Managers across our Firm are expected to:

Actively seek out developmental opportunities for growth, act as strong brand ambassadors for the firm as well as share their knowledge and experience with others.
Respect the needs of their colleagues and build up cooperative relationships.
Understand the goals of our internal and external stakeholder to set personal priorities as well as align their teams’ work to achieve the objectives.
Constantly challenge themselves, collaborate with others to deliver on tasks and take accountability for the results.
Build productive relationships and communicate effectively in order to positively influence teams and other stakeholders.
Offer insights based on a solid understanding of what makes Deloitte successful.
Project integrity and confidence while motivating others through team collaboration as well as recognising individual strengths, differences, and contributions.
Understand disruptive trends and promote potential opportunities for improvement.

Requirements:

Proficiency in programming languages such as Python, Golang, Java, or similar, focusing on operational efficiency.
Demonstrated experience in system architecture and design, prioritizing reliability, and scalability.
Strong understanding of SRE principles, including SLOs, SLIs, toil reduction, and incident post-mortems.
Experience with cloud environments (e.g., AWS, Azure, Google Cloud) and their operational management.
Strong expertise in Linux system administration.
Proven experience in troubleshooting application support issues with a focus on performance and connectivity.
Familiarity with networking concepts and effective troubleshooting techniques.
Excellent problem-solving abilities and a proactive approach to operational challenges.
Ability to work independently while effectively collaborating within a team environment.

Preferred Skills:

Familiarity with monitoring tools and performance optimization techniques.
Experience in scripting or automation for system administration tasks.
Knowledge of networking concepts and troubleshooting methodologies.
Hands-on knowledge of cloud platforms (e.g., AWS, Azure, Google Cloud) and their services.
Familiarity with DevOps practices and frameworks, including CI/CD, infrastructure as code, and containerization.

Due to volume of applications, we regret that only shortlisted candidates will be notified.

Please note that Deloitte will never reach out to you directly via messaging platforms to offer you employment opportunities or request for money or your personal information. Kindly apply for roles that you are interested in via this official Deloitte website.

Similar Jobs

1d ago

Telecommunication Drive Test Engineer

OG META GROUP SDN BHD