SRE (site reliability engineering) is a discipline used by software engineering and IT teams to proactively build and maintain more reliable services. SRE is a functional way to apply software development solutions to IT operations problems. From IT monitoring to software delivery to incident response – site reliability engineers are focused on building and monitoring anything in production that improves service resiliency without harming development speed.
To supervise the behavior of distributed applications and track the origin of service failures and downtime, developers often use traditional monitoring technologies and tools. However, this approach can fall short in its ability to measure the overall health of modern cloud-native architectures, which can span multiple hosting environments and encompass hundreds of microservices.
At ServiceNow, we’ve adjusted to the changing times, encouraging our employees to work in the most efficient and safe ways available to them to embrace the new world of work. We’ve embraced the distributed work model. Three of our global employees are proving it’s possible to adapt to new working environments and thrive—no matter where they are. Their jobs have impacted their personal decisions to work remotely full time, in the office full time, or a hybrid of both.