Monthly Archive

Webinar Recap: How Observability Impacts SRE, Development, and Security Teams

Jan 31, 2023 By Mezmo In Mezmo

In today’s fast paced and constantly evolving digital landscape, observability has become a critical component of effective software development. Companies are relying more on and using machine and telemetry data to fix customer problems, refine software and applications, and enhance security. However, while more data has empowered teams with more insights, the value derived from that data isn’t keeping pace with this growth. So how can these teams derive more value from telemetry data?

Read Post

Mezmo

Read more about Webinar Recap: How Observability Impacts SRE, Development, and Security Teams

Analytics in Squadcast | Visualize Team and Organization Level Analytics | MTTA MTTR | Squadcast

Jan 31, 2023 By Squadcast In Squadcast

Analyzing incident data plays a key role to do better SRE. Squadcast's Analytics Dashboard helps you analyze the performance of your Organization/ Team, for a given time period. It also gives you more insight into past outages that affected your systems.

View Video

Squadcast

Read more about Analytics in Squadcast | Visualize Team and Organization Level Analytics | MTTA MTTR | Squadcast

What are Network Operation Centers (NOC) and how do NOC teams work?

Jan 30, 2023 By Vishal Padghan In Squadcast

Modern-day markets are highly competitive and in order to foster stronger customer relations, we see businesses striving hard to be always available and operational. Hence, businesses invest heavily to ensure higher uptime and to have dedicated teams that constantly monitor the performance of an organization's IT resources. In this blog, we will explore what NOC teams are and why they are important.

Read Post

Squadcast

Read more about What are Network Operation Centers (NOC) and how do NOC teams work?

SRE Dashboards

Jan 26, 2023 By Cortex In Cortex

One of the most important features of any software tool or web application is its reliability. Businesses that offer slow or unreliable software services always risk losing customers to better, more competent service providers. This makes it important for businesses to constantly monitor and enhance the performance and reliability of their digital systems.

Read Post

Cortex

Read more about SRE Dashboards

Runbook Automation as a Baseline for Controllability and Observability

Jan 23, 2023 By Amalya Shnaps In MoovingON

Some of the highest priorities for engineers - from NOC Engineers, DevOps & Site Reliability Engineers - are the automation and optimization of their production environments. Many companies today face tough challenges with their Network Operations Centers (NOCs) or production environments. These challenges fall into the hands of engineering teams.

Read Post

MoovingON

Read more about Runbook Automation as a Baseline for Controllability and Observability

What are Webhooks and why should developers use them?

Jan 20, 2023 By Vardhan NS In Squadcast

Webhooks and APIs are a developer-friendly approach to building modern-day web applications. In this blog, we explain what a webhook is, do a detailed webhooks vs. API comparison, and explain why we recommend developers use them with Squadcast.

Read Post

Squadcast

Read more about What are Webhooks and why should developers use them?

Reliability and SRE in the 2022 State of DevOps Report

Jan 18, 2023 By Dave Stanke In Google Operations

Learn more about the connection between SRE, DevOps and reliability.

Read Post

Google Operations

Read more about Reliability and SRE in the 2022 State of DevOps Report

SRE Trends from AWS re:Invent 2022

Jan 18, 2023 By Squared Up In Squared Up

In November/December 2022 I attended AWS re:Invent in Las Vegas. It was certainly an experience for this small town kid from New Zealand, and one that I took a lot away from. While I was at the conference, I took the time to walk around and take notes. In this article I will share the trends that I observed which I think will have an impact on SRE work in 2023 and beyond, including: ...and others.

Read Post

Squared Up

Read more about SRE Trends from AWS re:Invent 2022

Understanding Site Reliability Engineering (SRE)

Jan 16, 2023 By Makenzie Buenning In NinjaOne

Success in this modern age of digital services and operations is found when businesses are able to prioritize effective digital processes. Because of this, IT teams are constantly looking for ways to improve their IT operations by making them efficient, reliable, and scalable. One way this is accomplished is through site reliability engineering (SRE). LinkedIn listed SRE as the 21st fastest growing job in the U.S. in January 2022. What is SRE, and why is it in such high demand?

Read Post

NinjaOne

Read more about Understanding Site Reliability Engineering (SRE)

A practical guide for implementing SLO

Jan 12, 2023 By Prathamesh Sonpatki, In Last9

How to set Service Level Objectives with 3 steps guide.

Read Post

Last9

Read more about A practical guide for implementing SLO

Why SREs need better visibility, not more tools

Jan 11, 2023 By LogicMonitor In LogicMonitor

As a site reliability engineer (SRE), you juggle a lot of moving targets. You keep tabs on your operational environment’s health and maximize service levels, all while trying to scale your business and exceed client expectations. To hold it all together, you’ve likely implemented a hybrid cloud strategy to keep a watchful eye over everything: your on-premises infrastructure, containers, and numerous cloud deployments.

Read Post

LogicMonitor

Read more about Why SREs need better visibility, not more tools

Introducing Levitate: 'uplifting' your metrics woes because self-management sucks like gravity

Jan 11, 2023 By Nishant Modak In Last9

Managing your own time series database is painful. We’ve moved from servers to services, and yet, monitoring metrics data is primitive. Our managed time series database powers mission-critical workloads for monitoring, at a fraction of the cost.

Read Post

Last9

Read more about Introducing Levitate: 'uplifting' your metrics woes because self-management sucks like gravity

SRE Report 2023: Are we Aligned? Yes. No. Maybe.

Jan 10, 2023 By Denton Chikura In Catchpoint

Each year of the SRE Report, there’s a trend or anti-pattern that leaps out and makes us pause and reflect. Last year, for example, we found a huge drop in global toil levels. With the whole world working from home for a full year, it made sense that global toil levels would drop, right? But this year, despite the great reopening underway, toil levels dropped even further - it's a paradox, one which no doubt will require its own scrutiny.

Read Post

Catchpoint

Read more about SRE Report 2023: Are we Aligned? Yes. No. Maybe.

Lessons from the CircleCI Security Incident

Jan 9, 2023 By Quentin Rousseau In Rootly

In some respects, security and reliability are competing priorities. Security controls may reduce reliability, and responding to security incidents may require mission-critical systems to be paused or shut down until they're secure. The recent security incident involving CircleCI, however, shows that it's not always necessary to choose between prioritizing security or reliability.

Read Post