Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

Go from 0 to monitoring in minutes with Netdata

Netdata is zero-configuration monitoring. It’s a principle that we’ve stood behind since the project’s beginning, when it was only our CEO Costa trying to solve a “painful, real-world problem,” and it’s one we stand by today. Our insistence on zero-configuration guides every product decision we make, every grooming process, and every React component our frontend teams design.

Scaling Fleet and Kubernetes to a Million Clusters

We created the Fleet Project to provide centralized GitOps-style management of a large number of Kubernetes clusters. A key design goal of Fleet is to be able to manage 1 million geographically distributed clusters. When we architected Fleet, we wanted to use a standard Kubernetes controller architecture. This meant in order to scale, we needed to prove we could scale Kubernetes much farther than we ever had.

Monitor DNS with Datadog

DNS is a critical component of your infrastructure, enabling your services to reach the endpoints they rely on and connecting your users to your web applications from anywhere in the world. In order to keep your DNS healthy and performant, you need complete visibility into both internal and external DNS resolution. Datadog is excited to announce new DNS monitoring features that help you troubleshoot DNS end-to-end, so you can ensure your applications’ performance and availability.

JFrog Xray Quick Scan Guide

JFrog Xray is a Software Composition Analysis tool (SCA) which is tightly integrated with JFrog Artifactory to ensure security and compliance governance for the organization binaries throughout the SDLC. This video will take you through configuring your JFrog Platform instance to start displaying security and license information about the artifacts in your JFrog Artifactory as fast as possible.

Key Network Monitoring Challenges Every Remote Team Faces

Remote teams are not a new concept. Several organizations have been outsourcing development and support tasks to nearshore and offshore bases for more than a decade. And remote working is gradually increasing given the benefits it gives, like high productivity levels, lower costs, and access to a global talent pool. With the recent COVID-19 outbreak, virtual teams and remote working have truly become mainstream and are being embraced by both employers and employees alike.

Maximize Monitoring in Rancher 2.5 with Prometheus

We dedicate a lot of space in our blog to the topic of monitoring. That’s because when you’re managing Kubernetes clusters, things can change quickly. It’s important that you have tools to monitor the health and resource metrics of your clusters. In Rancher 2.5, we introduced a new version of our monitoring based on the Prometheus Operator, which provides Kubernetes-native deployment and management of Prometheus and related monitoring components.

Knowing your systems and how they can fail: Twilio and AWS talk at Chaos Conf 2020

Get started with Gremlin's Chaos Engineering tools to safely, securely, and simply inject failure into your systems to find weaknesses before they cause customer-facing issues. This year’s Chaos Conf was packed full of incredible talks from some of the industry’s foremost experts on Chaos Engineering.

A Migration of 1000 Workloads Starts with a Single...

You’ve heard that old Chinese proverb that says a journey of a thousand miles begins with a single step. It’s sound advice … except if the journey you’re talking about is the one from the data center to the cloud. With cloud deployment at the center of virtually any digital transformation effort, the journey itself can have a profound impact on the successful outcomes you seek at the destination. So you have to get it right. But what does that actually mean?

Generate process metrics to analyze historical trends in resource consumption

Your application’s health depends on the performance of its underlying infrastructure. Unexpectedly heavy processes can deprive your services of the resources they need to run reliably and efficiently, and prevent other workloads from executing. If one of your applications is triggering a high CPU or RSS memory utilization alert, the issue has likely occurred before.

Infinity Welcomes Careful Versioning

Our distinguished competition took a puzzling position of “Imagine there’s no versions”. At Cloudsmith we think that’s crazy. Software versions are what makes software development possible. They make deployment possible. They make distribution possible. How else can you understand and navigate complex dependency trees or be sure your code will interact with a third party’s? Hint - you can’t! You must care about versions. And updates.