Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on DevOps, CI/CD, Automation and related technologies.

What Can a Service Mesh Do for Your Kubernetes Environment? with Tony Pope-Cruz

Explore the essentials of Kubernetes management with Tony Pope-Cruz from @dynatrace in this detailed walkthrough. Understand how to avoid common pitfalls in Kubernetes deployments, such as mismanagement of resources that can lead to significant outages. Gain insights into how service meshes provide robust solutions for traffic management, service reliability, and observability.

Three roles you need for reliability success

It’s one thing to say that reliability is a priority for your organization, and a whole other thing to make actual, demonstrable improvements in the availability of your applications. Sadly, it’s common for organizations to invest time, money, and effort into improving reliability only to barely nudge the needle on incidents and downtime. But there are hundreds of companies successfully improving their reliability posture—and doing it at enterprise scale.

Manage incidents seamlessly with the Datadog Slack integration

Modern, distributed application architectures pose particular challenges when it comes to coordinating incident management. DevOps, SREs, and security teams—often spread out across separate locations and time zones, and equipped with limited knowledge of each other’s services—must work quickly to collaboratively triage, troubleshoot, and mitigate customer impact.

Console Connect recognised as gold-tier Google Verified Peering Provider

In the cloud-centric world of today, enterprises rely on highly available connectivity for access to public-facing Google Cloud apps, Google Workspace, or Google APIs. Some also need to access latency sensitive Secure Access Service Edge (SASE) solutions, combining security and networking services on one cloud platform.

Balancing AI Workloads and Energy Demands with DCIM Software

AI-driven processes, including machine learning models and data processing, require significant computational resources which can lead to increased energy consumption and heightened operational costs. The complexity of these workloads, which often involve real-time data analysis and continuous model training, exacerbates the need for robust data center management.

Resolve Actions vs. DIY Automation: Which is Really Better?

When it comes to IT automation, the choice between a service orchestration and automation platform like Resolve Actions or building your own automation engine or workflow tool is a pivotal decision. While each option presents its own set of strengths and considerations, Resolve Actions emerges as a powerful solution for organizations looking to streamline their automation efforts. Here’s why.

Five ways Gremlin helps organizations meet DORA requirements

Enacted by the European Union, the Digital Operational Resilience Act (DORA) establishes new standards for digital operational resilience in the financial sector. DORA changes the financial sector's approach to digital security and resilience by imposing stringent Information and Communication Technology (ICT) risk management, incident reporting, third-party risk management, and regular testing.

Practical lessons for AI-enabled companies

We went live with our first set of AI-enabled features a few months ago. Needless to say, we learned a lot along the way, as this was the first time we had experimented with generative AI. Here, I'll share some of what we've learned as we’ve grappled with using LLMs to power new products at incident.io. This will be most applicable to the application layer, AI-enabled but not AI companies.