Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Containers, Kubernetes, Docker and related technologies.

DevOps Team Extension: Qovery engineers watching your infrastructure every month

DevOps Team Extension gives growing engineering teams a named Qovery engineer, a continuous infrastructure scan, and a monthly remediation plan, so they can scale reliability, security and compliance without hiring a DevOps team first. Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Replacing NGINX with Envoy Gateway on Hundreds of Clusters: What the Plan Missed

In February we planned to move every Qovery-managed cluster from ingress-nginx to Envoy Gateway by the end of March. It took until October. Here is the rollout model, every class of problem we hit, and what I would do differently. Benjamin is a staff engineer at Qovery focused on infrastructure automation, Kubernetes internals, and building the deployment engine that powers thousands of clusters.

Inside a DevOps Team Extension engagement: an infrastructure review and a Typesense architecture session

How Qovery assessed a growing B2B SaaS company's infrastructure across about 150 controls, what the review with their team changed, and why a Typesense high-availability project turned into a sharding plan. Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Digital Sovereignty Panel: Who Really Controls Your Data? | Civo Navigate London

Your data sits in the UK. But who has the final say over it? As AI raises the stakes on where data lives and who controls it, digital sovereignty has moved from policy debate to boardroom priority. At Civo Navigate London, Lord Paul Drayson, Simon Hansford, Richard Woodfield of DataVita, Dan Chester of VAST Data and Ed Barker of Havishams explore what sovereignty really means for governments, businesses and the infrastructure behind them, and what it will take for the UK to adopt AI without giving up control.

Why AI SRE Agents Need More Than Observability Data with Groundcover | Civo Navigate London

Your AI agent can read every log, metric and trace in seconds. It still can't tell you why things broke. At Civo Navigate London, James Herbert of groundcover explains why AI SRE agents need that context to travel the whole development lifecycle. He shows how agents embedded in the tools where software is built and run can move from isolated tasks to investigating and acting with the full picture.

How we investigate Sentry errors with an AI agent

We built an AI agent on Qovery to investigate Sentry alerts before our team picks them up. Here’s how the workflow runs, what it delivers, and where engineers still need to step in. Rémi is a staff frontend engineer at Qovery. He writes about frontend architecture, developer experience, and building scalable UI systems for platform engineering tools.

Building AI SRE Agents, Part 3: Autonomous in the Cloud

Your agent has earned trust in shadow mode. Now it runs on its own: an alert fires, the agent starts, investigates and proposes a fix before anyone opens a laptop. Here is what it takes to make that safe, scalable and better every week. This is the third article in a series on taking an AI SRE agent from a weekend experiment to production. Part 1 built a local, read-only agent on a throwaway cluster and refined it against a synthetic eval set.

We Stuck Minecraft on a Kubernetes Cluster and Observed it with Open Source

Why did we do this? Not important: jump to 1:31 to see the Pis. TL;DR: We stuck Minecraft on kubernetes running on 4 Raspberry Pis in a 3D-Printed case, and monitored it with open source observability. Huge thanks to Percona DBA Ivan Zaitsev for putting the demo together for Percona University, Montevideo. Coroot automatically collects and visualizes all your telemetry data: logs, metrics, traces, profiles, and a complete map of your services. With the complete context of eBPF, it can diagnose the exact cause of an incident in seconds, and show you the exact commands to fix it.

Context Engineering for AI Agents: What to Feed an Agent and What to Leave Out

Picture a Monday at 03:10 UTC. An agent investigating an out-of-memory alert on payment-service works out the pattern: it runs out of memory every Monday between 03:00 and 04:00, and the spike lines up with the batch reconciliation job. The next Monday the alert fires again. A different agent picks it up and starts from zero, because nothing it can see holds what the first one learned.