Operations | Monitoring | ITSM | DevOps | Cloud

One Cloud, Every Environment: The New Civo Dashboard | Civo Navigate London

Your infrastructure lives everywhere. Your view of it shouldn't. Civo CTO Dinesh Majrekar gives the first look at the new Civo dashboard. It's built around one idea: however many regions, clusters or environments you run, managing them should feel like one cloud. Civo makes complexity a thing of the past.

Managing Claude Code Sessions Through Lynx

At Tigera, we spend a lot of time thinking about agent security: identity, policy, runtime controls, and the record left behind after an agent acts. Coding agents create an interesting problem because, in most organizations, they didn’t arrive through the front door. Few companies ran a platform evaluation and rolled Claude Code out to 500 developers. Developers installed it themselves.

Can India's sovereign cloud keep up with what India is building?

The conversation around sovereign cloud in India is getting louder, which is welcome, but as more providers enter the space, I keep seeing the same architectural pattern, and it's worth being direct about what it misses. The pattern is familiar... take a cloud platform, host it in India, manage it end to end, call it sovereign, and let data residency carry the rest of the argument. That works for web apps and standard enterprise workloads.

The Unveiling: NVIDIA Vera Rubin Comes to the UK | Civo Navigate London

At some point, you have to stop talking about it and start building it. Civo CEO Mark Boost unveils Civo's plan for the UK: 40 edge data centres with a combined gigawatt of capacity, built for low-latency inference and NVIDIA Vera Rubin NVL72, the successor to Blackwell. Mark takes us inside a live digital twin of the rack. There are 72 GPUs wired so tightly together that they behave as a single machine, with so much power that the only answer is liquid in, liquid out. The rack is the computer.

What tools detect and resolve Kubernetes resource contention automatically?

Table of Contents Resource contention happens when workloads compete for more capacity than a node or cluster can provide. For CPU and memory, the symptoms are familiar: CPU throttling, OOM kills, and noisy neighbors slowing latency-sensitive services. These are largely solved problems, addressed by accurate requests and limits, Quality of Service classes, and autoscalers like VPA and HPA that adjust sizing and replicas as demand changes. GPU contention is different, and far more expensive to get wrong.

Cycle and Cherry Servers Webinar: Sovereign Bare Metal with Cloud Simplicity

Cycle teamed up with bare metal service provider Cherry Servers to discuss how rising political tension and controversy involving the United States is pushing many European organisations to take a closer look at where their data is hosted and handled. Many companies have been forced to look into alternative options to ensure their data is not hosted in or accessible by anyone outside of Europe.

What platforms support container auto-scaling and policy-driven resource management?

Table of Contents Autoscaling and policy-driven resource management are two sides of the same coin. Autoscaling adjusts capacity as demand changes, while policies define the boundaries it operates within: who can use how much, which workloads can be changed, and what safeguards must be respected. Without autoscaling, clusters are either overprovisioned or overwhelmed. Without policies, autoscaling can create runaway costs, noisy neighbors, or disruptive changes to critical services.

Heroku to AWS in One Command, With an Agent Doing the Work (Webinar Replay)

Replay and recap of our live session: an AI agent reads a Heroku Rails app and deploys the full stack to AWS through Qovery from one prompt. Chapters, timestamps, the four ways teams leave Heroku, and the steps a human should still own. Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Best Kubernetes monitoring tools compared in 2026: A buyer's guide for SRE, platform engineering, and ITOps teams

Monitoring Kubernetes in production is structurally different from monitoring traditional servers: pods are ephemeral, nodes autoscale, and the layered resource model—containers, pods, nodes, namespaces, deployments, services—creates an observability challenge that tools built for static infrastructure were never designed to handle.

Civo Navigate London 2026 Wrap-Up

This week marks the end of our fourth Civo Navigate London event, and it feels like a good moment to say that this one had a slightly different energy from the ones before it. Over the past four years, we have hosted 10 Civo Navigate events across North America, India, and Europe, and each one has taught us something new about what this community actually wants from a day like this. London 2026 was no exception, and I think this year's lineup pushed that a little further than usual.

The Clearinghouse For AI Agents Has A Blind Spot

Jamin Ball’s recent piece, “Systems of Record Won the SaaS Era — Clearinghouses Will Win the Agents Era,” is the cleanest articulation I’ve seen of where the durable moat goes next. His argument is simple and, I think, correct: the SaaS era rewarded whoever owned the system of record, and the agent era will reward whoever owns the clearinghouse.

[WEBINAR] Introducing the Komodor Agentic Operations Platform

Join Komodor CTO and co-founder Itiel Shwartz for a live look at the newly launched Komodor Agentic Operations Platform. Itiel will guide us through where agentic AI operations are headed in 2026 and beyond, and how Komodor got here: years spent resolving incidents in some of the world’s largest production environments, and what that experience revealed about what makes agents valuable in production.

[DEMO] Komodor Agentic Operations Platform

The Komodor Agentic Operations Platform allows enterprises to confidently implement autonomous operations in mission-critical environments, fully governed from the first run. Get started instantly with pre-built, end-to-end agentic workflows, along with the shared infrastructure, tools, MCP gateways and integrations needed to build, run, and optimize your own.

Self-Improving Agents: A Practical Guide to Continuous Learning

We build agents to take work off engineers’ plates. Then we give those engineers a new manual job: reading failed runs and babysitting prompts. Agents will improve themselves automatically. We’re not there yet, but this is the future I’m betting on. We’ve been working on this ourselves at Komodor over the past year. We know how hard it is to turn a failure into an improvement that holds up beyond a few examples.

17: There's No Life Without AI: Agents, MCP, and the Future of Automation With Viktor Farcic

On this episode of Kubex Talks, technology critic Viktor Farcic returns to talk with Andrew Hillier about the rapidly changing landscape in tech. Viktor has gone from AI skeptic to believer, claiming that there really is no life without AI anymore, from a professional standpoint.

Top 10 Managed Kubernetes Services

Kubernetes has become a standard foundation for modern containerized applications, but operating clusters still requires significant engineering work. In the CNCF’s 2025 annual cloud native survey, published in January 2026, 82% of container users reported running Kubernetes in production. As adoption matures, the question for many teams is no longer whether to use Kubernetes, but how much of the operational burden they want to own.

Open Sourcing Kubex's GPU Process Exporter: Gain Visibility in Your Shared GPUs

Table of Contents GPU sharing with NVIDIA hardware is becoming easier to adopt in Kubernetes but it hasn’t been easier to observe. Time-slicing lets multiple workloads share the same GPU. MPS allows CUDA workloads to execute concurrently. Schedulers like KAI make it easier to manage these shared environments. But sharing a GPU introduces a problem that is easy to underestimate.

How to automate Docker Registry creation with Harness Pipelines and Terraform

Provision a fresh Docker Registry with Terraform, build your container image into it, and deploy to Kubernetes in one One pipeline. One click. It provisions a fresh Docker Registry with Terraform, builds your container image into it, and deploys that image to Kubernetes. Every run creates a uniquely named registry, so you never hit naming conflicts. Creating Docker registries by hand every time you spin up a new service or environment gets tedious fast.

HITL for autonomous agents: Where does the human go?

Human approval is easy when you are sitting in front of the agent. For an agent running by itself in a cluster, almost none of that holds. You’re in a meeting and your agent is running in a cluster. It has a service account, it has been asked to keep a service healthy, and it has just worked out that the right fix is to roll back a database migration. Nobody is watching it. That was rather the point of deploying it. You want to get notified to approve such an important action.

Option to Keep Restored or Cloned Virtual Machines Powered Off

Starting with Harvester v1.9.0, virtual machines created from a snapshot, clone, or backup no longer power on automatically. To keep a virtual machine stopped after creation, select the **Remain halted** option on the UI or set `spec.haltAfterRestore: true` in your `VirtualMachineRestore` CRD or virtual machine manifest.

Monitor TAS and gang scheduling for AI training in Kubernetes

Distributed AI training workloads impose complex scheduling requirements that Kubernetes’s built-in scheduler can’t meet. Kubernetes schedules pods individually and independently, but distributed training introduces two requirements that break this model: Pods must land on hardware with the right inter-GPU bandwidth, and all pods must be scheduled simultaneously. If either requirement goes unmet, training stalls or runs far below the hardware’s potential.

Stop Building Your Own Agent Infrastructure. Meet Agent Tasks

Agent Tasks let you run one-time or scheduled AI agents on your own infrastructure, with network guardrails, centralized MCP access and custom agent environment. Alessandro leads product at Qovery. He drives the changelog, roadmap, and product strategy - turning customer feedback into platform capabilities.

Your existing kit just became more valuable

Hardware costs are rising. But Civo Product Director Russ Smith has a different take: your existing kit just became more valuable. The hyperscalers competing for the same DRAM and compute as you still have to pass that cost on eventually. At high utilisation rates, your resource rental overtakes purchase cost. Typically in under a year.

AI Agents on Kubernetes 101: From Laptop Script to Production Pod

In short, this is a beginner’s guide to deploying an AI agent on Kubernetes. You will containerize an agent, store its API key as a Kubernetes secret, write a deployment with health probes and resource limits, expose it with a service, and lock down its network egress, in that order, with a working manifest at every step. On a local kind cluster the whole walkthrough takes about an hour.

Container hardening isn't a substitute for artifact management

Hardened base images are a great secure foundation. They're minimal, security-vetted, and have few dependencies to worry about. But almost nobody ships a bare base image. Teams build on top of it. This video cover whys that "on top of it" layer is where the risk actually lives: Skip the base image hardening and you're building on a shaky foundation. Skip artifact management and you're leaving everything built on top of that foundation ungoverned. A strong posture uses both.

Making Shared GPUs Even Safer with Kubex and HAMi-core

Table of Contents A few months ago, we introduced Kubex support for the KAI Scheduler to improve GPU sharing for production inference workloads. The basic model is simple: The KAI Scheduler handles placement and GPU sharing. Kubex continuously observes usage and adjusts those allocations as demand changes. KAI provides the scheduling foundation. It lets multiple workloads share a GPU while accounting for the amount of GPU each workload requests. Kubex then closes the loop.

We renovated the Civo Community Slack: Here's what changed and why

The Civo Community Slack has become home to over 35,000 engineers, platform teams, students, and practitioners. It’s one of the things we’re most proud of, a genuine space where the people who use Civo and the people who built Civo are in the same room. Since starting the Civo Community Slack, we’ve shipped an entirely new brand, launched Konstruct, and expanded our AI infrastructure.

Railway vs Render vs Your Own Cloud Account: What Actually Fits a Scaleup Outgrowing Managed PaaS

An honest 2026 comparison of Railway, Render, Fly.io and deploying into your own AWS, GCP, Azure or Scaleway account - with the six measurable signals you have outgrown managed PaaS, a priced cost model, and a four-step framework with if-then verdicts. Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Moving Beyond OOM Kills: Introducing Memory QoS in Kubernetes 1.37

Table of Contents For most of Kubernetes’ history, memory management has been a blunt instrument. Cross your limit, and the kernel kills your container. There has been no equivalent to CPU throttling, no graceful backpressure, just a hard stop. With Kubernetes 1.37, that changes: Memory QoS, built on cgroups v2, graduates to Beta and is enabled by default.

Shipped: Rightsize Kubernetes workloads without leaving your MCP client

Changing a Kubernetes resource request takes two numbers: what the workload requests, and what it uses. The CloudZero MCP server now returns both, by cluster, namespace, or workload. This gives you a number you can defend. Usage comes back as P95 over the date range you query, 30 days by default. When an engineering lead asks whether a service runs on a smaller request, that is the figure that settles it. Over-provisioning and under-provisioning show up on the same query.

Cycle achieves SOC 2 Type 1: Strengthening our commitment to data security and system availability

Cycle recently went through a System and Organization Controls (SOC) 2 Type 1 audit and we've got the report. It’s an important step in our continuous commitment to data security and system availability. But are we just checking boxes for the sake of boxes or is there more behind it?