Operations | Monitoring | ITSM | DevOps | Cloud

The latest News and Information on Cloud monitoring, security and related technologies.

Shipped: Monthly cost comparison in Explorer gets a glow up

Months have different numbers of days, and a monthly cost chart built on raw totals mixes that calendar difference into the trend. A 28-day February next to a 31-day March shows a 10.7% increase even when daily spend never moved. The same math works in reverse: real growth in a short month can look flat, hiding an increase worth investigating. That costs you time in two places. The first is triage.

AI budgeting: how to plan and forecast AI spend

AI budgeting is the process of planning, allocating, and forecasting an organization's AI spend: model and API costs, AI infrastructure, tooling, and the people running it all. It differs from traditional budgeting because AI spend is usage-based, scales with product success rather than headcount, and often spans multiple providers.

Building an End-to-End Drone Ecosystem: The Technologies That Need to Work Together

Commercial drone technology is rarely a single application running alongside an aircraft. A complete solution may include flight software, onboard sensors, telemetry, cloud infrastructure, web and mobile interfaces, data processing pipelines, analytics tools, and integrations with existing business systems.

How to right-size your existing Claude skills

You shipped a skill. It worked. You closed the tab. That’s the whole problem. Model choice is a decision you make once, at the moment you’re least equipped to make it: before the skill is even authored. Then you never revisit it, because the skill stopped being interesting the day you got it working. So go back and check. Here’s how.

JFrog Artifactory Now Integrates Natively with Artifact Registry in Google Cloud

Teams running containerized workloads on Google Cloud have long relied on JFrog as their single source of truth for container images. The missing piece has been getting Google Cloud’s own runtime services — like Cloud Run and Google Kubernetes Engine (GKE) — to pull directly from JFrog for every container image pull. I’m happy to say that the gap is now closed. Artifact Registry in Google Cloud has introduced a new repository mode called Connector that addresses this requirement.

Peak Cloud: Decentralising for resilience

For more than a decade, the prevailing wisdom in enterprise IT was simple: move everything to the public cloud. Hyperscale platforms promised unlimited scalability, lower costs, agility and freedom from the burdens of managing infrastructure. Cloud-first has been rapidly gaining momentum as the de facto path to a modern digital footprint. Until now.

Private cloud disaster recovery: How to design for business continuity without public cloud dependency

Disaster recovery (DR) is one area where organizations often assume public cloud has the answer already. Multi-region deployments, managed backup services, automated failover - the hyperscaler catalog is full of DR-flavored offerings, and the marketing suggests that resilience is a solved problem once you're on cloud infrastructure. For many workloads, this is roughly true.

The Architecture Question That Never Dies: From BPMN and M&A to MCP

Twenty years ago at RMIT, I became preoccupied with a question that sounded technical but was really about corporate value: could you predict how difficult a company would be to acquire by looking at the shape of its APIs? It was 2006. I was completing Honours in a Bachelor of Applied Science in Software Engineering, and the brief for my research project was unusually open: find an impactful software research hypothesis that hasn’t been done before.

Cloud Incident Management: Process, Tools, and Practices

How do you resolve an outage your organization has no authority to fix? A managed database drops into read-only mode and stops accepting writes. There's no host to reach, no configuration file to edit, and no restart command available to your engineers. Cloud incident management begins at that boundary, where the response depends on a support channel and a provider status page. Plenty of what you already know still applies here.