Operations | Monitoring | ITSM | DevOps | Cloud

Introducing Infrastructure Knowledge: Teach Netdata AI What Your Metrics Can't Show

Netdata AI sees everything your infrastructure does: every metric, every anomaly, every alert. It does not see what your infrastructure is: which services matter, which host is supposed to run hot, who owns what, what your team considers normal. Without that context, “CPU at 91%” is just a finding. With it, it might be a machine doing exactly its job.

Monitor test health at a glance in Bitbucket Tests

When a team relies on automated tests in CI/CD, knowing that tests ran is only the beginning. Understanding whether the suite is healthy, which tests need attention, and how a specific test has behaved over time — that’s what drives action. Bitbucket Tests is evolving to make those answers easier to find and give you tools to improve your test health.

Assisted, Augmented or Agentic? Choose Your Splunk Starting Point

Episode two of Beyond the Thread explores how organizations can leverage a solid data foundation for AI-driven actions. Hosted by Courtney Wright and featuring experts Greg Ainsley-Malik and Sonal Pardeshi, the discussion delves into the Cisco Data Fabric, powered by the Splunk platform, and its role in transforming machine data into actionable insights. The episode highlights the journey towards agentic operations, addressing the challenges faced in moving from AI-ready data to effective implementations, and examines different adoption strategies that organizations may pursue.

Agentic Operations Start with Context: Build the Right Data Foundation

Episode 1, "Beyond the Thread: Deconstructing the Cisco Data Fabric Powered by the Splunk Platform," explores the intersection of data strategy and operational efficiency. Hosted by Splunk's Courtney Wright, the session features insights from experts Keith McClellan and Michael Sondag on the complexities organizations face in data management and operational models.

How to Build an HR PTO AI Agent with Resolve Agent Lab

See how to build an HR PTO agent with Resolve Agent Lab. In this Resolve Reels demo, we create a purpose-built AI agent by adding automation skills, instructions, conversation starters, and guardrails. The agent can answer PTO questions, check balances, account for calendar conflicts, and submit requests through systems like Workday or ADP. See how Resolve helps teams build AI agents that take action across enterprise systems.

AI Agents on Kubernetes 101: From Laptop Script to Production Pod

In short, this is a beginner’s guide to deploying an AI agent on Kubernetes. You will containerize an agent, store its API key as a Kubernetes secret, write a deployment with health probes and resource limits, expose it with a service, and lock down its network egress, in that order, with a working manifest at every step. On a local kind cluster the whole walkthrough takes about an hour.

Connect Codex to CircleCI: Fix Failing CI Without Leaving Your Terminal

Connect Codex to CircleCI and give your coding agent direct access to the CI feedback it needs to keep working. In this tutorial, we’ll walk through setting up the CircleCI CLI and CircleCI plugin for Codex, then show how Codex can check pipeline results, validate your CircleCI config, diagnose failed builds, trigger new pipelines, and keep iterating on a fix until CI is green. Instead of bouncing between your terminal and CircleCI to copy logs and errors back to your agent, you can bring the full CI feedback loop directly into your Codex session.

Your existing kit just became more valuable

Hardware costs are rising. But Civo Product Director Russ Smith has a different take: your existing kit just became more valuable. The hyperscalers competing for the same DRAM and compute as you still have to pass that cost on eventually. At high utilisation rates, your resource rental overtakes purchase cost. Typically in under a year.

Continous ORT Testing with Harness

Most Operational Readiness Testing (ORT) programs follow the same ritual. A checklist gets filled out. Someone runs a load test in a war room the week before launch. A failover drill gets scheduled, and everyone hopes it goes cleanly. Then the release is shipped, and testing is done. But with Harness, you can make this process continuous, and your service resilience is protected by the same ORT checklist for every small change in your SDLC.