Operations | Monitoring | ITSM | DevOps | Cloud

Find code faster: Introducing our new & improved search experience

Finding code across your repositories in Bitbucket just got a major upgrade. We’ve rolled out code search in open beta for Bitbucket Cloud: a faster, more integrated search experience built to help you find code across your workspace without interrupting your workflow. For many developers, search is one of the fastest ways to explore an unfamiliar codebase, investigate an incident, audit API usage, or scope a refactor.

Cloud Migration Strategies, the 6 Rs, and How to Avoid Getting Stuck Mid-Move

Consider a platform team that spends four months building a thorough migration plan. Its pilot, a stateless order-status API, runs on AWS within three weeks. Six months later, that API is still the only workload in the cloud. In this scenario, the blockers are not exotic technical failures.

Incident Response Metrics Worth Tracking (Beyond MTTR)

Most engineering teams track exactly one incident response metric, and it is usually MTTR. It appears on the quarterly slide, it goes up or down by a few minutes, someone says "we need to bring that down," and nothing about the next incident changes. The problem is not that teams measure the wrong thing out of laziness. The problem is that incident response metrics are genuinely hard to design, and a single average duration is the easiest number to produce from an incident tracker.

A Better Way to Monitor Every Digital Journey with LogicMonitor Synthetics and Internet Performance Monitoring

LogicMonitor Synthetics and Internet Performance Monitoring helps ITOps teams catch digital experience issues earlier with outside-in visibility across apps, networks, APIs, and SaaS.

ilert AI SRE is generally available

When you get paged at 3am, it takes about 30 seconds for the notification to reach you and maybe two minutes until you're in front of a laptop, awake enough to read. What you see then is usually a raw alert. A metric name, a threshold, a link to a dashboard. Then the ritual starts: open the dashboard, check what deployed in the last few hours, grep the logs for the first error, ask in Slack whether anyone touched the database. ‍ Most of that time is search.

You Can Have Your Pi and Kepler It Too

One of the features I have been wanting in Kepler for a long time was the ability to use Pi as my agent when spinning up tasks. Pi is such a minimal harness that it doesn’t prompt for approval for every little thing and it’s system prompt let’s the model just be itself. That minimalism comes at a cost, though. Pi doesn’t have ACP support out of the box, so that means we haven’t been able to officially support it in Kepler, yet.

Digital Experience Monitoring with Grafana Cloud: Session Replay, synthetic checks, and faster investigations

When something breaks in production, the questions that matter most are also the toughest to answer from metrics alone: who was affected, what did they actually see, and is this worth waking someone up for? Answering those questions requires a fuller picture of the issue and its impact on your users. That’s where Digital Experience Monitoring (DEM) in Grafana Cloud comes in.

Installing CFEngine with Ansible

You started with Ansible, and for a long time it was the only thing you needed. With a handful of playbooks and an inventory file, you got the job done. However, as the fleet grew, runs went from taking minutes to hours, and parts of the inventory were unreachable at any given moment. This is not an Ansible flaw. Push-based configuration and continuous state enforcement are simply two different jobs. This blog post is all about getting you started with a hybrid system.