San Francisco, CA, USA
2016
  |  By Fred Hebert
Roughly a year ago, I left Honeycomb’s SRE team to join the newly formed Tenant team, which works on our Private Cloud offering. This team held some significant challenges on its roadmap if it wanted to demonstrate that the offering was possible, would be worth the cost, and could be done without representing a heavy tax on the rest of the organization.
  |  By Austin Parker
A year ago, I wrote “It’s the End of Observability (and I Feel Fine).” The upshot of that post was that AI was about to fundamentally change the way we approach systems design and operation in the future. In the grand tradition, I’d like to revisit my claims from then and see how my predictions panned out.
  |  By Liz Fong-Jones
In this two-part blog series, I give a detailed report-out on how our Honeycomb engineering team 2.5x-ed our throughput using AI without breaking everything or lowering our standards for quality. Part 1 explains how we did it and shows data about how that ramp-up happened. Part 2 shares what we learned.
  |  By Liz Fong-Jones
In this two-part blog series, I give a detailed report-out on how our Honeycomb engineering team 2.5x-ed our throughput using AI without breaking everything or lowering our standards for quality. Part 1 explains how we did it and shows data about how that ramp-up happened. In this blog, I share what we learned. The “platform engineering” frame and the “autonomy, ownership, feedback loops” frame are the same frame, spoken in two different vocabularies.
  |  By Josh Parsons
We recently wrapped up a large-scale, multi-month Kafka migration project. We used to run self-hosted Confluent Platform and ZooKeeper as clusters of AWS EC2 instances, and now all of our Kafka clusters run open-source Apache Kafka 4.1.1 running in KRaft mode and deployed to AWS EKS.
  |  By Charity Majors
The world is especially hard right now. The future of the software engineering profession looks more uncertain than ever. Execs are under heavy pressure to turn AI into magic results, and teams are fighting product competition and AI-induced burnout on one side, melting mental models and hellish oncall on the other side. Observability was supposed to be a solved problem by now.
  |  By Juliana Gomez
One of the hardest challenges facing platform teams is wrangling the rising volume of PRs looking to add drift to the systems we've invested in. It's impossible to catch them all, so it's more important than ever to invest in building stronger guardrails so our product teams can keep building quickly and catch issues before they merge to main. Linters are a great tool to reach for first.
  |  By Moses Mendoza
This post was co-written with Staff Software Engineer Martin Holman. Honeycomb Canvas is a collaborative investigation environment. When something goes wrong in production, multiple engineers might join the same Canvas to debug it together. Each person has their own AI agent, so they can pursue their own conversation thread and line of inquiry. This creates an opportunity for coordination.
  |  By Moses Mendoza
"Hello world, this is your agent speaking!" The agent loop! The LLM is calling tools, the answers are sensible, and the sky's the limit. Now, as you look forward to production, you look for a composable toolset, something that can grow with your use case and system needs. That's what we created with Honeycomb Canvas: a collaborative investigation space where AI agents help you understand, fix, and learn about your system.
  |  By Reid Savage
As of today, I’ve drafted this post upwards of 10 times – it’s old enough that the version I first started working on was called “Reflections on 1 Year of SRE Management” (I’m currently at 2.5 years). But everything I learned during that first year became critical for the next.
  |  By Honeycomb
At Slack, between 100 to 200 users per day use Honeycomb for client observability, tracing, instrumentation, analysis of performance, frontend issues, investigating incidents, or just looking into production issues.
  |  By Honeycomb
Watch Nathen Harvey's full talk at O11yCon 2026, Honeycomb's observability conference, and enjoy Christine Yen's intro as well.
  |  By Honeycomb
In this demo, Liz and Kale talk through a slow query that Liz couldn't get out of her head. During a conference, she set out to solve it... and ended up finding two more bugs to fix with, Honeycomb MCP, and Honeycomb Canvas.
  |  By Honeycomb
In her talk at O11yCon 2026, Nishi Bhonsle of Salesforce talked about,, and provided some great examples of how Honeycomb has helped Salesforce issues in seconds. Here's a 4-minute highlight reel.
  |  By Honeycomb
Honeycomb and Embrace are extending the rigorous, data-driven practice that Honeycomb pioneered for foundational to mobile and web, giving, site reliability, and platform teams a complete, correlated picture of system health. The strategic partnership makes understanding performance and reliability for every user and every screen part of the observability practice, bringing new depth and standardization to how teams measure end user impact.
  |  By Honeycomb
Watch a full replay of all sessions on Day 3 of Honeycomb's Innovation Week.
  |  By Honeycomb
Honeycomb has shipped a production integration with Amazon Bedrock AgentCore, surfacing agent telemetry directly in Agent Timeline, Honeycomb's trace view for behavior. It's available now and built on.
  |  By Honeycomb
Watch this video to see the re-imagined Canvas in action, where auto-investigation has already ranked your hypotheses before you open the tab, multiplayer agents build on each other's work in real time, and a custom skill encoding your team's own runbook can reprioritize the entire incident before you've had your morning coffee.
  |  By Honeycomb
Watch this video to see Agent Timeline in action: one conversation ID, one view, every agent invocation, LLM call, tool call, and downstream trace, so you stop stitching tabs together and start finding the failure in seconds.
  |  By Honeycomb
Watch a full replay of all sessions and demos on Day 2 of Honeycomb's Innovation Week.
  |  By Honeycomb
Honeycomb is an event-based observability tool, but you can-and should-use metrics alongside your events. Fortunately, Honeycomb can analyze both types of data at the same time. When maturing from metrics-based application monitoring to an observability-based development practice, there are considerations that can make the transformation easier for you and your team.
  |  By Honeycomb
Evaluating observability tools can be a daunting task when you're unfamiliar with key considerations and possibilities. This guide steps through various capabilities for observability tooling and why they matter.
  |  By Honeycomb
This document discusses the history, concept, goals, and approaches to achieving observability in today's software industry, with an eye to the future benefits and potential evolution of the software development practice as a whole.

Honeycomb is a tool for introspecting and interrogating your production systems. We can gather data from any source—from your clients (mobile, IoT, browsers), vendored software, or your own code. Single-node debugging tools miss crucial details in a world where infrastructure is dynamic and ephemeral. Honeycomb is a new type of tool, designed and evolved to meet the real needs of platforms, microservices, serverless apps, and complex systems.

Honeycomb provides full stack observability—designed for high cardinality data and collaborative problem solving, enabling engineers to deeply understand and debug production software together. Founded on the experience of debugging problems at the scale of millions of apps serving tens of millions of users, we empower every engineer to instrument and query the behavior of their system.