G2 reports that nearly 70% of IT professionals state that IT environments have increased in complexity compared to just two years ago. More complex IT environments suggest the need for increased monitoring and management to ensure that all components efficiently work together. When you add more applications and software to IT management, it can complicate the management process if solutions don’t integrate or there isn’t structure with how the applications work together.
Incident management is easily one of the most annoying things anyone has to ever deal with. There will always be only a handful of people who would ever want to walk into the building on fire to mitigate. That’s the same with most engineering teams. Only a handful are willing to get in, find the root cause, and mitigate the incident.
Heroku is a cloud provider well known for its simplicity and its support out of the box for multiple programming languages. When thinking about consuming logs from applications hosted in Heroku, Grafana Loki is a great choice. But in the past, shipping logs from Heroku to any Loki instance required ad-hoc scripts to fiddle with Heroku’s logs format and send them. This can be a time-consuming experience.
If you’ve ever had a website or service go down as you were using it, then you’ll understand the irritation of a generic error message and a plea to “Be patient!” (if you’re lucky). It’s almost like they know they’re not telling you the full story. The companies that are on top of their outage game will have a prepared link or redirect to their Status Page (or at least, have one prominently displayed on their pages and social media) for times like these.
With distributed IT Operations becoming the norm, most enterprise teams struggle with communication and collaboration within and across the organization. Without the proper tools, staying on top of incidents can be challenging, quickly resulting in outages taking longer to resolve. The overall effect: increase in downtime-related costs and decrease in performance and availability of services making mean time to resolve (MTTR) worse.