PagerDuty is an end-to-end incident management and response platform that provides developers, IT operations, and business stakeholders the insights they need to resolve and prevent business-impacting incidents quickly. PagerDuty makes it easy to monitor your infrastructure, set up on-call schedules, establish escalation policies, create automated workflows, and alert the right people at the right time.
Catalytic guides the team efficiently through business processes, automates mundane tasks and provides real-time visibility into operations.
Rootly is a Slack-native incident management app and platform that automates tedious manual work during incidents. 1,000's of users around the world trust Rootly daily to create a simple, consistent, and delightful incident management process. Some things you can do and automate with Rootly: - 🧪 Creating dedicated incident channels, Zoom rooms, and Jira tickets - 🧑🏾🤝🧑🏾 Looping in the right teams/responders and assign roles (e.g. Commander) - ⏰ Setting reminders and tasks (e.g. upda
LibreNMS is an open-source, fully featured network monitoring system that provides real-time visibility into network performance, device health, and security. It supports a wide range of devices using SNMP and other protocols, making it a flexible solution for IT teams managing complex networks. Key Features and Functionality: - Automatic Network Discovery: Automatically detects and adds new devices, reducing manual setup. - Customizable Alerts & Notifications: Alerts via email, Slack, Mi
CloudThinker is an autonomous AI-powered cloud operations platform built for engineering teams that need to move fast without sacrificing reliability, security, or cost efficiency. Rather than another dashboard you have to interpret yourself, CloudThinker's AI agents work continuously across your stack, managing cloud resources, reviewing code, responding to incidents, and optimizing spend with minimal human intervention required. Six core platform modules: AI Code Review Agent — automated PR
PerfAgents Uncloud is an AI based synthetic monitoring tool offered as software-as-a-service. It is designed to monitor the average availability and average response time metrics for websites, web applications, REST API endpoints, and web transaction flows. PerfAgents provides the infrastructure to run scripts from multiple locations around the world to find the availability and response time from these locations. The scripts can be written using any of the following open-source frameworks:
Avoryx replaces the sprawl of separate tools — Jira, Freshdesk, GitBook, PagerDuty, your CRM, DocuSign and more — with one AI-native platform where a ticket, the fix, the customer, the contract and the revenue are the same record, not six disconnected apps. Unlike point tools that only run your team, Avoryx also runs your money: a revenue back-office that invoices, meters, recognizes and reconciles on one audited ledger. The AI turns support tickets into fixes and pull requests, human-gated. Bu
Slik Protect offers comprehensive visibility into the security posture of your cloud infrastructure and databases. With automated, customizable backup schedules, you can ensure that your entire infrastructure is safeguarded and can be restored at any moment. Our agentless scanning technology provides complete visibility into every aspect of your cloud environment, eliminating blind spots. Slik seamlessly integrates with AWS, Azure, GCP, and Alibaba Cloud, covering virtual machines, containers,
TokenTimer is a privacy-focused SaaS platform that helps teams and individuals prevent outages and compliance failures caused by expired assets such as API keys, TLS certificates, secrets, licenses, and contracts. Unlike generic monitoring or password-management tools, TokenTimer focuses exclusively on expiration lifecycle management. Our platform centralizes all expiration tracking in one secure dashboard and sends proactive, threshold-based alerts via Email, Slack, Microsoft Teams, Discord, W
BlazeRunner is a premier provider of comprehensive software testing solutions, dedicated to ensuring the reliability and performance of your applications. With a team of seasoned experts, BlazeRunner offers a suite of services designed to meet the diverse needs of modern software development. Key Features and Functionality: - Performance Testing: Conducts various tests, including load, endurance, scalability, stress, spike, and volume, to assess and enhance system performance. - API Testing &
RigD Inc. offers a collaborative platform tailored for DevOps and IT Operations teams, integrating machine learning to enhance incident management and operational efficiency. By automating routine tasks and facilitating real-time collaboration within environments like Slack, RigD empowers teams to resolve incidents swiftly and reduce manual workloads. The platform seamlessly integrates with tools such as PagerDuty and Jira, ensuring a cohesive workflow across various systems. Key Features and
Collaborative on-call for engineering teams. View your PagerDuty schedules, connect channels to a rotation team, make overrides, get help with coverage, and more, all right from Slack, for the whole team.
Most uptime monitors only check if your server responds. upsonar goes further — it loads your page and discovers every external dependency: CDN-hosted files, third-party scripts, web fonts, and API endpoints. If any of them fail, you get alerted before your users notice. Availability checks run from multiple global regions simultaneously, so regional CDN failures and localized outages don't go undetected. Beyond uptime, upsonar monitors SSL certificates approaching expiration, tracks domain exp
AmAzE Uptime monitors your websites and APIs up to every 60 seconds, no agents, no code changes, nothing to install. The moment something breaks, you get an instant email with the full picture: HTTP status code, error message, response headers, and a body snapshot. Not just "site is down." Exactly what broke and when. SSL certificates are monitored automatically on every HTTPS endpoint and you're alerted before they expire, not after. Scheduled maintenance? Set a window and alerts pause themsel
Anyshift Annie is your AI Site Reliability Engineer, deployed as a Slack-native and MCP-compatible agent. Annie investigates incidents, answers infrastructure questions, and audits changes against a versioned knowledge graph that spans every layer of your stack: cloud (AWS, GCP, Azure), infrastructure-as-code (Terraform), containers (Kubernetes), and monitoring (Datadog, PagerDuty, incident.io). Unlike telemetry-based AIOps tools that infer cause from logs and metrics, Anyshift reads the topolo
ProdView is a privacy-first workforce analytics platform aimed at engineering managers and IT teams. It tracks aggregate productivity signal — productive hours, focus time, app categories, idle ratio, and fleet activity — across macOS, Windows, and Linux via a single lightweight agent (~28MB), deliberately without keystroke logging, screen recording, or message/file access. The core positioning is "a management tool, not a surveillance camera": employees see a bit-for-bit identical view of their
Larm is an uptime monitoring platform that checks websites, APIs, and services from multiple global locations and only alerts you when multiple probes confirm a problem. This multi-location verification eliminates the false positives that plague most monitoring tools. Larm supports HTTP, TCP, DNS, SSL certificate, and heartbeat/cron monitoring. Every HTTP check captures a full request waterfall (DNS, TCP, TLS, TTFB, transfer), giving you the data you need to debug performance issues, not just k
OpsBrief is an AI-powered operations intelligence platform designed to consolidate incidents, releases, deployments, and campaigns from various tools into a single, daily digest. By integrating with platforms such as Slack, Microsoft Teams, GitHub, PagerDuty, Datadog, and more, OpsBrief provides teams with a unified view of critical events, reducing noise and enhancing situational awareness across departments. Key Features and Functionality: - Unified Event Collection: Aggregates events from m