Flashduty is an incident response platform for engineering, DevOps, and SRE teams. It consolidates alerts from monitoring tools, cloud platforms, and custom sources into a single stream, reduces duplicate and low-signal notifications through rule-based and intelligent grouping, and routes what remains to whoever is on call.
Alert aggregation and noise reduction. Flashduty ingests alerts from 50+ sources including Prometheus, Grafana, Zabbix, Nagios, Sentry, Splunk, AWS CloudWatch, Azure Monitor, Google Cloud Monitoring, Alibaba Cloud, Tencent Cloud, and Huawei Cloud. Incoming events pass through a pipeline that normalizes fields, enriches labels, filters known-noise patterns, and groups related alerts into one incident thread. Storm detection and flapping detection suppress repeated firing from an unstable source.
On-call management. Schedules support rotations, multi-level escalation policies, primary and backup roles, severity- and time-window-based routing, and follow-the-sun handoffs across time zones. Escalation continues automatically until someone acknowledges.
Notification delivery. Nine channels are supported: mobile push, phone call, SMS, email, Slack, Microsoft Teams, Feishu/Lark, DingTalk, and WeCom. Phone and SMS act as fallbacks when chat channels go unacknowledged, and each user sets per-severity preferences.
Incident response. A war room can be opened in Slack, Feishu, DingTalk, or WeCom in one step. Acknowledging, escalating, and resolving work from inside the chat, and those actions sync back to the incident timeline. After resolution, Flashduty drafts a postmortem from the alert payload, timeline, and war-room discussion for the team to edit collaboratively, with action items tracked to tickets.
Status pages. Public and internal status pages report component-level availability tied to incident history, with custom domains, branding, email and RSS subscriptions, and chat notifications for internal pages.
AI SRE (beta). An agent correlates telemetry, investigates an incident, proposes likely root causes, and recommends next steps. It supports reusable skill flows, knowledge packs, MCP tool access, and either a hosted sandbox or a self-hosted runner inside the customer network.
Integrations and automation. Bidirectional sync with Jira, ServiceNow, and ServiceDesk Plus; inbound and outbound webhooks; custom actions; an open API with 288 endpoints; plus a Go SDK, CLI, Terraform provider, and MCP server.
Deployment. Available as a hosted service with a free tier for up to five users, Standard at $14 per user per month, and Professional at $29 per user per month. A private deployment option covers on-premises data retention, network isolation, and LDAP integration. Interface and documentation are available in English and Chinese.