
We run observability for dozens of isolated production environments across AWS, OCI, and Azure with a small DevOps/SRE team, and Sumo Logic is the only reason that scales. Everything lands in one org: Kubernetes clusters on EKS and OKE, cloud audit logs, endpoint and identity telemetry, on-prem Windows infrastructure. Incident triage starts in one search box instead of a "which cluster is that log in?" hunt. The API-first design is the underrated part — we generate Terraform from the live org via the API and manage our entire collector estate as code, and onboarding a new client environment dropped from 5–6 weeks to about one. Scheduled Views keep our SLA reporting cheap, the AI anomaly monitors catch what static thresholds miss, and the Flex pricing model with tiered partitions finally let us match spend to data value. Review collected by and hosted on G2.com.
The query language has a real learning curve for engineers coming from other stacks, and the Terraform provider occasionally lags behind newer platform features, which matters if you manage everything as code like we do. Content management (folders, dashboards) via API is clunkier than collector management. None of these are dealbreakers — but they're the rough edges. Review collected by and hosted on G2.com.