I've been using Databricks as part of our data engineering workflow to build and maintain ETL pipelines, analyze large datasets, and support reporting requirements. One of the things I like most is that it brings data engineering, analytics, and notebooks into a single workspace. Instead of switching between multiple tools, I can write PySpark code, validate transformations, collaborate with teammates, and schedule jobs from the same platform. This has made day-to-day development more organized, especially when working on multiple data pipelines.
Another feature I rely on frequently is the notebook environment. It's convenient for developing and testing transformations before moving them into production. During development, I often use notebooks to inspect sample data, troubleshoot failed transformations, and validate business logic with SQL and PySpark. The ability to mix code, markdown documentation, and query results in one place also makes it easier for team members to understand the implementation during code reviews or knowledge transfer sessions.
I also appreciate the platform's scalability. Some of our data processing jobs involve millions of records, and Databricks handles distributed processing efficiently without requiring us to manage the underlying infrastructure directly. Features like cluster management, job scheduling, and integration with cloud storage reduce operational overhead. That said, cluster startup times can occasionally delay quick debugging sessions, and managing compute resources carefully is important to avoid unnecessary costs. Overall, Databricks has helped simplify large scale data processing while giving enough flexibility for both development and production workloads. Review collected by and hosted on G2.com.
While Databricks has been reliable for our data engineering workloads, there are a few areas where I think it could be improved. One challenge I've experienced is cluster startup time. When I only need to test a small code change or validate a transformation, waiting for a cluster to start can interrupt the development flow. It's not a major issue for scheduled production jobs, but during active development and debugging, those extra minutes add up.
Another limitation is cost management. Since compute resources are tied to cluster usage, it's important to monitor cluster configurations and ensure they are shut down when not needed. We've had situations where development clusters remained active longer than expected, resulting in higher cloud costs. The platform provides tools to manage this, but it still requires teams to establish good governance and usage policies. I also found that some configuration settings for jobs, permissions, and clusters have a learning curve, especially for new team members who are unfamiliar with the Databricks environment.
From a day to day perspective, debugging distributed Spark jobs can sometimes be challenging. While the logs provide useful information, identifying the exact cause of failures often requires navigating through multiple execution logs and Spark UI details. For straightforward issues this isn't a problem, but troubleshooting more complex pipeline failures can take time. Despite these limitations, none of them outweigh the benefits of the platform, and most challenges can be managed with proper cluster configuration, monitoring, and team practices. Review collected by and hosted on G2.com.
We're glad to hear that Databricks has streamlined your data engineering workflow and provided a single workspace for analytics and notebooks. We understand the importance of having a convenient environment for developing and testing transformations, and we appreciate your feedback on the cluster startup times and cost management. We're continuously working to improve the platform and provide better tools for managing compute resources and reducing operational overhead. Thank you for sharing your experience with Databricks!