Databricks Reviews (1,358)

View 1 Video Reviews
Reviews

Databricks Reviews (1,358)

View 1 Video Reviews
4.6
1,358 reviews

What do users say?

Generated using AI from real user reviews
Users consistently praise Databricks for its ease of use and powerful scalability, which streamline data engineering and analytics workflows. The platform's ability to integrate various tools and facilitate collaboration across teams enhances productivity and accelerates project timelines. However, some users note that managing costs can be challenging, particularly with compute resources.

Pros & Cons

Generated from real user reviews
View All Pros and Cons
Search reviews
Filter Reviews
Clear Results
G2 reviews are authentic and verified.
Krupa P.
KP
Krupa P.
Software Engineer
Mid-Market (51-1000 emp.)
"Very powerful tool with Spark and big data."
4.5/5
What do you like best about Databricks?

In fact, the most valuable thing about Databricks is that you do not require worrying about looking after the Spark infrastructure. Previously, it took us so much time to configure clusters manually and here, in a few clicks, you can spin up a cluster.

The collaborative notebooks are also very much helpful. My teammates and I are able to collaborate in the same notebook and write Python or Scala or SQL in the same location and share the output in a short time. The connection to AWS and Git is also very fluid, and thus pushing code to production is not demanding a lot of effort at the moment. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

The most significant issue of mine is the cluster start time. There are also cases that I would simply need to make a minor change in the code and the cluster can take about 5-7 minutes to spin up a cold start. It actually disrupts the development. Review collected by and hosted on G2.com.

Response from Jess Darnell of Databricks

We're glad to hear that you find Databricks to be a powerful tool for managing Spark and big data. The ease of spinning up clusters and the collaborative notebooks are indeed some of the key features that our users appreciate.

AP
Aruna P.
Senior IT Manager
Small-Business (50 or fewer emp.)
"This is very powerful for big data and machine learning but watch the cluster costs!"
5/5
What do you like best about Databricks?

The best thing is that we don't have to do any infrastructure to manage now. My team was spending too much time on setting up Apache Spark cluster, managing yarn, and memory crashes on-premise before. With Databricks, we could – within 2-3 clicks – spin up a cluster; collaborative notebooks are very nice! Data engineers and data scientists share the same notebook, so they can collaborate on the same notebook, at the same time. We can have Python, Scala and SQL together in one place without changing any environments. Another super solid feature is delta lake; those provide us with transactions over raw data, this saved us from a lot of data corruption issues since in the past. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

Frankly, its cost is quite high. They charge for DBUs (Databricks Units) and then the cloud provider charge (we're using AWS). Unless you are monitoring, the bill will shoot like a rocket. At times, my developers forget to shutdown the clusters and, if auto-termination is not configured correctly, it's running all night and we get an interest big wake up call in the billing portal the next morning. Review collected by and hosted on G2.com.

Response from Aunalisa Arellano of Databricks

We appreciate your feedback on the benefits of using Databricks for your data management and processing needs. It's great to hear that it has helped to centralize your data and improve processing speed. We understand your concerns about the cost and the potential for unexpected billing spikes. We recommend closely monitoring cluster usage and considering auto-termination to help manage costs. Please feel free to reach out directly to your account team for further questions or feedback.

JF
Joseph F.
Cloud Engineer
Mid-Market (51-1000 emp.)
"Databricks Notebooks Make Collaboration Seamless Across Python, SQL, and Scala"
5/5
What do you like best about Databricks?

Databricks collaborative notebooks are really useful and let me work in whatever language I need to meet my requirements effectively. The ability to mix Python, SQL and even Scala within a dashboard makes collaboration and teamwork much smoothet. I also appreciate how easily it integrates with other tools and cloud platforms, so it fits into my existing workflows without very little friction. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

I like their customer support and the frequent updates are a big reason this has become my favorite for data management, I also appreciate how well it integrates with external tools like Power BI for reporting its really good. Review collected by and hosted on G2.com.

Response from Janelle Glover of Databricks

It's great to hear that Databricks is simplifying cross-team collaboration and improving development cycles for you. We strive to provide a platform that reduces infrastructure and analytics overhead, allowing teams to focus on their core objectives.

Sachin G.
SG
Sachin G.
Machine Learning Engineer
Mid-Market (51-1000 emp.)
"Eliminates the fragmentation tax for ML teams, but Unity Catalog migration takes patience"
5/5
What do you like best about Databricks?

Managing end-to-end machine learning pipelines, specifically training and deploying multi-agent models and recommendation engines.What I appreciate most about Databricks is how it completely eliminates the coordination overhead—the fragmentation tax—between our data engineering and data science teams. Before Databricks, we were losing hours every day moving data between unmanaged data lakes, proprietary data warehouses, and our isolated machine learning compute clusters. Having MLflow natively managed inside the Databricks workspace is a massive advantage for my day-to-day workflow. I no longer have to worry about setting up tracking servers or maintaining infrastructure just to log my training metrics, because Databricks handles the automatic updates and maintenance seamlessly. Every experiment is automatically tracked, and the model registry seamlessly handles version control, making the handoff from experimentation to production deployment incredibly smooth. Additionally, the recent updates to MLflow for evaluating GenAI agents, specifically the ability to use trace-derived baselines to generate runnable evaluation scripts, have saved me countless hours of manual assembly. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

The transition to Unity Catalog has been a significant hurdle for our team. Upgrading our legacy workspace to support Unity Catalog's centralized access control and lineage tracking involved a steep learning curve, especially when dealing with privilege inheritance and ensuring the correct schema privileges were granted across the board. Furthermore, while the platform beautifully abstracts away a lot of DevOps work, it can obscure underlying infrastructure costs. It is far too easy for an engineer to spin up an oversized compute cluster for a simple exploratory data analysis task, leading to sudden and severe spikes in our monthly cloud bill. You have to be extremely disciplined with setting strict auto-termination policies and cluster management rules to keep costs in check. The user interface can also feel a bit tedious at times, requiring you to click through multiple layers in the Catalog Explorer just to view the model details page and trace table-to-model lineage. Review collected by and hosted on G2.com.

Response from Aunalisa Arellano of Databricks

Thank you for sharing your detailed feedback on your experience with Databricks. We are thrilled to hear that you are enjoying the benefits of managing end-to-end machine learning pipelines seamlessly. We understand that the transition to Unity Catalog has presented challenges for your team, and we appreciate your patience as you navigate this process. Your insights on infrastructure costs and user interface are valuable, and we will share this feedback with our team for further improvement.

We are committed to providing a platform that streamlines your workflows and enhances productivity. If you have any specific concerns or need assistance with the Unity Catalog migration or any other aspect of Databricks, please feel free to reach out to us. We are here to support you every step of the way. Thank you for choosing Databricks!

Jagdish S.
JS
Jagdish S.
Associate Data Scientist
Mid-Market (51-1000 emp.)
"Phenomenal Spark Performance, Frustrating UX, and Eye-Watering Bills"
5/5
What do you like best about Databricks?

I run a data science team at a mid-sized company where we handle everything from messy data pipelines to heavy-duty machine learning. Databricks is the core engine of our stack. We use it to ingest raw customer telemetry, clean it up, and run massive PySpark jobs to train our predictive models. We also rely heavily on its MLflow integration to manage our model registry and handle deployments. Essentially, it's the infrastructure playground where all our heavy data lifting happens.The sheer raw performance is unmatched. If you are dealing with massive, bloated datasets that choke local machines or standard cloud instances, Databricks handles them like a beast. The managed Spark environment takes away a massive chunk of the infrastructure headaches involved in setting up clusters from scratch. From a pure data science perspective, having collaborative notebooks where my team can jump in, write Python or SQL concurrently, and instantly visualize data without switching tools is a massive plus. The MLflow integration is also fantastic; being able to track hyperparameters, log artifacts, and register models in the exact same workspace where the data actually lives saves us from fragmented tool sprawl and keeps our MLOps pipelines incredibly tight. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

The user experience can be deeply frustrating, and the platform often feels like a collection of entirely different tools taped together. The UI is clunky, unintuitive, and constantly changing, which means you waste time just trying to navigate the workspace. Debugging a failed Spark job is also an absolute nightmare—you have to dig through endless layers of convoluted driver and executor logs just to find a simple syntax or out-of-memory error. But my absolute biggest issue is the pricing structure. The billing is completely opaque. They charge you Databricks Units (DBUs) on top of your standard cloud provider's compute costs, and if a junior dev accidentally leaves a high-concurrency cluster running over the weekend without auto-termination strictly configured, you will face an eye-watering bill on Monday. Review collected by and hosted on G2.com.

Response from Jess Darnell of Databricks

We're thrilled to hear that Databricks has been instrumental in improving your data science team's workflow and performance. We understand your frustration with the user experience and pricing structure, and we are constantly working to improve these aspects of our platform. Your feedback is valuable to us and will be shared with our team for further consideration.

Jatin P.
JP
Jatin P.
Production Manager
Pharmaceuticals
Enterprise (> 1000 emp.)
"Unified AI and Data Engineering Platform with Smooth Cost Control"
5/5
What do you like best about Databricks?

I appreciate how Databricks brings together data engineering, analytics, and machine learning processes in a single, governed workspace. The data reliability features like automatic versioning, transaction support, and quality controls are great for maintaining consistency and audit readiness without extra manual effort. For AI-related work, I find the experiment tracking, model deployment, and governance capabilities helpful for scaling efforts securely while meeting compliance standards. I also like the cost monitoring and cluster data management tools, which provide better visibility and help control expenses as usage grows across departments. The detailed breakdowns by job, cluster, user, and workload type, along with budget and alerts, are particularly useful. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

There is a learning curve when first adopting Databricks, especially for teams transitioning from traditional setups. The initial setup was a little difficult for these teams. Review collected by and hosted on G2.com.

Response from Jess Darnell of Databricks

We're glad to hear that you appreciate the unified workspace and data reliability features of Databricks, as well as the AI-related capabilities and cost monitoring tools. We understand that the learning curve and initial setup may be challenging for some teams, and we're continuously working to improve the onboarding process to make it easier for new users.

Anita P.
AP
Anita P.
Business Intelligence Analyst
Mid-Market (51-1000 emp.)
"Unified Scalable Data Processing and Machine Learning Platform"
4/5
What do you like best about Databricks?

As a Data Scientist working for a mid-size company, my main use case for Databricks is as the central engine for all of our data processing and predictive modeling pipeline. I use it every day to pull raw dirty data from our cloud storage, explore it with complicated SQL queries and then create and train machine learning models with PySpark and Python. Basically it gives our data engineering and data science teams a common place to play on the same huge data sets at the same time without having to endlessly exchange files or credentials.From a day-to-day workflow perspective, I love the fluidity of the collaborative notebook environment. The ability to work with different languages in the same workplace is a great advantage. I can perform an optimized SQL query to pull in a hefty data set in one cell, then process it in the next using PySpark, and visualize it with Python libraries straight after. This fully removes the need to constantly bounce between different tools or IDEs. Another big victory for my daily work is the out-of-the-box connection with MLflow. It makes it very easy to roll back to a previous version, automatically tracks hyperparameter tuning, compares several model runs, and manages the full lifespan of a model. I really enjoy how Databricks takes away the effort of managing Spark clusters, you can spin up a distributed cluster with a few clicks, and focus on writing algorithms vs playing DevOps. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

And despite all its potential, working with Databricks does come with certain daily difficulties. What is most important for a mid-sized company like us is the aggressive pricing model for compute costs. The monthly payment can get out of control very rapidly, if you’re not compulsively watching your cluster configurations and auto-termination settings especially if a high-memory cluster is unintentionally left operating over the weekend. Another major pain point is the built-in Git integration. Databricks Repos has been helpful however managing complicated merge conflicts or branch management still feels unexpectedly clumsy compared to a regular local IDE like VS Code. Lastly, the learning curve is rather severe for new employees. The user interface might be complicated and debugging distributed computing failures can be a major bottleneck for young data scientists getting up to speed. Review collected by and hosted on G2.com.

Response from Aunalisa Arellano of Databricks

We're glad to hear that Databricks has been able to streamline your data processing and predictive modeling pipeline, and that you find the collaborative notebook environment and multi-language support advantageous for your day-to-day workflow.

Ranjit P.
RP
Ranjit P.
Cloud Engineer
Information Technology and Services
Mid-Market (51-1000 emp.)
"Managed Spark Clusters and Collaborative Notebooks That Just Work"
4.5/5
What do you like best about Databricks?

The best thing about Databricks is the managed Spark clusters. Earlier, setting up Apache Spark manually on AWS or Azure was a big headache. Now, with Databricks, I can spin up a cluster with just a few clicks. The auto-scaling feature works very well, when processing heavy data workloads, it automatically adds nodes and reduces them when done, which saves some cloud costs.

Also, the collaborative notebooks are amazing. My team members and I can work on the same Python or SQL code at the same time, just like Google Docs. The integration with Delta Lake is also a big plus because it gives ACID transactions directly on cloud storage, so data corruption issues are very rare now. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

The biggest issue is the pricing. Databricks DBUs Databricks Units are quite expensive, and if you are not careful with cluster configurations or leave a cluster running by mistake, the cloud bill will jump very high quickly. The cost management tools inside the platform could be much better. Review collected by and hosted on G2.com.

Response from Aunalisa Arellano of Databricks

It's great to hear that Databricks has helped to solve the challenges of data silos and slow ETL pipelines for your team. We are committed to providing a unified analytics platform that enables seamless collaboration and faster data processing for our users.

EC
Eleazar C.
Enterprise (> 1000 emp.)
"Comprehensive Ecosystem, Complex Setup"
5/5
What do you like best about Databricks?

What I like the most about Databricks is the whole ecosystem. It's not easy to have everything you need in a single platform that already has access to the data by its nature. You don't have to handle complex integrations for new projects like data engineering, machine learning, creating dashboards, or developing applications. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

I think Databricks can improve in the complexity. It gets difficult or tricky because there are plenty of things and features, and at some point, it becomes complicated to catch all of them. The user experience can improve, especially for stakeholders that are not 100% technical. It's not easy to set up; you need to set up a lot of things, and when you just start, it's really complicated to get things done. Review collected by and hosted on G2.com.

Response from Jess Darnell of Databricks

We appreciate your feedback on the complexity of Databricks. We are constantly working to improve the platform and make it more user-friendly, especially for those who are not fully technical. Thank you for bringing this to our attention.

Anupama J.
AJ
Anupama J.
Junior Data Analyst
Enterprise (> 1000 emp.)
"Prominent when scaling LLMs and pipelines, but be mindful of the cloud bill!"
4.5/5
What do you like best about Databricks?

As a researcher of AI, it seems like infrastructure is the number one problem, especially setting up clusters, building drivers, and scaling distributed training. Databricks takes care of all that by itself. I can easily and quickly deploy a cluster of nodes for GPUs with PyTorch and DeepSpeed preconfigured in a few clicks. This built-in MLflow is a lifesaver to keep track of experiments. All the hyperparameters or architecture changes with respect to an embedding model are automatically being tracked every time. ESSENTIAL: I no longer have to struggle to get clean and versioned datasets from data engineers for training purposes when working with Delta Lake. Getting around those feature stores is also very easy with the Unity Catalog. Review collected by and hosted on G2.com.

What do you dislike about Databricks?

First, it's really expensive, brother. On an extremely large A100 GPU cluster, if you, or someone on your team, forget to configure the auto-terminate, you are going to have a very bleak day with finance tomorrow. Expenses can add up quickly. Additionally, although they are too lightweight to be an ideal platform for distributed deep learning, the debugging workflow may be tedious. The intersection of the computing nodes makes it difficult to find the exact PyTorch-Out-Of-Memory or CUDA-Out-Of-Memory error occurring in the Spark logs. I also feel like the native MLflow UI in Databricks isn't as advanced and specialized as some of the tools like Weights & Biases. Review collected by and hosted on G2.com.

Response from Jess Darnell of Databricks

Thank you for your feedback - we appreciate your review!