
What I like most about OLake is its focus on reliable data ingestion into modern lakehouse destinations, rather than trying to be yet another generic ETL wrapper. I’ve mainly used it for Postgres ingestion, CDC/incremental sync behavior, and testing Parquet destinations, and those workflows made it clear to me that OLake is addressing a real infrastructure problem.
For my use cases, the Postgres integration, CDC support, and the project’s direction around Parquet/Iceberg destinations are the most valuable parts. They help teams move data into open formats on object storage without getting locked into a single warehouse. I also appreciate how approachable the project feels for open-source contributors: the separation between sources, destinations, state, and recovery logic makes the codebase easier to understand, and the maintainers are active and helpful when it comes to guiding implementation details. Review collected by and hosted on G2.com.
The main downside is that OLake still feels like a fast-moving open-source project, so in a few areas you end up needing to read the docs, GitHub issues, or even the codebase to fully understand the expected behavior. For example, more advanced topics like CDC state handling, retries, destination recovery, and object-storage behavior can be difficult to reason about at first.
I’d like to see more end-to-end examples, architecture diagrams, and troubleshooting guides that cover common production scenarios. The overall direction feels strong, but clearer onboarding and practical reference material would make it easier for new users and contributors to adopt OLake more quickly. Review collected by and hosted on G2.com.