Recommendations to others considering Apache Pig:
You have UDFs which you want to parallelize and utilize for large amounts of data, then you are in luck. Use Pig as a base pipeline where it does the hard work and you just apply your UDF in the step that you want.
Lazy evaluation: unless you do not produce an output file or does not output any message, it does not get evaluated. This has an advantage in the logical plan, it could optimize the program beginning to end and optimizer could produce an efficient plan to execute.
Enjoys everything that Hadoop offers, parallelization, fault-tolerancy with many relational database features.
If you want to do apply some statistics to your dataset. Functional programming paradigm fits quite naturally to pipeline processes, so I expect it to be quite successful. Review collected by and hosted on G2.com.
What problems is Apache Pig solving and how is that benefiting you?
Data Analysis for the raw data we have. Initial data exploration has been useful with pig. Review collected by and hosted on G2.com.