All Paths Lead To A Federation Data Lake

Is your organization constrained by 2nd platform data warehouse technologies with limited or no budget to move forward towards 3rd platform agile technologies such as a Data Lake? As an EMC customer you have the advantage of leveraging existing EMC investments to develop a Federation Data Lake at minimal cost. Additionally, the Federation Data Lake will generate healthy returns, as it is packaged up with the expertise needed to immediately execute on data lake uses cases such as data warehouse ETL offloading and archiving.

Data Lake

With the release of William Schmarzo’s Five Tactics to Modernize Your Existing Data Warehouse, I wanted to explore whether the Dean of Big Data views data warehouse modernization tactics or paths ultimately leading to a Federation Data Lake.

1.  What is a Data Lake and who should care?

Continue reading

A Novel Idea: Practical Advice From A Big Data Practitioner

Big Data: Understanding How Data Powers Big Business is yet another Big Data book to hit the market. What makes this book unique? There is practical advice and hands on exercises so that you end up with a Big Data action plan unique to your business after completion of the book. I spoke to the author, EMC’s own Big Data’s preeminent expert William Schmarzo, to explain the goals of his book and why organizations grappling with Big Data should pick it up.


1.  What makes you a Big Data expert in providing practical advice for developing Big Data strategies?

Continue reading

Dear BI Users: Your Hadoop SQL Wish Has Finally Come True

To accelerate the value of Big Data, many products have been developed to make data managed in Hadoop much easier to access and analyze through SQL.  First there was Hive, which provides a SQL query abstraction layer by converting SQL queries into MapReduce jobs.  More recently, Cloudera announced Impala which bypasses MapReduce to enable interactive queries on data stored in Hadoop using the same variant of SQL that Hive uses.  And today, EMC Greenplum announced Pivotal HD, the only high performing, true SQL query engine on top of Hadoop.  Don’t be confused by these approaches, as there is a common thread – to leverage Hadoop as a Big Data platform for running SQL queries.  The major difference with Pivotal HD is that now there is a single, scalable, flexible, and cost-effective data platform for all of your analytic needs.



I spoke with Greenplum Chief Scientist Milind Bhandarkar to explain this breakthrough SQL interface to Hadoop.

1. How does Pivotal HD provide a true, high performing SQL interface to Hadoop?

Continue reading

Want To Become A Data Scientist? EMC Can Train You in 5 Days.

Everyone agrees that there is a shortage of Data Scientists. If not addressed soon, Big Data breakthroughs in areas such as healthcare, renewable energy, public sector, etc will decelerate.  I am proud to say that EMC is doing its part to solve the problem by fostering Data Science development with training and certification, hands on expertiseweb events, internships, and more.  For example, EMC Education Services offers a 5-day Data Science and Big Data Analytics  training and certification,  designed to enable immediate and effective participation in big data and other analytics projects.

As a Big Data citizen, I want to motivate those thinking about moving into the world of Data Science, to take action and get trained. I met with Barry Heller, a developer for EMC’s Data Science curriculum, who leverages his extensive education and past experience as an EMC Data Scientist for curriculum development.  If Barry’s story resonates and you relate in some way, I hope it inspires you to start a career in Data Science.

1) How many people have completed the EMC Data Science and Big Data Analytics training since its creation early this year?

Continue reading