---
title: "EMC Big Data Play Continues: Greenplum Acquisition"
date: 2010-07-07T16:51:11Z
modified: 2010-07-07T16:51:11Z
permalink: "https://redmonk.com/jgovernor/emc-big-data-play-continues-greenplum-acquisition/"
type: post
status: publish
excerpt: ""
wpid: 2825
categories:
  - Uncategorized
tags:
  - Uncategorized
  - EMC
  - hadoop
  - mapreduce
timestamp: 2010-07-07T16:51:11Z
---

I wrote recently about [VMware’s emerging Data Management play](http://www.redmonk.com/jgovernor/2010/04/21/vmwares-springsource-redis-and-rabbit-acquisitions-a-database-play-is-emerging/) after the announcement the firm was [hiring Redis lead developer Salvatore Sanfillipo](http://blogs.vmware.com/console/2010/03/vmware-hires-key-developer-for-redis.html).

> While \[CEO Paul\] Maritz may say VMware isn’t getting into the database business, he means not the _relational_ database market. The fact is application development has been dominated by relational- Oracle on distributed, IBM on the mainframe – models. Cloud apps are changing that. As alternative data stores become natural targets for new application workloads VMware does indeed plan to become a database player, or NoSQL player, or data store, or whatever you want to call it.
> 
> We have been forcing round holes into square pegs with object/relational mapping for years, but the approach is breaking down. Tools and datastores are becoming heterodox. something RedMonk has heralded for years.

Now comes another interesting piece of the puzzle. EMC is acquiring Greenplum – and building a new division around the business, dubbed Data Computing Product Division. While Redis is a “NoSQL” data store, Greenplum represents a massively parallel processing architecture designed to take advantage of the new multicore architectures with pots of RAM: its designed to process data into chunks for parallel processing across these cores. While Greenplum has a somewhat traditional “datawarehouse” play – it also supports MapReduce processing. EMC will be competing with the firms like [Hadoop packager Cloudera](http://redmonk.com/sogrady/2009/10/02/hadoopworld/) \[client\] and [its partners such as IBM](http://redmonk.com/sogrady/2010/03/16/rod-smith-interviews/) \[client\]. Greenplum customers include Linkedin, which uses the system to support its new “People You May Know” function.

There is a grand convergence beginning between NoSQL and distributed cache systems (see [Mike Gualtieri’s Elastic Cache piece](http://blogs.forrester.com/mike_gualtieri)). It seems EMC plans to be a driver, not a fast follower. The Hadoop wave is just about ready to crash onto the enterprise, driven by the likes of EMC and IBM. Chuck Hollis, for example, [points out Greenplum would make a great pre-packaged component VBlock](http://chucksblog.emc.com/chucks_blog/2010/07/emc-to-acquire-greenplum.html) for VMware/EMC/Cisco’s VCE alliance – aimed at customers building private clouds. Of course Cisco is likely to make its own Big Data play anytime soon… That’s the thing with emergent, convergent markets- they sure make partnering hard. But for the customer the cost of analysing some types of data is set to fall by an order of magnitude, while query performance improves by an order of magnitude. Things are getting very interesting indeed.