The answer is "no". The question is "will key/value cloud computing data bases replace relational"? RDBs are very good at modeling your data, but have scaling problems. The cloud computing solutions scale well but are really poor at modeling data; you have to do it yourself. Example problem: you commit an update to a record. Subsequent reads may not see the change until the DB gets around to committing it some time later. Can't run your airlines or banks that way. The eventual solution will have RDB-like modeling with cloud-like scaling, replication, robustness, and all that good stuff. The real solution will allow you to request high data integrity along with I-dont-care-what-happens for storing your RSS feeds or slashdot comments.
RDBs are very good at modeling your data, but have scaling problems.
I think the scaling problems of relational databases a vastly overstated. Right now, using off the shelf kit, you could build a relational database handling 10,000 commits/sec and 100T of data, without using "sharding" or any nonsense like that. Too many people think MySQL and its limitations are representative of this technology; it isn't, not by a long way.
Too many people think MySQL and its limitations are representative of this technology; it isn't, not by a long way.
This is about as true as it is sad, but when your only exposure to relational databases is the least mature and featured one on the market, and everyone is too busy sharding and remaining oblivious about it, this is the kind of discussion you get.
"Sharding" is shorthand for "MySQL's concurrent write performance is abysmal, let's spend more on trying to make it work than it would have cost to just buy Oracle".
Not too mention that sharding probably is one of those things which might seem clever when your database engine doesn't even have hash joins.
I mean, when your DB can't join large result-sets quickly and efficiently in process, why care about the loss of joinability when you spread your data across several servers?
DB2 is free up to something like eight gig of memory usage. I don't understand why people are so wedded to MySQL and Postgres now that DB2 and Oracle have free and very cheap versions.
I can't speak to DB2, but I am really sick of the weird complexity of Oracle. It's power comes at a fairly high cost of complexity, whatever the pricing structure looks like.
It sure is mired in complexity. I am an Oracle programmer by day and once you put it under load there are some pile of bugs that start to show up! That said, Oracle can do so e amazing work, if you have people on hand who can make it do it!
Whoever solves the "scalability or consistency: choose one" problem will make a boatload of money. From what I can tell this is a very hard problem, so software that makes it seamless to operate at multiple levels of this tradeoff curve -- perhaps at the same time, for different parts of your data -- would also be a big win.
I think people are doing it, they are just not slicing it up in a pay-as-you-go format for end users.
The best talk I ever saw in this area was by Paul Strong from eBay (they obviously have a very strong requirement for both scalability and consistency). He talked about how they eventually ripped out all transactions and stored procedures from the RDBMS layers and built their own giant, cross cluster, consistent system to make the zillion customer/auction problem finally managable.
For the less giant problems, I think Amazon could charge a premium for an API-managed RDBMS solution, they have all the tools for making one in place (as opposed to requiring you go build it yourself with EBS and EC2 nodes).
In addition to a cloud-hosted solution, packaged software (along the lines of Hadoop) would be great too. Not everyone wants their corporate data out in the cloud, but I suspect they would all be happy to skip writing one-off "fix RDBMS scalability" layers.
It is a very difficult problem. I worked for a short time for a company that was trying to solve consistency, distributed redundancy, and scalability. It succeded consistency and redundancy, but failed because transaction chatter caused it to fail at scaling. I think that within an application or enterprise schemas will be partitioned among types of database by whether the data must always be consistent(RDB's), are mostly readonly and can be updated leisurely (contact-lists, personnel records), and those that can afford to be inconsistent (archives, web chatter).
I see this as an area where Microsoft could pull out a big win. Imagine if, instead of their half-baked BigTable clone that they just released, they had instead put 1000 or so brains on the problem of distributing a single SQL Server database across N machines.
Scaling out to a dozen DB servers that master/slave their way to scalability is no fun, but it's solved. The problem is that you're currently required to do it yourself. It would rock to be able to outsource that to the cloud.
I want to toss my ASP.NET application up into the Microsoft Cloud, where it will figure out how many webservers it needs to spread itself across and how many database servers it needs to fire up to handle the load it's seeing. And I want it to pretend like it's a single webserver talking to a single DB instance on a single box.
Say what you will about Microsoft, but they have the skills to pull that off. I sure hope they're working on it.
Imagine if, instead of their half-baked BigTable clone that they just released, they had instead put 1000 or so brains on the problem of distributing a single SQL Server database across N machines.
Unfortunately brains don't scale, so you're better off using one huge brain, like Michael Stonebraker.