Yes, very strong agree. I currently use HDF5 on my home machine that sucks in a fairly large amount of data (it's in the tens of TB, and adds 20GB daily). Before I set up HDF5 my database experience was mostly limited to Postgres (although more recently MySQL). I explored a variety of options including Hadoop and Cassandra but I just didn't really want to have more than one node for this exact reason, and I couldn't see a compelling advantage to either without that sort of workflow. Were I not working specifically with timeseries, I probably would have thrown it into postgres.
There's a particular bias I come across where people who genuinely have big data want to set it up in ways that is not necessarily performant because Hadoop is basically the most recognizable tool. If the data processing I'm doing generated a lot of output data I might consider a different flow, but there just isn't much of a reason: most of the data is inert for long periods of time, the output insights are fairly small and the actual processing has to occur very quickly and with low latency.
There's a particular bias I come across where people who genuinely have big data want to set it up in ways that is not necessarily performant because Hadoop is basically the most recognizable tool. If the data processing I'm doing generated a lot of output data I might consider a different flow, but there just isn't much of a reason: most of the data is inert for long periods of time, the output insights are fairly small and the actual processing has to occur very quickly and with low latency.