Most of that data is not the content of the tweet itself, but the metadata associated with it. When I last checked, we were storing about a kilobyte of data for every tweet.
Also, tweets are limited to 140 characters, not bytes - chinese tweets typically take about 200-250 bytes, for example.
I have access to the Twitter gardenhose (which is equal to slightly less than 10% of the full volume). These are the RX and TX statistics from the machine that I've been using to gather data for a few months now:
It works out at around 70GB/day, so I'd actually think that the full firehose would use considerably more data than 400GB/day (likely closer to 800GB).
Woah, 100 GB per second???!!
What exactly do they do?
Edit: Never mind, this is what they are doing
"We would like to develop some kind of ‘google' brain where we can zoom in and out, see it from different perspectives and understand how brain structure and function is related."
Pick a topic you are familiar with. Open up a twitter search for it. Wait a while. See the inevitable storm of tweet-spam that sort of looks like social sharing.
Thousands of fake accounts are tweeting out nonsense all the time. Another example, there are multiple accounts that tweet items from HN (and presumably lots of other rss feeds).
EDIT: That's 11.5k tweets/sec. How do you get eleven thousand people to tweet every second?