Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

The way I use Twitter is to download tweets to a local database, including images/videos/unshortened URL and then I view the data in my own UI from local DB.

It circumvents the tweet/account removals done by users or Twitter. Also I can do any kind of search/data processing I want.

I guess I'm not alone. You can't hide your data once they are made public.

It's interesting what is possible once you start treating web services as a data source. Sometimes there is much more exposed than is visible or easily consumable on the page and you can filter/search/consume the way you want without all the clutter/feeding algorithms around the original page/service.



How many petabytes does something like this take up?


I download profiles I'm interested in of course. Not everything. 70 profiles with 75000 tweets take 5GB for JSON + images + linked pages content + 60GB for videos.

I didn't optimize for space at all, twitter's json is quite wasteful.


The internet says there are about 500 million tweets per day. If we assume 500 bytes for storing each, this would be 250 Gigabytes per day.


Average tweet JSON object size is 5300B.


Should compress nicely though.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: