Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I've thought about doing this before actually, don't know how they'd feel about someone scraping their content though.


It's not like bulbapedia own the fundamental content - do they place a public license on the wiki pages?

Ideally bulbapedia would provide mediawiki dumps for this, but they don't, and they've gone on record saying they don't intend to. They did leave the mediawiki API open though if you want to crawl a clean rip of each page's wikitext - the default mediawiki API guidelines are also intact, which say that single-threaded crawls should be acceptable in almost all instances, but you should warn the site owner before initiating a multi-threaded scrape.


It looks like all of Bulbapedia's content is licensed under the Creative Commons Attribution-NonCommercial-ShareAlike license. http://bulbapedia.bulbagarden.net/wiki/Bulbapedia:Copyrights




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: